Top-of-rack switch reboot, rack SIN-A04
Netzwerk — Transit & Peering — SIN
Zeitleiste
-
Behoben
Rack A04 has been stable since 12:16 UTC. Duration: 8 minutes. Running workloads were not interrupted; only reachability was lost. Affected machines receive the SLA extension automatically.
-
Ursache erkannt
The switch has reloaded and its uplinks are up. It rebooted on a software watchdog; we have the crash file and will send it to the vendor.
-
Untersuchung läuft
The top-of-rack switch in rack A04 of SIN rebooted unexpectedly at 12:08 UTC. The 7 machines in that rack are unreachable while it comes back up.
Nachbesprechung des Vorfalls ·
Was passiert ist und was sich geändert hat
Ursache
The top-of-rack switch in SIN-A04 hit a memory leak in its control-plane process that the vendor has since fixed; the watchdog reloaded the switch when the process stopped responding.
Auswirkung
7 machines were unreachable for 8 minutes. GPU workloads continued to run and no data was lost.
Was wir geändert haben
- The fixed switch firmware was rolled out to every rack in the fleet during the following maintenance windows.
- The watchdog now reports a warning at 80% memory use instead of only reloading at 100%, giving us time to move the rack to its backup path first.
Zeiten werden in Ihrer Zeitzone angezeigt ().
Verwandte Seiten
- ServicestatusLive-Status jeder gpuserver.io-Komponente und jedes Rechenzentrums, alle 60 Sekunden aus fünf Städten gemessen, mit 90-Tage-Verfügbarkeit und Protokoll seit 2022.
- VorfallshistorieJeder Vorfall und jedes Wartungsfenster auf gpuserver.io seit September 2022, mit Zeitleiste und Nachbesprechung des Vorfalls.
- Vereinbarung zum ServicelevelDie 99,9-%-Zusage: was als Ausfallzeit zählt, wie sie von außen gemessen wird, fünf Minuten Laufzeit zurück je verlorener Minute, und die Ausschlüsse vollständig.
- NetzwerkPorts ohne Volumenbegrenzung bis 25 Gbit/s, zwei Carrier und ein IX pro Standort, ständige DDoS-Filterung, geroutetes IPv6 — und vier Dinge, die wir nicht anbieten.