Top-of-rack switch reboot, rack AMS-E08
Rete — transito e peering — AMS
Cronologia
-
Risolto
Rack E08 has been stable since 06:44 UTC. Duration: 7 minutes. Running workloads were not interrupted; only reachability was lost. Affected machines receive the SLA extension automatically.
-
Causa individuata
The switch has reloaded and its uplinks are up. It rebooted on a software watchdog; we have the crash file and will send it to the vendor.
-
In analisi
The top-of-rack switch in rack E08 of AMS rebooted unexpectedly at 06:37 UTC. The 7 machines in that rack are unreachable while it comes back up.
Analisi post-incidente ·
Cosa è successo e cosa è cambiato
Causa
The top-of-rack switch in AMS-E08 hit a memory leak in its control-plane process that the vendor has since fixed; the watchdog reloaded the switch when the process stopped responding.
Impatto
7 machines were unreachable for 7 minutes. GPU workloads continued to run and no data was lost.
Cosa abbiamo cambiato
- The fixed switch firmware was rolled out to every rack in the fleet during the following maintenance windows.
- The watchdog now reports a warning at 80% memory use instead of only reloading at 100%, giving us time to move the rack to its backup path first.
Gli orari sono mostrati nel fuso orario locale ().
Pagine correlate
- Stato del servizioStato di ogni componente e data center gpuserver.io, sondati ogni 60 secondi da cinque città: disponibilità su 90 giorni e registro incidenti dal 2022.
- Cronologia incidentiOgni incidente e finestra di manutenzione su gpuserver.io dall’inizio del monitoraggio, a settembre 2022, mese per mese, con cronologia e analisi post-incidente.
- Accordo sul livello di servizioL’impegno del 99,9%: cosa conta come downtime, come viene misurato dall’esterno, cinque minuti di durata in più per ogni minuto perso, e le esclusioni per intero.
- RetePorte senza contatore fino a 25 Gbit/s, due carrier e un IX per sede, DDoS sempre filtrato, IPv6 instradato — e quattro cose che non offriamo, dette subito.