Top-of-rack switch reboot, rack SIN-A04
Red — tránsito y peering — SIN
Cronología
-
Resuelto
Rack A04 has been stable since 12:16 UTC. Duration: 8 minutes. Running workloads were not interrupted; only reachability was lost. Affected machines receive the SLA extension automatically.
-
Causa identificada
The switch has reloaded and its uplinks are up. It rebooted on a software watchdog; we have the crash file and will send it to the vendor.
-
Investigando
The top-of-rack switch in rack A04 of SIN rebooted unexpectedly at 12:08 UTC. The 7 machines in that rack are unreachable while it comes back up.
Revisión posterior al incidente ·
Qué ocurrió y qué cambió
Causa
The top-of-rack switch in SIN-A04 hit a memory leak in its control-plane process that the vendor has since fixed; the watchdog reloaded the switch when the process stopped responding.
Impacto
7 machines were unreachable for 8 minutes. GPU workloads continued to run and no data was lost.
Qué cambiamos
- The fixed switch firmware was rolled out to every rack in the fleet during the following maintenance windows.
- The watchdog now reports a warning at 80% memory use instead of only reloading at 100%, giving us time to move the rack to its backup path first.
Las horas se muestran en su zona horaria ().
Páginas relacionadas
- Estado del servicioEstado en vivo de gpuserver.io: componentes y centros de datos sondeados cada 60 s desde 5 ciudades, con disponibilidad de 90 días e incidentes desde 2022.
- Historial de incidentesTodos los incidentes y ventanas de mantenimiento de gpuserver.io desde septiembre de 2022, mes a mes, con cronología y revisión posterior al incidente.
- Acuerdo de nivel de servicioEl compromiso del 99,9 %: qué cuenta como inactividad, cómo se mide desde fuera, cinco minutos de plazo por cada minuto perdido, y las exclusiones completas.
- RedPuertos sin contador hasta 25 Gbit/s, dos operadores y un IX por sede, filtrado DDoS permanente, IPv6 enrutado — y las cuatro cosas que no ofrecemos, dichas ya.