Top-of-rack switch reboot, rack AMS-E08
Réseau — transit et peering — AMS
Chronologie
-
Résolu
Rack E08 has been stable since 06:44 UTC. Duration: 7 minutes. Running workloads were not interrupted; only reachability was lost. Affected machines receive the SLA extension automatically.
-
Cause identifiée
The switch has reloaded and its uplinks are up. It rebooted on a software watchdog; we have the crash file and will send it to the vendor.
-
Investigation en cours
The top-of-rack switch in rack E08 of AMS rebooted unexpectedly at 06:37 UTC. The 7 machines in that rack are unreachable while it comes back up.
Revue post-incident ·
Ce qui s’est passé, et ce qui a changé
Cause
The top-of-rack switch in AMS-E08 hit a memory leak in its control-plane process that the vendor has since fixed; the watchdog reloaded the switch when the process stopped responding.
Impact
7 machines were unreachable for 7 minutes. GPU workloads continued to run and no data was lost.
Ce que nous avons changé
- The fixed switch firmware was rolled out to every rack in the fleet during the following maintenance windows.
- The watchdog now reports a warning at 80% memory use instead of only reloading at 100%, giving us time to move the rack to its backup path first.
Les heures sont affichées dans votre fuseau horaire ().
Pages liées
- État du serviceÉtat en direct des composants et centres de données gpuserver.io, sondés toutes les 60 s depuis cinq villes : disponibilité sur 90 jours et incidents depuis 2022.
- Historique des incidentsChaque incident et fenêtre de maintenance sur gpuserver.io depuis septembre 2022, mois par mois, avec chronologie et revue post-incident.
- Accord de niveau de serviceL’engagement de 99,9 % : la définition d’une panne, la mesure externe, cinq minutes de durée d’engagement rendues par minute perdue, et les exclusions en détail.
- RéseauPorts sans compteur jusqu’à 25 Gbit/s, deux opérateurs et un IX par site, filtrage DDoS permanent, IPv6 routé — et quatre choses que nous n’offrons pas, en clair.