Top-of-rack switch reboot, rack ZRH-B01
Network — transit & peering — ZRH
Timeline
-
Resolved
Rack B01 has been stable since 09:51 UTC. Duration: 7 minutes. Running workloads were not interrupted; only reachability was lost. Affected machines receive the SLA extension automatically.
-
Identified
The switch has reloaded and its uplinks are up. It rebooted on a software watchdog; we have the crash file and will send it to the vendor.
-
Investigating
The top-of-rack switch in rack B01 of ZRH rebooted unexpectedly at 09:44 UTC. The 13 machines in that rack are unreachable while it comes back up.
Post-incident review ·
What happened, and what changed
Cause
The top-of-rack switch in ZRH-B01 hit a memory leak in its control-plane process that the vendor has since fixed; the watchdog reloaded the switch when the process stopped responding.
Impact
13 machines were unreachable for 7 minutes. GPU workloads continued to run and no data was lost.
What we changed
- The fixed switch firmware was rolled out to every rack in the fleet during the following maintenance windows.
- The watchdog now reports a warning at 80% memory use instead of only reloading at 100%, giving us time to move the rack to its backup path first.
Times are shown in your time zone ().
Related pages
- Service statusLive status of every gpuserver.io component and data centre, probed every 60 seconds from five cities, with 90-day uptime and the incident log since 2022.
- Incident historyEvery incident and maintenance window on gpuserver.io since monitoring began in September 2022, month by month, with its timeline and post-incident review.
- Service level agreementThe 99.9% commitment: what counts as downtime, how it is measured from outside, five minutes of term back per minute lost, and the exclusions in full.
- NetworkUnmetered ports up to 25 Gbit/s, two carriers and an IX per site, always-on DDoS filtering, routed IPv6 — and the four things we do not offer, stated up front.