Top-of-rack switch reboot, rack ZRH-B01
Ağ — transit ve peering — ZRH
Zaman çizelgesi
-
Çözüldü
Rack B01 has been stable since 09:51 UTC. Duration: 7 minutes. Running workloads were not interrupted; only reachability was lost. Affected machines receive the SLA extension automatically.
-
Neden belirlendi
The switch has reloaded and its uplinks are up. It rebooted on a software watchdog; we have the crash file and will send it to the vendor.
-
İnceleniyor
The top-of-rack switch in rack B01 of ZRH rebooted unexpectedly at 09:44 UTC. The 13 machines in that rack are unreachable while it comes back up.
Olay sonrası inceleme ·
Ne oldu ve ne değişti
Neden
The top-of-rack switch in ZRH-B01 hit a memory leak in its control-plane process that the vendor has since fixed; the watchdog reloaded the switch when the process stopped responding.
Etki
13 machines were unreachable for 7 minutes. GPU workloads continued to run and no data was lost.
Neyi değiştirdik
- The fixed switch firmware was rolled out to every rack in the fleet during the following maintenance windows.
- The watchdog now reports a warning at 80% memory use instead of only reloading at 100%, giving us time to move the rack to its backup path first.
Saatler, saat diliminizde () gösterilir.
İlgili sayfalar
- Hizmet durumugpuserver.io'nun her bileşeni ve veri merkezi, beş şehirden 60 saniyede bir yoklanır; 90 günlük kullanılabilirlik ve 2022'den beri olay logu burada yer alır.
- Olay geçmişigpuserver.io'da izlemenin başladığı Eylül 2022'den bu yana yaşanan her olay ve bakım penceresi, ay ay, zaman çizelgesi ve olay sonrası incelemesiyle birlikte.
- Hizmet seviyesi sözleşmesi%99,9 taahhüdü: neyin kesinti sayıldığı, dışarıdan nasıl ölçüldüğü, kaybedilen her dakika için beş dakika süre iadesi ve istisnaların tamamı.
- Ağ25 Gbit/s'e kadar sayaçsız portlar, site başına iki operatör ve bir IX, sürekli DDoS filtreleme, yönlendirilen IPv6 — ve baştan belirtilen, sunmadığımız dört şey.