Top-of-rack switch reboot, rack AMS-E08
ネットワーク — トランジット・ピアリング — AMS
タイムライン
-
解決済み
Rack E08 has been stable since 06:44 UTC. Duration: 7 minutes. Running workloads were not interrupted; only reachability was lost. Affected machines receive the SLA extension automatically.
-
原因特定
The switch has reloaded and its uplinks are up. It rebooted on a software watchdog; we have the crash file and will send it to the vendor.
-
調査中
The top-of-rack switch in rack E08 of AMS rebooted unexpectedly at 06:37 UTC. The 7 machines in that rack are unreachable while it comes back up.
事後レビュー ·
発生した事象と変更内容
原因
The top-of-rack switch in AMS-E08 hit a memory leak in its control-plane process that the vendor has since fixed; the watchdog reloaded the switch when the process stopped responding.
影響度
7 machines were unreachable for 7 minutes. GPU workloads continued to run and no data was lost.
変更内容
- The fixed switch firmware was rolled out to every rack in the fleet during the following maintenance windows.
- The watchdog now reports a warning at 80% memory use instead of only reloading at 100%, giving us time to move the rack to its backup path first.
時刻はお使いのタイムゾーン()で表示されます。
関連ページ
- サービスステータスgpuserver.ioの全コンポーネントと全データセンターのライブステータス。5都市から60秒ごとに監視プローブで測定し、90日間の稼働率と2022年以降のインシデントログを掲載。
- インシデント履歴gpuserver.ioで2022年9月の監視開始以降に発生した、すべてのインシデントとメンテナンス時間帯を月別に掲載。タイムラインと事後レビューも収録しています。
- サービスレベル契約稼働率99.9%の約束:何がダウンタイムとみなされるか、外部からどう測定するか、失った1分につき契約期間を5分延長する仕組み、そして適用除外の全内容。
- ネットワーク最大25Gbit/sの無制限ポート、拠点ごとに2つのキャリアと1つのIX、常時稼働のDDoSフィルタリング、ルーテッドIPv6。そして、提供しない4つのことを最初に明示します。