Top-of-rack switch reboot, rack SIN-A04
ネットワーク — トランジット・ピアリング — SIN
タイムライン
-
解決済み
Rack A04 has been stable since 12:16 UTC. Duration: 8 minutes. Running workloads were not interrupted; only reachability was lost. Affected machines receive the SLA extension automatically.
-
原因特定
The switch has reloaded and its uplinks are up. It rebooted on a software watchdog; we have the crash file and will send it to the vendor.
-
調査中
The top-of-rack switch in rack A04 of SIN rebooted unexpectedly at 12:08 UTC. The 7 machines in that rack are unreachable while it comes back up.
事後レビュー ·
発生した事象と変更内容
原因
The top-of-rack switch in SIN-A04 hit a memory leak in its control-plane process that the vendor has since fixed; the watchdog reloaded the switch when the process stopped responding.
影響度
7 machines were unreachable for 8 minutes. GPU workloads continued to run and no data was lost.
変更内容
- The fixed switch firmware was rolled out to every rack in the fleet during the following maintenance windows.
- The watchdog now reports a warning at 80% memory use instead of only reloading at 100%, giving us time to move the rack to its backup path first.
時刻はお使いのタイムゾーン()で表示されます。
関連ページ
- サービスステータスgpuserver.ioの全コンポーネントと全データセンターのライブステータス。5都市から60秒ごとに監視プローブで測定し、90日間の稼働率と2022年以降のインシデントログを掲載。
- インシデント履歴gpuserver.ioで2022年9月の監視開始以降に発生した、すべてのインシデントとメンテナンス時間帯を月別に掲載。タイムラインと事後レビューも収録しています。
- サービスレベル契約稼働率99.9%の約束:何がダウンタイムとみなされるか、外部からどう測定するか、失った1分につき契約期間を5分延長する仕組み、そして適用除外の全内容。
- ネットワーク最大25Gbit/sの無制限ポート、拠点ごとに2つのキャリアと1つのIX、常時稼働のDDoSフィルタリング、ルーテッドIPv6。そして、提供しない4つのことを最初に明示します。