PDU fault in rack AMS-B06
電源・冷却 — AMS · アウトオブバンド管理 — AMS
タイムライン
-
解決済み
Every machine in rack B06 has been up since 02:52 UTC. Duration: 16 minutes. The PDU is being replaced; the 8 machines that lost power receive the SLA extension automatically. Data on local disks is intact; the machines simply rebooted.
-
原因特定
The A-side PDU in rack B06 failed with a breaker fault. The 8 machines that rebooted are the ones whose B-side supply was not seated correctly after a recent hardware change. All are powered from the B-side now and are booting.
-
調査中
8 machines in rack B06 of AMS lost power at 02:36 UTC. Machines in the rack are dual-fed, but these 8 rebooted when one PDU failed. Engineers are at the rack.
事後レビュー ·
発生した事象と変更内容
原因
The A-side PDU in AMS-B06 suffered a breaker fault. Every machine in the rack has two power supplies on separate feeds, but on 8 of them the B-side cord had been left partially seated during a hardware change the previous week, so they lost power when the A side went.
影響度
8 machines rebooted and were unavailable for 16 minutes. No data was lost; the machines came back to the state they were installed in, with local disks intact.
変更内容
- Every hardware change now ends with a power-redundancy test: each feed is dropped in turn and the machine must stay up.
- The faulty PDU model is being phased out across the fleet.
- Term extensions were applied to the 8 machines the same day.
時刻はお使いのタイムゾーン()で表示されます。
関連ページ
- サービスステータスgpuserver.ioの全コンポーネントと全データセンターのライブステータス。5都市から60秒ごとに監視プローブで測定し、90日間の稼働率と2022年以降のインシデントログを掲載。
- インシデント履歴gpuserver.ioで2022年9月の監視開始以降に発生した、すべてのインシデントとメンテナンス時間帯を月別に掲載。タイムラインと事後レビューも収録しています。
- サービスレベル契約稼働率99.9%の約束:何がダウンタイムとみなされるか、外部からどう測定するか、失った1分につき契約期間を5分延長する仕組み、そして適用除外の全内容。
- ネットワーク最大25Gbit/sの無制限ポート、拠点ごとに2つのキャリアと1つのIX、常時稼働のDDoSフィルタリング、ルーテッドIPv6。そして、提供しない4つのことを最初に明示します。