Loss of external connectivity in AMS
ネットワーク — トランジット・ピアリング — AMS · アウトオブバンド管理 — AMS
タイムライン
-
解決済み
AMS has been fully reachable since 18:00 UTC. Duration: 10 minutes. Workloads kept running; every machine in the site receives the SLA extension automatically.
-
原因特定
A configuration push to the two core switches in AMS was applied in the wrong order and removed the external VLAN from both before re-adding it. The push has been rolled back.
-
調査中
All machines in AMS became unreachable from outside at 17:50 UTC. Both transit carriers and the exchange are down simultaneously, which points at something inside the site rather than at a carrier.
事後レビュー ·
発生した事象と変更内容
原因
A configuration change intended to be applied to one core switch at a time was pushed to both at once by the automation, because the change set was tagged with the site rather than with a single device. The change removed and re-added the external VLAN; with both switches doing it simultaneously there was no path out of the site for 10 minutes.
影響度
Every machine in AMS was unreachable from the internet for 10 minutes. Out-of-band access went down with it because the management network shares the same uplinks. No workload was interrupted and no data was lost.
変更内容
- The automation now refuses to target more than one core device per site in a single change, without exception.
- The out-of-band network in every site now has an independent uplink so that we retain console access during an event of this kind.
- Term extensions were applied to every machine in AMS.
時刻はお使いのタイムゾーン()で表示されます。
関連ページ
- サービスステータスgpuserver.ioの全コンポーネントと全データセンターのライブステータス。5都市から60秒ごとに監視プローブで測定し、90日間の稼働率と2022年以降のインシデントログを掲載。
- インシデント履歴gpuserver.ioで2022年9月の監視開始以降に発生した、すべてのインシデントとメンテナンス時間帯を月別に掲載。タイムラインと事後レビューも収録しています。
- サービスレベル契約稼働率99.9%の約束:何がダウンタイムとみなされるか、外部からどう測定するか、失った1分につき契約期間を5分延長する仕組み、そして適用除外の全内容。
- ネットワーク最大25Gbit/sの無制限ポート、拠点ごとに2つのキャリアと1つのIX、常時稼働のDDoSフィルタリング、ルーテッドIPv6。そして、提供しない4つのことを最初に明示します。