Loss of external connectivity in AMS
网络——转接与对等互联——AMS · 带外管理——AMS
时间线
-
已解决
AMS has been fully reachable since 18:00 UTC. Duration: 10 minutes. Workloads kept running; every machine in the site receives the SLA extension automatically.
-
已确定原因
A configuration push to the two core switches in AMS was applied in the wrong order and removed the external VLAN from both before re-adding it. The push has been rolled back.
-
调查中
All machines in AMS became unreachable from outside at 17:50 UTC. Both transit carriers and the exchange are down simultaneously, which points at something inside the site rather than at a carrier.
事故复盘 ·
发生了什么,改变了什么
原因
A configuration change intended to be applied to one core switch at a time was pushed to both at once by the automation, because the change set was tagged with the site rather than with a single device. The change removed and re-added the external VLAN; with both switches doing it simultaneously there was no path out of the site for 10 minutes.
影响
Every machine in AMS was unreachable from the internet for 10 minutes. Out-of-band access went down with it because the management network shares the same uplinks. No workload was interrupted and no data was lost.
我们改变了什么
- The automation now refuses to target more than one core device per site in a single change, without exception.
- The out-of-band network in every site now has an independent uplink so that we retain console access during an event of this kind.
- Term extensions were applied to every machine in AMS.
时间以您所在时区()显示。