Loss of external connectivity in AMS
네트워크 — 트랜싯 및 피어링 — AMS · 대역 외 관리 — AMS
타임라인
-
해결됨
AMS has been fully reachable since 18:00 UTC. Duration: 10 minutes. Workloads kept running; every machine in the site receives the SLA extension automatically.
-
원인 확인
A configuration push to the two core switches in AMS was applied in the wrong order and removed the external VLAN from both before re-adding it. The push has been rolled back.
-
조사 중
All machines in AMS became unreachable from outside at 17:50 UTC. Both transit carriers and the exchange are down simultaneously, which points at something inside the site rather than at a carrier.
사후 검토 ·
발생한 일과 변경 사항
원인
A configuration change intended to be applied to one core switch at a time was pushed to both at once by the automation, because the change set was tagged with the site rather than with a single device. The change removed and re-added the external VLAN; with both switches doing it simultaneously there was no path out of the site for 10 minutes.
영향도
Every machine in AMS was unreachable from the internet for 10 minutes. Out-of-band access went down with it because the management network shares the same uplinks. No workload was interrupted and no data was lost.
변경 사항
- The automation now refuses to target more than one core device per site in a single change, without exception.
- The out-of-band network in every site now has an independent uplink so that we retain console access during an event of this kind.
- Term extensions were applied to every machine in AMS.
시간은 현지 시간대() 기준으로 표시됩니다.
관련 페이지
- 서비스 상태gpuserver.io의 모든 구성 요소와 데이터센터에 대한 실시간 상태, 다섯 개 도시에서 60초마다 점검, 90일 가동률과 2022년 이후 장애 기록 포함.
- 장애 이력gpuserver.io에서 2022년 9월 모니터링을 시작한 이후 발생한 모든 장애와 점검 시간대, 월별 정리, 각 사건의 타임라인과 사후 검토 포함.
- 서비스 수준 협약99.9% 약정: 무엇이 다운타임인지, 외부에서 어떻게 측정하는지, 손실 1분당 이용 기간 5분을 어떻게 돌려주는지, 그리고 면책 사항 전체.
- 네트워크최대 25 Gbit/s 무제한 포트, 거점마다 회선 사업자 2곳과 IX 1곳, 상시 DDoS 방어, 라우팅된 IPv6 — 그리고 당사가 제공하지 않는 4가지를 먼저 밝힙니다.