PDU fault in rack AMS-B06
전력 및 냉각 — AMS · 대역 외 관리 — AMS
타임라인
-
해결됨
Every machine in rack B06 has been up since 02:52 UTC. Duration: 16 minutes. The PDU is being replaced; the 8 machines that lost power receive the SLA extension automatically. Data on local disks is intact; the machines simply rebooted.
-
원인 확인
The A-side PDU in rack B06 failed with a breaker fault. The 8 machines that rebooted are the ones whose B-side supply was not seated correctly after a recent hardware change. All are powered from the B-side now and are booting.
-
조사 중
8 machines in rack B06 of AMS lost power at 02:36 UTC. Machines in the rack are dual-fed, but these 8 rebooted when one PDU failed. Engineers are at the rack.
사후 검토 ·
발생한 일과 변경 사항
원인
The A-side PDU in AMS-B06 suffered a breaker fault. Every machine in the rack has two power supplies on separate feeds, but on 8 of them the B-side cord had been left partially seated during a hardware change the previous week, so they lost power when the A side went.
영향도
8 machines rebooted and were unavailable for 16 minutes. No data was lost; the machines came back to the state they were installed in, with local disks intact.
변경 사항
- Every hardware change now ends with a power-redundancy test: each feed is dropped in turn and the machine must stay up.
- The faulty PDU model is being phased out across the fleet.
- Term extensions were applied to the 8 machines the same day.
시간은 현지 시간대() 기준으로 표시됩니다.
관련 페이지
- 서비스 상태gpuserver.io의 모든 구성 요소와 데이터센터에 대한 실시간 상태, 다섯 개 도시에서 60초마다 점검, 90일 가동률과 2022년 이후 장애 기록 포함.
- 장애 이력gpuserver.io에서 2022년 9월 모니터링을 시작한 이후 발생한 모든 장애와 점검 시간대, 월별 정리, 각 사건의 타임라인과 사후 검토 포함.
- 서비스 수준 협약99.9% 약정: 무엇이 다운타임인지, 외부에서 어떻게 측정하는지, 손실 1분당 이용 기간 5분을 어떻게 돌려주는지, 그리고 면책 사항 전체.
- 네트워크최대 25 Gbit/s 무제한 포트, 거점마다 회선 사업자 2곳과 IX 1곳, 상시 DDoS 방어, 라우팅된 IPv6 — 그리고 당사가 제공하지 않는 4가지를 먼저 밝힙니다.