Cooling alarm in YUL, row D; some nodes thermally throttled
전력 및 냉각 — YUL
타임라인
-
해결됨
Row D is back at target temperature and no GPU is throttling. Duration: 36 minutes. Affected workloads ran slower during the event but were not interrupted.
-
모니터링 중
The tripped unit is back in service and inlet temperatures are falling. GPU clocks are returning to normal.
-
원인 확인
One of the two CRAC units serving row D tripped on a pressure alarm. The remaining unit is holding the row a few degrees above target. The facility has an engineer on the unit.
-
조사 중
Inlet temperatures in row D of YUL rose above the alert threshold at 20:12 UTC. GPUs on 17 machines in that row have reduced their clocks to stay within limits. Machines remain reachable and workloads continue at lower performance.
시간은 현지 시간대() 기준으로 표시됩니다.
관련 페이지
- 서비스 상태gpuserver.io의 모든 구성 요소와 데이터센터에 대한 실시간 상태, 다섯 개 도시에서 60초마다 점검, 90일 가동률과 2022년 이후 장애 기록 포함.
- 장애 이력gpuserver.io에서 2022년 9월 모니터링을 시작한 이후 발생한 모든 장애와 점검 시간대, 월별 정리, 각 사건의 타임라인과 사후 검토 포함.
- 서비스 수준 협약99.9% 약정: 무엇이 다운타임인지, 외부에서 어떻게 측정하는지, 손실 1분당 이용 기간 5분을 어떻게 돌려주는지, 그리고 면책 사항 전체.
- 네트워크최대 25 Gbit/s 무제한 포트, 거점마다 회선 사업자 2곳과 IX 1곳, 상시 DDoS 방어, 라우팅된 IPv6 — 그리고 당사가 제공하지 않는 4가지를 먼저 밝힙니다.