Cooling alarm in YUL, row D; some nodes thermally throttled
電源・冷却 — YUL
タイムライン
-
解決済み
Row D is back at target temperature and no GPU is throttling. Duration: 36 minutes. Affected workloads ran slower during the event but were not interrupted.
-
監視中
The tripped unit is back in service and inlet temperatures are falling. GPU clocks are returning to normal.
-
原因特定
One of the two CRAC units serving row D tripped on a pressure alarm. The remaining unit is holding the row a few degrees above target. The facility has an engineer on the unit.
-
調査中
Inlet temperatures in row D of YUL rose above the alert threshold at 20:12 UTC. GPUs on 17 machines in that row have reduced their clocks to stay within limits. Machines remain reachable and workloads continue at lower performance.
時刻はお使いのタイムゾーン()で表示されます。
関連ページ
- サービスステータスgpuserver.ioの全コンポーネントと全データセンターのライブステータス。5都市から60秒ごとに監視プローブで測定し、90日間の稼働率と2022年以降のインシデントログを掲載。
- インシデント履歴gpuserver.ioで2022年9月の監視開始以降に発生した、すべてのインシデントとメンテナンス時間帯を月別に掲載。タイムラインと事後レビューも収録しています。
- サービスレベル契約稼働率99.9%の約束:何がダウンタイムとみなされるか、外部からどう測定するか、失った1分につき契約期間を5分延長する仕組み、そして適用除外の全内容。
- ネットワーク最大25Gbit/sの無制限ポート、拠点ごとに2つのキャリアと1つのIX、常時稼働のDDoSフィルタリング、ルーテッドIPv6。そして、提供しない4つのことを最初に明示します。