Cooling alarm in YUL, row D; some nodes thermally throttled
Power & cooling — YUL
Timeline
-
Resolved
Row D is back at target temperature and no GPU is throttling. Duration: 36 minutes. Affected workloads ran slower during the event but were not interrupted.
-
Monitoring
The tripped unit is back in service and inlet temperatures are falling. GPU clocks are returning to normal.
-
Identified
One of the two CRAC units serving row D tripped on a pressure alarm. The remaining unit is holding the row a few degrees above target. The facility has an engineer on the unit.
-
Investigating
Inlet temperatures in row D of YUL rose above the alert threshold at 20:12 UTC. GPUs on 17 machines in that row have reduced their clocks to stay within limits. Machines remain reachable and workloads continue at lower performance.
Times are shown in your time zone ().
Related pages
- Service statusLive status of every gpuserver.io component and data centre, probed every 60 seconds from five cities, with 90-day uptime and the incident log since 2022.
- Incident historyEvery incident and maintenance window on gpuserver.io since monitoring began in September 2022, month by month, with its timeline and post-incident review.
- Service level agreementThe 99.9% commitment: what counts as downtime, how it is measured from outside, five minutes of term back per minute lost, and the exclusions in full.
- NetworkUnmetered ports up to 25 Gbit/s, two carriers and an IX per site, always-on DDoS filtering, routed IPv6 — and the four things we do not offer, stated up front.