BMC sensor read errors on H100 PCIe nodes in YUL
Out-of-band management — YUL
Timeline
-
Resolved
Sensor data is complete for every node. Duration: 44 minutes. The fixed firmware is being rolled out during regular maintenance; no host reboot is required.
-
Monitoring
All 6 BMCs have been reset and report cleanly. Watching for recurrence.
-
Identified
A BMC firmware version on 6 nodes has a bug in its sensor polling that surfaces after a long uptime. Resetting the BMC clears it without touching the host; we are doing that node by node.
-
Investigating
The BMCs on some H100 PCIe nodes in YUL are returning sensor read errors since 09:58 UTC. Power and temperature readings in the customer area are missing for those machines. The machines themselves are unaffected.
Times are shown in your time zone ().
Related pages
- Service statusLive status of every gpuserver.io component and data centre, probed every 60 seconds from five cities, with 90-day uptime and the incident log since 2022.
- Incident historyEvery incident and maintenance window on gpuserver.io since monitoring began in September 2022, month by month, with its timeline and post-incident review.
- Service level agreementThe 99.9% commitment: what counts as downtime, how it is measured from outside, five minutes of term back per minute lost, and the exclusions in full.
- NetworkUnmetered ports up to 25 Gbit/s, two carriers and an IX per site, always-on DDoS filtering, routed IPv6 — and the four things we do not offer, stated up front.