Provisioning paused: driver package failed validation
Provisioning pipeline
Timeline
-
Resolved
Resolved at 18:02 UTC. Duration: 33 minutes. 4 orders were delivered late; each customer has been told, and the delay is added to their term.
-
Monitoring
Provisioning has resumed and the queue is clearing. Each affected node is being re-provisioned from scratch.
-
Identified
The package was corrupted on our mirror during a sync. We have re-fetched it from the upstream source and its checksum matches.
-
Investigating
New provisioning is paused since 17:29 UTC: the NVIDIA driver package (580.65.06) failed its checksum validation on every node it was applied to. We stop rather than ship a machine with an unverified driver.
Post-incident review ·
What happened, and what changed
Cause
A sync from the upstream driver mirror was interrupted and left a truncated 580.65.06 package with a stale checksum file, which our validation step then rejected on every node.
Impact
Provisioning stopped for 33 minutes. 4 orders were delivered late. No machine was shipped with a bad driver.
What we changed
- Mirror syncs are now atomic: a package and its checksum are published together or not at all.
- The provisioning pipeline retries a validation failure against the upstream source before stopping.
Times are shown in your time zone ().
Related pages
- Service statusLive status of every gpuserver.io component and data centre, probed every 60 seconds from five cities, with 90-day uptime and the incident log since 2022.
- Incident historyEvery incident and maintenance window on gpuserver.io since monitoring began in September 2022, month by month, with its timeline and post-incident review.
- Service level agreementThe 99.9% commitment: what counts as downtime, how it is measured from outside, five minutes of term back per minute lost, and the exclusions in full.
- NetworkUnmetered ports up to 25 Gbit/s, two carriers and an IX per site, always-on DDoS filtering, routed IPv6 — and the four things we do not offer, stated up front.