All 6 data centres operational

Paid in crypto · No identity check · Root access in under 5 minutes

Major Resolved Review published

Top-of-rack switch reboot, rack SIN-A04

Network — transit & peering — SIN

Timeline

  1. Resolved

    Rack A04 has been stable since 12:16 UTC. Duration: 8 minutes. Running workloads were not interrupted; only reachability was lost. Affected machines receive the SLA extension automatically.

  2. Identified

    The switch has reloaded and its uplinks are up. It rebooted on a software watchdog; we have the crash file and will send it to the vendor.

  3. Investigating

    The top-of-rack switch in rack A04 of SIN rebooted unexpectedly at 12:08 UTC. The 7 machines in that rack are unreachable while it comes back up.

Post-incident review ·

What happened, and what changed

Cause

The top-of-rack switch in SIN-A04 hit a memory leak in its control-plane process that the vendor has since fixed; the watchdog reloaded the switch when the process stopped responding.

Impact

7 machines were unreachable for 8 minutes. GPU workloads continued to run and no data was lost.

What we changed

  • The fixed switch firmware was rolled out to every rack in the fleet during the following maintenance windows.
  • The watchdog now reports a warning at 80% memory use instead of only reloading at 100%, giving us time to move the rack to its backup path first.

Sign in

Console, invoices and out-of-band access.

No account yet?

There is no separate sign-up. Your account is created while you place your first order — you choose the email and the password on the payment step, and the console is open by the time the machine is.

Configure a server

Language