Loss of external connectivity in AMS
Rete — transito e peering — AMS · Gestione fuori banda — AMS
Cronologia
-
Risolto
AMS has been fully reachable since 18:00 UTC. Duration: 10 minutes. Workloads kept running; every machine in the site receives the SLA extension automatically.
-
Causa individuata
A configuration push to the two core switches in AMS was applied in the wrong order and removed the external VLAN from both before re-adding it. The push has been rolled back.
-
In analisi
All machines in AMS became unreachable from outside at 17:50 UTC. Both transit carriers and the exchange are down simultaneously, which points at something inside the site rather than at a carrier.
Analisi post-incidente ·
Cosa è successo e cosa è cambiato
Causa
A configuration change intended to be applied to one core switch at a time was pushed to both at once by the automation, because the change set was tagged with the site rather than with a single device. The change removed and re-added the external VLAN; with both switches doing it simultaneously there was no path out of the site for 10 minutes.
Impatto
Every machine in AMS was unreachable from the internet for 10 minutes. Out-of-band access went down with it because the management network shares the same uplinks. No workload was interrupted and no data was lost.
Cosa abbiamo cambiato
- The automation now refuses to target more than one core device per site in a single change, without exception.
- The out-of-band network in every site now has an independent uplink so that we retain console access during an event of this kind.
- Term extensions were applied to every machine in AMS.
Gli orari sono mostrati nel fuso orario locale ().
Pagine correlate
- Stato del servizioStato di ogni componente e data center gpuserver.io, sondati ogni 60 secondi da cinque città: disponibilità su 90 giorni e registro incidenti dal 2022.
- Cronologia incidentiOgni incidente e finestra di manutenzione su gpuserver.io dall’inizio del monitoraggio, a settembre 2022, mese per mese, con cronologia e analisi post-incidente.
- Accordo sul livello di servizioL’impegno del 99,9%: cosa conta come downtime, come viene misurato dall’esterno, cinque minuti di durata in più per ogni minuto perso, e le esclusioni per intero.
- RetePorte senza contatore fino a 25 Gbit/s, due carrier e un IX per sede, DDoS sempre filtrato, IPv6 instradato — e quattro cose che non offriamo, dette subito.