Loss of external connectivity in AMS
Réseau — transit et peering — AMS · Gestion hors bande — AMS
Chronologie
-
Résolu
AMS has been fully reachable since 18:00 UTC. Duration: 10 minutes. Workloads kept running; every machine in the site receives the SLA extension automatically.
-
Cause identifiée
A configuration push to the two core switches in AMS was applied in the wrong order and removed the external VLAN from both before re-adding it. The push has been rolled back.
-
Investigation en cours
All machines in AMS became unreachable from outside at 17:50 UTC. Both transit carriers and the exchange are down simultaneously, which points at something inside the site rather than at a carrier.
Revue post-incident ·
Ce qui s’est passé, et ce qui a changé
Cause
A configuration change intended to be applied to one core switch at a time was pushed to both at once by the automation, because the change set was tagged with the site rather than with a single device. The change removed and re-added the external VLAN; with both switches doing it simultaneously there was no path out of the site for 10 minutes.
Impact
Every machine in AMS was unreachable from the internet for 10 minutes. Out-of-band access went down with it because the management network shares the same uplinks. No workload was interrupted and no data was lost.
Ce que nous avons changé
- The automation now refuses to target more than one core device per site in a single change, without exception.
- The out-of-band network in every site now has an independent uplink so that we retain console access during an event of this kind.
- Term extensions were applied to every machine in AMS.
Les heures sont affichées dans votre fuseau horaire ().
Pages liées
- État du serviceÉtat en direct des composants et centres de données gpuserver.io, sondés toutes les 60 s depuis cinq villes : disponibilité sur 90 jours et incidents depuis 2022.
- Historique des incidentsChaque incident et fenêtre de maintenance sur gpuserver.io depuis septembre 2022, mois par mois, avec chronologie et revue post-incident.
- Accord de niveau de serviceL’engagement de 99,9 % : la définition d’une panne, la mesure externe, cinq minutes de durée d’engagement rendues par minute perdue, et les exclusions en détail.
- RéseauPorts sans compteur jusqu’à 25 Gbit/s, deux opérateurs et un IX par site, filtrage DDoS permanent, IPv6 routé — et quatre choses que nous n’offrons pas, en clair.