Skip to content

AWS US-East-1 outage (October 2025): when recovery itself becomes the failure

A DNS race triggered the October 2025 AWS US-East-1 cascade. The recovery overload, with thousands of services reconnecting at once, turned a blip into 15 hours of disruption.

On 19–20 October 2025, a large share of AWS US-East-1 services went down for roughly 15 hours. The trigger was a DNS race condition, a fairly contained event. What turned it into a long, multi-service cascade came next: the recovery itself became the failure.

This post covers how congestive collapse works, why you need to test recovery paths as well as steady state, and how to reproduce the “everything reconnects at once” scenario in a MaxoPerf stress test before it happens to you.

What happened

Early on 19 October 2025, a DNS race condition cleared an internal endpoint table. A core fleet-management service lost contact with its dependencies. Services that relied on that coordination layer began to fail.

Up to this point, the event was serious but bounded. Then engineers started restoring the coordination service, and every compute instance that had lost its lease tried to re-establish it at the same time. The control plane received tens of thousands of re-registration requests at once and saturated. That saturation kept the outage running for 15 hours instead of 15 minutes.

The timeline

  • 19 Oct, ~02:00 UTC: A DNS race empties the endpoint table. The fleet-management service loses connectivity, and downstream AWS services start reporting errors. (AWS postmortem)
  • ~02:30 UTC: Engineers identify the root cause and start restoring the coordination service.
  • ~03:00–17:00 UTC: Recovery attempts stall again and again as the control plane is overwhelmed by simultaneous lease re-establishment from thousands of instances. This is classic congestive collapse.
  • Mitigation steps: Engineers throttled work queues, restarted affected hosts in batches, throttled dependent services, and disabled specific health-check failover behaviors to shrink the reconnection storm.
  • ~17:00 UTC: Services broadly restored after ~15 hours.

Why it happened

The DNS race was the trigger. On its own, it would likely have caused a short disruption. The long outage was a congestive collapse: mass reconnection at the same moment overwhelmed the very system that recovery depended on.

The missing defenses were:

  • No randomized reconnection delay. All instances tried to re-register in the same window. A random, jittered wait before each retry would have spread the burst into a stream the control plane could handle.
  • No admission control on the recovery path. The control plane accepted every reconnection attempt at once instead of queuing or rate-limiting them.
  • No load-shedding strategy. Once the control plane saturated, it had no way to reject the lowest-priority work and keep capacity for the most important re-registrations.

The engineers’ eventual mitigation (batched restarts, throttled work queues) was a manual version of the admission control the system should have had built in.

The failure pattern

This incident is a textbook retry storm / congestive collapse. The risk is not specific to AWS. Any system that coordinates state across a large fleet faces it. The pattern appears whenever:

  1. A dependency failure makes many clients lose their connection or session at the same moment.
  2. Those clients have no randomized backoff, so they all retry at once.
  3. Nothing rate-limits or protects the recovery path, so it takes the full burst.

You don’t need an AWS-scale fleet to hit this. A single microservice that loses its database connection, a set of workers that all lose a queue connection, or a mobile app with millions of users during a brief auth outage can each reproduce the same dynamics at smaller scale.

How it could have been prevented

The defenses are well documented:

Jittered exponential backoff. A client that loses a connection should wait before retrying, and the wait should include a random component. Retries then spread out over time instead of lining up. This alone breaks the burst into a stream the server can handle.

Admission control on recovery paths. Control planes and coordination services should queue or rate-limit incoming reconnection requests. Unbounded concurrent reconnects are the direct cause of congestive collapse.

Load shedding. A saturated system should reject or deprioritize lower-priority work so the most critical operations can proceed. A reconnection storm suits tiered priority: re-register first the services that others depend on.

Batch restarts. AWS’s eventual mitigation was batching instance restarts, a manual form of load shedding. Build it into the system design instead of applying it during an incident, and recovery drops from hours to minutes.

Pre-tested recovery paths. The surest way to know your recovery path holds under reconnection load is to test it before the incident.

How to test for this with MaxoPerf

Simulate a “mass simultaneous reconnect” on your own services and watch whether recovery degrades gracefully or collapses.

Engine: k6 or Taurus (JMeter).

Workload model: Open model. You want to set the arrival rate independently of service response time, because real reconnect storms are driven by time, not by capacity.

Profile (example):

  1. Baseline phase: 50 virtual users for 5 minutes to confirm a clean steady state.
  2. Disruption simulation: drop to 0 VUs for 30 seconds (the outage window in which connections are cut).
  3. Reconnect burst: ramp to 500 VUs within 10 seconds and hold for 5 minutes. This is your “everything reconnects at once” moment.
  4. Observe: does throughput recover, or does it spiral down (congestive collapse)?

Target: The session-management, lease-renewal or coordination endpoint in your own staging environment. Pick the path clients use to re-establish state after a disruption.

Execution locations: Run from ≥2 managed MaxoPerf regions at once to represent clients in different places reconnecting together.

What to read in MaxoPerf results:

  • Error rate at the reconnect burst: does it spike and recover, or stay high?
  • p95/p99 latency during the burst: a large jump that never comes back down means saturation.
  • Throughput plateau: if successful reconnects per second level off far below the incoming rate, the coordination service is backed up.

Run this scenario twice: once without backoff logic in your client, and once with jittered backoff on. The difference in error rate and recovery time is the evidence for your incident-response playbook.

For CI, run a lighter version of this recovery stress test (a shorter burst window and a lower VU count) as a performance gate on every release that touches your reconnection or session code.

Key takeaways

  • The AWS October 2025 outage lasted 15 hours because the recovery path had no protection against mass simultaneous reconnection. The initial DNS fault was not the severe part.
  • Congestive collapse is a predictable failure mode. Any system where many clients share a connection to a coordination layer is at risk.
  • Jittered backoff is non-negotiable. Synchronized retries turn recoverable faults into long outages.
  • Test the recovery path as well as steady state. A service that performs well under normal load may collapse when it has to absorb a reconnection burst, and you won’t know until you test it.
  • Admission control belongs in the system, not in the incident runbook. Batching restarts by hand during an incident is the expensive way to learn this.

Questions this article answers

What caused the AWS US-East-1 outage in October 2025?

A DNS race condition emptied a critical endpoint table, disrupting services. A recovery overload caused the long 15-hour outage: thousands of compute instances tried to re-establish leases at the same time and overwhelmed the control plane in a congestive collapse.

How can teams test for recovery overload and thundering-herd reconnects?

Use a stress test that models mass simultaneous reconnection. Ramp load fast after a simulated disruption and watch whether throughput degrades gracefully or collapses. Add jittered exponential backoff and admission-control logic, then verify the recovery path holds under that load.

What is congestive collapse in distributed systems?

Congestive collapse is when a system recovering from a fault receives more reconnection or retry traffic than it can handle, causing it to fail again. Each failed reconnection triggers another attempt, creating a feedback loop that prevents stable recovery.