Skip to content

A year of load failures: what broke, why, and how to test for it

Twenty real outages, ten failure patterns, one year. A guided tour of what load does to software in production, and how to test before it happens to you.

Between June 2025 and June 2026, load broke things again and again, at great cost, in patterns any team that ships software to the internet should recognise. Game launches, Black Friday, a live premiere, a government results portal, a payday window: each time, a real system that had passed normal testing collapsed under traffic that was, in hindsight, predictable.

This post looks back on that year. It collects twenty incidents, groups them under ten failure patterns, and maps each pattern to the load test that would have shown the risk before users found it. The companies range from government agencies to large cloud providers, and the lessons apply well beyond them.

If you are new to performance testing, read this as a field guide to what goes wrong and why. If you already run load tests, use it as a checklist and confirm your tests cover all ten patterns.

The ten failure classes

Every outage in this series falls into one of these ten categories. Each links to the academy test type that best addresses it.

Flash-sale concurrency. A small inventory window sends a rush of buyers to checkout at the same time and overwhelms the checkout and inventory-decrement path. No queue or rate limit protects the write path. → Spike test

Scheduled-release spike. Everyone knows the exact release minute. Without pre-orders to spread demand, all buyers arrive at once and pack the whole purchase or authentication flow into seconds. → Spike test

Thundering herd. A large population of users tries to log in or connect at the same moment, at a launch or right after a brief outage ends. The auth layer, connection pools or session store saturate before anything can warm up. → Spike test · Breakpoint/capacity test

Live-event peak. A broadcast or sporting event puts a hard concurrency cliff at kickoff with no time-shift: every viewer hits play in the same second. The system has to sustain that peak as well as survive the ramp. → Spike test · Soak/endurance test

Soak-growth collapse. The system handles the launch spike, then degrades as concurrency keeps climbing over hours or days. Slow resource leaks, pool exhaustion or cache pressure cause it, and a short test misses all three. → Soak/endurance test

Retry storm. A brief disruption makes clients reconnect at the same time. If the recovery path lacks backoff and jitter, the reconnection wave acts as a denial-of-service attack on the systems trying to recover. → Stress test

Capacity ceiling. Traffic grows past the system’s real maximum because autoscaling was not configured, replicas fell behind or load shedding was missing. Nobody measured the ceiling, so nobody raised it. → Breakpoint/capacity test · Scalability test

Dependency contention. A slow or degraded dependency (a database, a storage layer, an authentication backend) causes lock contention or connection exhaustion that spreads into the wider system. → Stress test

Config-limit cliff. An oversized input hits a hard-coded size, feature-count or quota limit. The system does not degrade gracefully. It panics or crashes outright. → Configuration test · Volume test

Seasonal demand spike. A known calendar event (tax season, exam results, payday) drives predictable peak traffic the system was not sized for, often because nobody tested that volume before the event. → Spike test · Volume test

The incident index

Flash-sale concurrency

C01: The Odyssey IMAX onsale crash. Demand for a limited run of 70mm screens sent AMC’s site into hour-long virtual queues and Fandango into error messages. AMC’s CEO called it the highest first-day studio-film sales since 2022 and said it all landed in a single buy window.

Scheduled-release spike

C05: Hollow Knight: Silksong storefront crash. With no pre-orders, 535,000 Steam players hit the purchase and authentication flow in the launch minute and crashed Steam, the PlayStation Store, eShop and Xbox at once.

C07: Stranger Things S5 premiere outage. A synchronised global play-start spike on Netflix outran resource allocation for roughly twenty minutes. The co-creator noted that a 30% bandwidth increase still was not enough.

Thundering herd

C03: Shopify Cyber Monday auth outage. On the busiest merchant day of the year, Shopify’s login, admin, POS and API login flows went down for around six hours. A shared authentication chokepoint failed under simultaneous merchant load.

C04: Battlefield 6 launch queues. Queues of up to 500,000 players formed at launch when simultaneous login demand outran provisioned capacity. EA scaled live and added admission queues to protect the experience.

C18: SSA “My Social Security” portal crashes. A new anti-fraud check earlier in the auth flow sent many more users into authentication at once. Public reporting says the system was not tested at high user volume before launch, which is the textbook missing load test.

Live-event peak

C08: Netflix NFL Christmas streaming issues. Buffering, quality drops and casting failures under huge concurrent live-sports load. Any system that must sustain a peak across millions of viewers at once, with no time-shift, faces this.

Soak-growth collapse

C06: ARC Raiders server overload. The game passed a pre-launch server stress test and ran clean for several days. Then login queues, matchmaking and voice broke as concurrency kept climbing. Only a long soak test shows this kind of gradual degradation.

Retry storm

C09: AWS US-EAST-1 cascade. A DNS issue set off a mass reconnection wave that overloaded the EC2 control plane in what the AWS postmortem called “congestive collapse”. The recovery itself became the load event.

C11: Google Cloud us-central1 retry flood. A crash-looping service was stabilised globally in about forty minutes. Recovery in the us-central1 region took three hours because clients retried without randomised exponential backoff and kept reloading the datastore.

Capacity ceiling

C12: GitHub sustained capacity failures. A months-long cluster of outages, attributed largely to traffic growing faster than capacity. Database read replicas fell behind under peak read load, and some services could not shed load from high-volume clients.

C14: Slack routing-config scale limit. Infrastructure growth quietly outgrew static routing configurations. Routing updates stopped reaching the web layer, and clients lost live routing data. Nobody measured the ceiling until they hit it.

C15: OpenAI/ChatGPT extended outage. Over ten hours, across many service components. No detailed public root-cause analysis was published, but AI inference capacity under heavy demand was cited as a contributing factor. It illustrates the need for headroom and graceful 429 shedding.

Dependency contention

C13: Clerk Cloud SQL migration contention. An unannounced live database migration spiked storage latency, increased lock contention, saturated compute and caused a wave of 429s until the migration finished. It was the second such failure in seven months.

C19: Fiserv/Zelle banking outage. A change at a shared core-banking platform spread across hundreds of institutions and more than sixty apps, including Zelle. Accounts locked and transfers failed during a Friday payday window.

Config-limit cliff

C10: Cloudflare Bot Management crash. A database configuration change made a data file grow to more than twice its expected size, past a hard feature-count limit. The proxy panicked on the oversized input instead of degrading gracefully, and Cloudflare’s core proxy went down globally for around six hours.

Seasonal demand spike

C02: Best Buy Black Friday crash. The site and app buckled under the morning surge, and Downdetector logged over 1,800 complaints. No official root cause was published, but the timing matched the peak shopping window exactly.

C16: NEET UG 2025 results-day overload. Over two million candidates tried to download scorecards at the same time on results day. The portal returned 503s, blank pages and failed logins. The spike was predictable and could have been sized and tested weeks in advance.

C17: IRS “Where’s My Refund?” outage. The site went down in the middle of the filing surge, in what was officially described as a maintenance window. Even planned downtime has to account for seasonal concurrency.

C20: UK payday banking outages. Several UK bank apps failed on a payday Friday. Experts pointed to shared third-party dependencies under month-end peak transaction volume: a predictable calendar spike met fragile dependency handling.

How to start testing for these patterns

The incidents above cluster around a handful of testing gaps. Here is how to close each one.

If you have never load-tested before: start with a load test at your expected peak concurrency. Hold it long enough to see steady-state behaviour, at least fifteen to thirty minutes. If nothing breaks, you have a baseline. If something breaks before you reach peak, you have your first finding.

For spike and flash-sale scenarios: run a spike test. Ramp from baseline to several times peak almost instantly, hold briefly, then drop. Note where the error rate leaves zero and where latency crosses your threshold. Target the path under pressure (checkout, authentication or inventory write), not only the homepage.

For thundering-herd and auth-under-load: spike the login and session path itself, not only read traffic. Pair it with a breakpoint test to find the ceiling. If you are shipping a new auth flow, test it at the volume of its first deployment before that deployment.

For slow-burn and soak-growth scenarios: run a soak/endurance test at and beyond projected peak for hours, not minutes. A test that passes in five minutes may fail at the two-hour mark as connection pools run out or memory pressure builds.

For retry-storm and recovery scenarios: stress test the reconnection path. Model many clients reconnecting at once and watch whether throughput collapses or degrades gracefully. Check that your backoff has jitter so reconnection waves do not arrive in lockstep.

For capacity ceiling and scaling gaps: run a scalability test. Step load up until you find the knee of the curve. Record the maximum sustainable throughput and which subsystem fails first. Schedule recurring runs so you know when growing traffic nears the old ceiling.

For dependency-contention scenarios: stress test while the dependency is degraded. In a staging environment, inject extra database or storage latency and note when lock contention and error rates rise.

For config-limit and boundary scenarios: run configuration and volume tests with oversized inputs: payload sizes, feature counts or record volumes at and above the documented limit. Confirm the system returns a clean error instead of crashing.

For seasonal and calendar peaks: size your spike to the known concurrency of the event, which is the population multiplied by your concurrency assumption, plus headroom. Run it several weeks before the date so you have time to act on the results. Rerun it closer to the event to catch changes since the first test.

MaxoPerf runs all of these patterns with Taurus, JMeter or k6 scenarios from managed cloud locations or private runners. You review results and run artifacts in a shared workspace and schedule recurring runs, so load testing becomes a routine instead of a pre-launch scramble. To attach a load gate to your release pipeline, start with the use-cases/ci-cd-performance-gates/ page.

Key takeaways

  • Traffic peaks are usually predictable. Black Friday, a game launch, an exam results page, a live premiere: each was on a calendar. The systems that failed were not tested at the traffic level the calendar implied.
  • The test must target the right path. Many teams load-test the read path and miss the auth or write path. Several incidents in this series involved authentication chokepoints that storefront tests would never have found.
  • Short tests miss slow degradation. A five-minute test can pass while a soak over hours shows memory leaks, pool exhaustion or queue back-pressure that only appear under sustained load.
  • Recovery is a load event. In two of the most severe outages here (the AWS and Google Cloud cases), the retry storm after the initial failure is what made the outage long. If clients reconnect without backoff, recovery can cause another outage.
  • Measure hard limits before production finds them. The Cloudflare incident is a clean example of a hard-coded ceiling that nobody tested against. Configuration and volume tests exist to surface these limits first.

Each of the twenty incidents above links to its full story: what happened, why it happened, and the test recipe that would have found the risk first.

Questions this article answers

What causes load-related outages?

Most load-related outages share a root cause: the system was never tested at the concurrency or throughput it actually received. That gap shows up as a spike that overwhelms a shared resource, a slow dependency that locks up under pressure, or a hard-coded limit that the team did not know existed.

How do you test for a traffic spike?

Run a spike test. Start at a comfortable baseline, ramp to many times that level almost instantly, hold for a few minutes, then drop. Watch where error rate departs from zero and where p95 latency crosses your threshold. Do this against a staging environment that mirrors production scaling behaviour.

What is a thundering herd?

A thundering herd happens when a large number of clients all try to connect or authenticate at the same moment, typically at a product launch or after a brief outage ends. The concentrated burst hits shared infrastructure (auth services, connection pools, caches) much harder than a smooth ramp would, because every client arrives before any capacity is warmed up.

How far in advance should you run a load test before a big event?

Run the test at least two to three weeks before the event so you have time to act on the results. A test the day before a launch tells you there is a problem but leaves no room to fix it, re-test, and deploy the fix safely.

What is the difference between a load test and a spike test?

A load test holds traffic at or just above expected peak for an extended period, checking that the system sustains that level without degrading. A spike test compresses a far larger traffic surge into a very short window, checking whether the system survives a sudden burst and recovers cleanly.