Skip to content

GitHub capacity outages 2025–2026 — what a year of failures teaches about load testing

GitHub suffered dozens of capacity-related outages between 2025 and early 2026. Here's what caused them and how to test for the same failure pattern.

Tens of millions of developers keep their code on GitHub. When it goes down, pull requests stall, deployments stop and CI pipelines fail around the world. Between mid-2025 and early 2026, a cluster of outages kept GitHub’s reliability in the headlines. The story behind most of them was simple: traffic grew faster than capacity.

This post walks through what happened, why the pattern repeats, and how to use load testing to catch the same failure class before it reaches production.

What happened

From October 2025 through April 2026, GitHub experienced a sustained cluster of availability incidents. The company’s own CTO acknowledged the problem publicly. He pointed to tight service coupling that let failures cascade, no way to shed load from high-volume clients, and services that needed manual scaling instead of automated headroom. The company launched a 10×–30× capacity expansion program in response.

On February 17, 2026, GitHub’s authentication path degraded under peak read load, and that incident shows the pattern well. Database read replicas fell behind the primary under sustained traffic, and token lookups became inconsistent. Developers got intermittent failures even with valid credentials. The team rerouted reads and added replicas, but those steps only happen after users already see the problem.

The timeline

  • October 2025 onward: The recurring outage cluster begins. Agentic and AI coding workflows drive a fast-growing traffic increase across repository, CI and API paths.
  • February 2, 2026: GitHub Actions degrades for approximately 3 hours and 40 minutes as job processing capacity saturates.
  • February 17, 2026: The authentication path is disrupted. Database read replicas lag under peak read load, and token lookups return inconsistent results. Engineering teams mitigate by rerouting reads and adding capacity.
  • April 2026: Third-party analysis confirms capacity as the leading cause, attributing 83 of 257 incidents to it. GitHub’s CTO publicly addresses the coupling and load-shedding gaps.

Why it happened

Three problems fed each other:

Traffic outpaced capacity. AI-assisted and agentic coding workflows sharply increased API and repository traffic. Services that lacked automatic scaling needed engineers to scale them by hand after symptoms appeared, so capacity was always chasing the problem.

Read replicas have a lag ceiling. Under sustained read-heavy load, database read replicas can fall behind the primary. Past a certain throughput, replication lag grows faster than the replicas can catch up. The system moves from “replicas occasionally a bit stale” to “reads returning wrong data”, and for authentication, wrong data means failed logins.

Tight coupling turned partial overload into cascades. When one service saturated, requests backed up into the services that depended on it. With no load-shedding rules to cut off a single high-volume client or queue from the rest of the system, a local overload spread into platform-wide degradation.

The failure pattern

This is a textbook capacity-ceiling failure. A system runs fine up to a throughput threshold, then degrades non-linearly once it crosses it. The threshold often stays invisible until production crosses it, because teams test at today’s traffic, not tomorrow’s.

The pattern is dangerous because it compounds. Replica lag causes authentication failures. Authentication failures cause retries. Retries add read load. The coupling GitHub’s CTO described is what turns a manageable overload into a widespread incident.

Any team whose read volume is growing is exposed to the same pattern, especially teams that serve machine-generated or automated clients.

How it could have been prevented

Autoscale critical paths. Manual scale-up is always late. Automated horizontal scaling, triggered on latency or queue depth, should act before users feel anything.

Separate replica read limits from primary write capacity. Replica lag becomes dangerous when the replication topology has no headroom. Over-provision read replicas relative to current peak and set health alerts that fire before lag reaches the danger zone.

Load-shed high-volume clients. Queue or rate-limit high-volume automated clients (bots, agentic workflows) before they saturate shared infrastructure. This rate limiting belongs in the traffic-management layer as well as the application layer.

Find the breakpoint before users do. A breakpoint test run on a regular basis tells you exactly where throughput plateaus and errors begin. Running it on a schedule means you learn about drift before your users report it.

How to test for this with MaxoPerf

Catching a capacity ceiling takes two parts: find the threshold, then monitor it over time.

Step 1: Breakpoint test

Pick the paths most likely to saturate first: your authentication endpoint, your primary API read path, or whatever your traffic analysis says is growing fastest. With k6 or Taurus in MaxoPerf, configure a breakpoint test that ramps concurrency in steps. For example, hold 50 virtual users for 2 minutes, then jump to 100, 200, 400, and so on. Run from at least two managed regions to model users in more than one geography.

Watch three signals in the run results:

  • Throughput plateau: the point where requests per second stop climbing even as VUs increase. That is your capacity ceiling.
  • p95/p99 latency inflection: latency stays flat while the system is healthy, then bends sharply at saturation. Note the VU level where the bend starts.
  • Error rate departure from zero: the first moment your error rate exceeds the baseline is the ceiling your users will actually hit.

Step 2: Scalability test

Once you know your breakpoint, run a scalability test at 50%, 100%, 150%, and 200% of your current peak. It shows if your system scales linearly or degrades non-linearly. The replica-lag pattern shows up here as latency climbing faster than VUs.

Step 3: Scheduled recurring runs

Traffic drift makes no noise. Schedule recurring breakpoint or scalability runs with MaxoPerf’s recurring test scheduler, weekly or after major releases. When your capacity floor starts falling relative to your target, you know before your users do. Run results are stored with each execution, so you can compare the current run with last week’s baseline without keeping records by hand.

If you run private infrastructure (self-hosted runners, internal staging environments), MaxoPerf’s private/BYOC execution locations let you run the same test against your internal stack with the same results workflow.

Key takeaways

  • Capacity-ceiling failures usually stay invisible until crossed. Traffic grows gradually, but the threshold crossing is sudden. Measure ahead of time instead of scaling after the fact.
  • Read replicas have a lag ceiling that depends on load. Any read-heavy path under sustained growth needs an explicit replica headroom target and an alert before lag becomes dangerous.
  • Autoscaling must fire before the symptom. Manual scale-up is always reactive. Trigger automation on leading indicators (queue depth, latency trend), not lagging ones (an error rate that is already rising).
  • Tight coupling amplifies every overload. Load-shedding rules and circuit breakers turn a partial overload into a managed degradation instead of a cascade.
  • A scheduled breakpoint run is the cheapest insurance. Running at multiples of current peak on a regular cadence catches drift before users report it.

Questions this article answers

Why did GitHub go down so often in 2025 and 2026?

Third-party analysis found that capacity constraints were the single leading cause of GitHub's outages during this period, accounting for 83 of 257 incidents. Traffic grew faster than the platform scaled.

How do I test for the same capacity-ceiling failure that affected GitHub?

Run a breakpoint test that steps virtual users up until throughput plateaus and error rate climbs, then schedule recurring runs so you catch any drift before your users do.

What is database replica lag and why does it cause authentication errors?

Under heavy read load, database read replicas can fall behind the primary and serve stale data. For authentication, that means token lookups return inconsistent results and requests fail even when credentials are valid.