Clerk Cloud SQL migration outage — when a database operation becomes a load event
A routine storage migration turned into a 26-minute outage for Clerk in March 2026. The dependency-contention failure pattern it exposed applies to any team running a live database.
Authentication and identity services sit on the critical path of almost every user action in a modern application. When they degrade, nothing else works. On March 10, 2026, Clerk, a developer authentication and user management platform, had a 26-minute outage. The cause was a routine database storage operation that collided with live production load.
The incident is a clear example of a failure class that catches many engineering teams off guard. In the dependency-contention pattern, a degraded or briefly slow dependency spreads into service-wide saturation.
What happened
According to Clerk’s own postmortem, a Cloud SQL live migration ran on the production database without prior notice to customers. The migration raised disk latency, and the latency spread upward. Database operations that normally finished quickly held locks longer, which increased lock contention across the connection pool. As contention grew, compute saturated. The system could not process incoming authentication requests fast enough, and callers saw the overflow as HTTP 429 responses.
The outage lasted approximately 26 minutes and ended on its own when the migration finished. It was the second similar incident in seven months. Afterwards, Clerk announced a move to maintenance windows with replica promotion instead of live migrations.
The timeline
- March 10, 2026, onset: A Cloud SQL live migration starts on the production database without advance notice to users.
- Elevated disk latency: Storage I/O slows, and database operations queue while they wait for the disk.
- Lock contention rises: Slower operations hold shared resources longer and block other queries. The contention compounds.
- Compute saturation: The connection pool fills. Worker threads sit waiting, and new requests cannot get in.
- 429 responses: Incoming authentication requests begin returning 429 errors as the service sheds load it cannot process.
- ~26 minutes in: The migration completes. Disk latency returns to normal, contention clears and the service recovers without manual intervention.
Why it happened
The direct cause was a scheduling choice: a live migration ran on the live database during production hours. The underlying weakness existed before that decision. The service had no buffer between a degraded dependency and user-visible failure.
When storage latency rose, every database operation that touched disk slowed down. Those operations held row locks and connection-pool slots longer than expected. Once enough of them stacked up, new requests could not get what they needed. The system had no way to queue gracefully or prioritize critical paths, so it saturated and returned errors.
This is dependency contention. A normally healthy dependency becomes slower for a while, and the service has no resilience layer to absorb the slow responses. The failure travels from the database layer up through compute and out to the API.
The failure pattern
Dependency-contention failures share a shape. The service works normally up to a certain dependency latency. Past that point, lock and connection exhaustion begin, and the service degrades fast. The nonlinearity is what surprises teams. The first 20ms of extra database latency is invisible to users. The next 20ms sends error rates to 100%.
The trigger does not have to be a migration, which is why this matters to any team running a live database. The same cascade can happen during a slow backup, an index rebuild, a noisy neighbor on shared infrastructure or any maintenance operation that slows storage for a while.
How it could have been prevented
Schedule maintenance operations away from production load. Plan migrations, index rebuilds and other storage-heavy operations for low-traffic windows. Or run them against a promoted replica with a traffic cut-over, not against the live primary while full traffic flows through it.
Test the degraded-dependency case before any migration. Before you run a migration in production, check in staging how the service behaves when database latency rises. If 50ms of extra disk latency exhausts the connection pool in the test, fix that before the real migration runs.
Add circuit breaking and queue depth limits. A service that detects a degraded dependency and sheds non-critical work (or queues it with a short timeout) turns a 26-minute outage into a brief degradation for some request types instead of total unavailability.
Track this class of risk against your error budget. A 26-minute outage of an authentication service burns a large part of a monthly error budget. Gate risky migrations on whether the error budget has enough headroom, and make that headroom visible before the change is scheduled.
How to test for this with MaxoPerf
Watch how your own service behaves when its database dependency degrades, before that happens in production under real load.
Set up a staging environment with controllable storage latency. Most cloud database services offer a test environment. You can also add artificial latency at the network layer or use database-proxy tools in staging. You need an environment where you control how long disk operations take.
Run a stress test against your staging authentication or primary API path. Configure a k6 or Taurus script in MaxoPerf that targets the endpoint most dependent on the database. Use an open workload model so the number of in-flight requests does not drop by itself as latency climbs. A representative profile:
- Ramp from 0 to your expected peak over 2 minutes.
- Hold at peak for 5 minutes with normal database latency.
- Add storage latency to simulate the migration: 50ms, then 100ms, then 200ms.
- Hold for 5 minutes at each latency level.
- Return to normal latency and observe recovery.
Watch these signals in MaxoPerf results:
- Error rate: the key indicator. Note the latency level where it leaves the baseline.
- p95/p99 response time: climbs before errors do. The point where p99 pulls away from p95 is often where lock contention starts to cascade.
- Throughput: successful requests per second plateau or drop before the error rate spikes. This is the connection-pool exhaustion signal.
Record the threshold and build a pre-migration checklist. Once you know your service starts returning errors at X ms of database latency, you have a concrete number for your migration plan: “this migration must not raise disk latency above X ms for more than Y seconds.” Run the stress test again after adding circuit breaking or connection-pool limits to confirm the fix works.
If your test or production environment is not reachable from the public internet, use MaxoPerf’s private/BYOC execution locations to run the same stress profile from inside your network, with the same test and results workflow.
Key takeaways
- A healthy database that runs slower than usual can take a healthy service down. The link between dependency latency and service error rate is often nonlinear, and you reach the dangerous zone quickly.
- Live migrations are load events. If you treat them as invisible background work, you miss how disk-latency spikes travel upward into lock contention and compute saturation.
- Test the degraded-dependency case as well as the happy path. A stress test with added database latency shows you the cascade point before a migration shows it to users.
- Circuit breaking and queue limits turn cascades into partial degradation. A service that detects contention and limits in-flight requests survives a slow dependency. One without that protection does not.
- Knowing your error budget prevents surprises. If you know your monthly budget and what a 26-minute authentication outage costs, you can decide when and how to run high-risk maintenance.
Questions this article answers
What caused the Clerk outage on March 10 2026?
A live database storage migration spiked disk latency, which increased lock contention, saturated compute, and caused incoming requests to return 429 errors for approximately 26 minutes.
How do I test for dependency-contention failures like the one that hit Clerk?
Run a stress test against your staging environment while deliberately elevating database or storage latency. Note when lock contention and error rates begin climbing. That is your system's dependency-contention threshold.
What is an error budget and how does it apply to database migrations?
An error budget is the acceptable failure rate defined by an SLO. Migrations that risk elevated error rates should be planned and tested within budget headroom, not assumed to be zero-impact.