Skip to content

Google Cloud outage (June 2025): the retry flood that extended a 40-minute fix into 3 hours

A null-pointer bug in new quota-policy code crash-looped Google Cloud's request-authorization service globally in June 2025. Most regions recovered in 40 minutes, but us-central1 spent 3 hours fighting a retry flood caused by missing backoff logic.

On 12 June 2025, Google Cloud’s request-authorization service crash-looped globally. This is the layer that checks quotas and policies for API calls across the platform. The bug was a null-pointer error introduced in a quota-policy code change, and for most of the world it was rolled back within about 40 minutes.

In us-central1, the outage ran for roughly three hours. The bug was no harder to fix there. The clients that had been failing during the outage all started retrying at once, without backoff, and the recovering authorization datastore could not absorb the burst.

This post shows how missing backoff logic turns a short incident into a long one, and how to test if your own services can recover under this pattern.

What happened

On the morning of 12 June 2025, Google Cloud engineers deployed a change to the quota-policy evaluation path inside their request-authorization service. The new code contained a null-pointer error that fired when it met policy entries with blank fields, and the new policy deployment produced exactly that condition right away.

The authorization service began crash-looping globally. Nearly every Google Cloud API call passes through request authorization, so the effect was broad: Compute, Storage, Kubernetes and many other services began returning errors to customers. The outage affected services across multiple regions simultaneously.

Engineers triggered a “red-button” rollback of the code change, and most regions stabilized within approximately 40 minutes.

The timeline

  • 12 Jun, ~09:45 UTC: The quota-policy change is applied, and new policy entries include blank fields. Authorization service begins crash-looping globally.
  • ~10:00–10:30 UTC: Customer-facing errors spread across Compute, Storage, GKE and other services. Error rates climb sharply.
  • ~10:25 UTC: Google engineers trigger a red-button rollback of the code change.
  • ~10:30–10:45 UTC: Most regions stabilize as the rollback takes effect. The authorization service recovers in most of Google Cloud.
  • ~10:45–13:00 UTC: us-central1 remains degraded. Retry traffic overwhelms the authorization datastore serving that region. The clients that failed during the crash-loop retry without randomized exponential backoff and flood the recovering service faster than it can process requests.
  • ~13:00 UTC: The us-central1 authorization service fully recovers after ~3 hours of degradation.

Why it happened

The initial crash-loop had a clear, fixable cause: a null-pointer error in new code, triggered by a policy with blank fields. Code review and integration testing should catch that kind of bug before deployment.

The longer outage in us-central1 had a different cause: retry behavior without randomized exponential backoff. When the authorization service crashed, every client started receiving errors and retrying, which is correct. But with no randomized wait before each retry, all clients sent their next request at nearly the same moment.

The authorization datastore had to absorb the combined demand of every failing client at once, at the exact moment it was trying to recover. It overloaded, which caused more failures, more retries and another round of the loop. The two causes were independent. Together they stretched a 40-minute fix into a 3-hour regional event.

The failure pattern

This is a retry storm, one of the most reliable ways to prolong an outage. The pattern appears whenever:

  1. A shared service becomes temporarily unavailable.
  2. A large number of clients start retrying at the same time.
  3. The retries are not spread over time (no jitter, no backoff, or too little backoff).
  4. The recovering service takes the full burst of accumulated retries, overloads again and cannot finish recovering.

The retry storm is the cloud-era thundering herd: all the failed requests come back at once. The more clients there are and the longer the outage lasts, the larger the burst when the service starts to recover.

Every SLO that promises fast recovery quietly assumes that client retries will not block recovery. Test that assumption.

How it could have been prevented

Jittered exponential backoff in every client. When a call fails, the client should wait before retrying, with a random component (jitter) so that many clients do not retry at the same moment. With jitter, the retry burst spreads into a stream the service can recover from.

Admission control on the recovering service. Even with client-side backoff, a recovering service should accept work at a controlled rate, queuing or rejecting excess requests, until it stabilizes.

Circuit breakers. Client-side circuit breakers stop retries entirely when error rates exceed a threshold, so recovery traffic arrives as a trickle instead of a burst.

Pre-deployment policy validation. The null-pointer bug was preventable. A test that applied the new code to a policy entry with blank fields would have shown the crash before deployment. Configuration changes to shared authorization services deserve the same care as code changes.

How to test for this with MaxoPerf

Simulate what happens to your service when it recovers from a brief outage and is hit right away by a retry burst from every client that was failing.

Engine: k6 or Taurus (JMeter).

Workload model: Open model. You want to set the arrival rate independently of service response time, because the number of queued-up clients drives a retry storm, not server capacity.

Profile (example):

Phase 1: Steady state. 200 virtual users for 5 minutes. Establish a clean baseline p95/p99 and a near-zero error rate.

Phase 2: Simulated outage. Return 503s for 2 minutes. Virtual users keep sending requests, and all of them fail. This is the accumulation window.

Phase 3: Recovery burst. Restore the endpoint. Ramp from 0 to 600 virtual users within 5 seconds, so every client retries at once. Hold 5 minutes and observe.

Phase 4: Control run with backoff. Repeat, but configure the client with jittered exponential backoff (e.g., 0.5–1.5s initial delay, doubling, capped at 30s). Compare error rate and recovery time with Phase 3.

Target: The authorization, session-validation or request-routing endpoint in your own staging environment, wherever clients pile up retries when a shared service is unavailable.

Execution locations: ≥2 managed MaxoPerf regions to model distributed clients resuming retries at the same time.

What to read in results:

  • Error rate in Phase 3: recovers quickly with backoff, stays high (congestive collapse) without it.
  • p99 latency: returns to baseline within seconds with good backoff, stays high without it.
  • Throughput plateau: successful req/s far below the incoming rate means the service is backed up.

Compare the no-backoff and jittered-backoff runs in the MaxoPerf results view. The difference in recovery time measures how well you withstand a retry storm. Add this as a release gate on anything touching retry configuration or shared service infrastructure.

Key takeaways

  • The June 2025 Google Cloud outage had two independent causes: a null-pointer code bug, and missing backoff logic in clients. The first was fixed in 40 minutes. The second kept us-central1 down for 3 hours.
  • Retry storms are predictable and testable. Clients that retry without backoff behave deterministically: they all retry at once, and the recovering service can’t handle it.
  • Jittered exponential backoff is not optional. Any client that calls a shared service and retries on failure must add a randomized delay. It is the basic defense against retry storms.
  • You only meet your SLO recovery time objective if your retry behavior cooperates. Test if your clients’ retry patterns help or block recovery. They often block it.
  • Stress test the recovery path as well as the happy path. A service that performs well at steady state may still fail to recover under a retry burst. Only a stress test built on realistic client retry behavior shows this.

Questions this article answers

What caused the Google Cloud outage in June 2025?

A code change introduced a null-pointer error in the quota-policy evaluation path of Google Cloud's request-authorization service. When a new policy was applied that included blank fields, the null-pointer caused the authorization service to crash in a loop globally. A "red-button" rollback stabilized most regions in about 40 minutes, but us-central1's authorization datastore was then overwhelmed by a retry flood from clients retrying without backoff, which extended recovery in that region to roughly 3 hours.

How can teams prevent retry floods from extending outages?

Implement jittered exponential backoff on every client that communicates with a shared service. When the service becomes unavailable, randomized wait intervals spread retry traffic over time instead of concentrating it into a burst. Pair this with admission control on the server side so the service can accept load at a safe rate as it recovers.

Why does missing backoff turn a short outage into a long one?

When a service goes down and clients retry without backoff, every client retries at nearly the same moment. The recovering service is immediately hit by more requests than it can handle, causing it to fail again. This loop continues until the load subsides on its own or engineers step in by hand. Neither is fast.