Skip to content

OpenAI ChatGPT outage June 2025 — an illustrative AI-capacity story

A 10-plus-hour ChatGPT outage in June 2025 affected free, Plus, API, and Sora users. No official root-cause analysis was published. Here is what the pattern suggests and how to test for it.

On June 10, 2025, ChatGPT and OpenAI’s API went down for more than ten hours. Free users, Plus subscribers, API customers, and Sora users were all hit, with approximately 21 service components reporting failures at the same time. OpenAI resolved the incident by fixing what it called “the underlying infrastructure issue.” It published no detailed root-cause analysis.

Important: the load cause discussed in this post is inferred, not confirmed. We include this incident as an illustration of how pressure on AI inference capacity can show up. It is not a case with a verified load root cause on record. Where cause is uncertain, this post says so plainly.

What happened

According to reporting at the time, the outage began on June 10, 2025 and lasted more than 10 hours. Users reported that ChatGPT was completely unavailable or returning errors, and the OpenAI API had the same problems. Free tier, paid Plus tier, API access, and Sora all failed at once. That scope points to a platform-level failure, not a front-end problem or one isolated service.

OpenAI’s public statements acknowledged the incident and said the infrastructure issue was fixed. The CEO mentioned heavy demand on the system in wider public comments around this period, though not as a direct explanation of this incident. OpenAI released no formal postmortem.

Given the scope, the duration, and how fast usage was growing at the time, load and capacity pressure on AI inference infrastructure is a plausible contributing factor. It is still an inference.

The timeline

  • June 10, 2025: a long outage begins. Free, Plus, API, and Sora services are unavailable or degraded.
  • ~21 components reporting failures: the status page shows broad impact at the same moment, which suggests a shared infrastructure layer rather than one isolated service.
  • 10+ hours of degradation: a long stretch before resolution, which suggests the fix needed more than a restart or a configuration rollback.
  • Resolution: OpenAI engineers fix the underlying infrastructure issue, and services come back.
  • No public postmortem: no detailed root-cause analysis has been published as of this writing.

Why it happened

With no official postmortem, this section is speculation. Read it as an illustrative analysis, not a verified account.

AI inference workloads make unusual demands on infrastructure. Traditional API requests finish in milliseconds. Inference requests are compute-heavy, slow, and stateful in ways that make horizontal scaling harder than adding servers. GPU resources are expensive and slow to provision, and demand for AI services grew faster in 2025 than most capacity planning cycles could keep up with.

If capacity pressure did contribute here, the likely dynamics would be these. Inference requests take longer under heavy load. Queues back up. Delayed responses make clients retry, which adds more load, which delays responses further. Eventually the system cannot serve new requests at all, which would explain the broad component failures.

To repeat: this is the inferred pattern, not the confirmed cause. If you build on AI inference APIs, take it as a reason to test for capacity boundaries. It says nothing about what OpenAI’s internal postmortem would have found.

The failure pattern

We classify this incident as a capacity-ceiling failure of the AI/GPU pool sub-type. The pattern applies to any team building on AI inference infrastructure, whether a third-party provider or a self-hosted model serving system.

Capacity ceiling failures hit AI inference especially hard because inference requests are often synchronous and visible to the user. When a web search or database query fails, the user sees an error. When an inference request fails, the user’s workflow stops. Inference downtime is usually more visible and does more damage than the same downtime in supporting services.

If you build applications on AI inference APIs, you need to know what your application does when the inference provider returns a 429 or 529 response. Does it queue the work, show a clear message, and retry with backoff? Or does it show the user a confusing error?

How it could have been prevented

What follows is general engineering practice. It makes no claim about what OpenAI should have done differently.

Provision inference capacity ahead of demand signals. GPU and accelerator lead times are long. If you wait for utilization metrics to show saturation before ordering more capacity, you will always trail the demand curve. Plan AI infrastructure capacity from demand forecasts, not current utilization.

Shed load before you hit the ceiling. A well-designed AI service starts returning 429 or 529 responses before it becomes completely unavailable. A system that sheds load in a controlled way, accepting some requests and declining others with a clear retry-after signal, holds up far better than one that tries to accept everything until it falls over.

Build client-side retry logic with backoff. Applications that call AI inference APIs should expect 429 responses and handle them with exponential backoff and jitter. Do not retry immediately. A client that retries at once on a 429 makes the capacity problem worse.

Test how your application behaves when the inference provider degrades. “Will the provider stay up?” is not a question you can act on. “When the provider is slow or shedding load, what does my application do?” is.

How to test for this with MaxoPerf

There are two separate tests to run. They apply if you operate your own AI inference infrastructure, or if you want to know how your application holds up when the inference API it depends on is under pressure.

Test 1: breakpoint test on your inference endpoint

If you host or control your own model-serving infrastructure, run a breakpoint test against your inference endpoint with MaxoPerf and k6 or Taurus. Set up an open workload model that steps requests per second upward until you reach the ceiling:

  • Start at a low sustained RPS you know is safe, such as 10 RPS.
  • Step up by 10–20 RPS every two minutes.
  • Keep going until 429/529 responses start, p99 latency passes your tolerance, or throughput levels off even though RPS keeps rising.

The point where errors start is your capacity ceiling. The distance between that ceiling and your expected peak is your headroom, or the gap you have to close.

Test 2: application resilience test with simulated inference degradation

This test targets your own application, not the inference provider. Run your application in a staging environment where you can configure the inference API (or a mock of it) to return slow responses or 429s. Then load test your application with MaxoPerf at normal expected traffic:

  • Hold normal traffic for 5 minutes.
  • Switch the inference mock to return 429s with a Retry-After: 5 header.
  • Hold for 5 minutes.
  • Restore normal inference responses.
  • Watch the recovery.

Watch in results:

  • Does the error rate your end users see go up when the inference provider returns 429s? If so, by how much?
  • Does throughput recover cleanly after the inference provider recovers, or does a backlog keep your application degraded?
  • Are the error messages users see meaningful, or do they expose the provider’s internal error codes?

A test gives you concrete answers to these questions. Without one, you have none.

Run the tests from MaxoPerf’s managed cloud locations to spread load across regions the way real users are spread, or use private/BYOC execution locations if your inference infrastructure is only reachable from inside your network.

Key takeaways

  • OpenAI published no official root-cause analysis for this incident. The load cause discussed here is inferred and illustrative, not confirmed.
  • Pressure on AI inference capacity is a real and growing risk. Demand for AI services in 2025 grew faster than infrastructure provisioning cycles. If you build on inference APIs, plan for provider degradation as a normal operating scenario.
  • Shedding load in a controlled way matters more than total capacity. Clients cope far better with a system that returns 429s clearly and early than with one that accepts every request until it becomes unavailable.
  • You can test how your application behaves during inference degradation. You do not have to wait for a provider outage. A mock inference backend set to return 429s tells you everything.
  • Breakpoint testing shows the headroom gap. If you control your own inference infrastructure, act on the distance between your capacity ceiling (found with a breakpoint test) and your traffic forecast.

Questions this article answers

What caused the ChatGPT outage on June 10 2025?

OpenAI did not publish a detailed root-cause analysis. The load cause is inferred, not confirmed. OpenAI resolved the incident by fixing an underlying infrastructure issue and the CEO referenced heavy demand, but no official postmortem was released.

How should teams test AI inference APIs for capacity limits?

Run a breakpoint test that steps requests per second upward until the API begins returning 429 or 529 responses. That identifies the sustainable throughput ceiling before users encounter it.

Why do AI services use 429 and 529 status codes during capacity pressure?

HTTP 429 means too many requests and is the standard rate-limit signal. Some AI providers use 529 to indicate service overload specifically. Both are forms of load shedding: the server is telling clients to slow down.