Skip to content

Cloudflare outage (November 2025): the hard size limit that crashed a global proxy

A database query change caused an oversized configuration file to exceed a hard internal limit, crashing Cloudflare's core proxy globally for ~3 hours in November 2025. The lesson is boundary testing before shipping config changes.

On 18 November 2025, Cloudflare’s core proxy service crashed globally and cut off traffic for a large share of internet users for roughly three hours. The root cause was not a traffic surge. A database query change produced a configuration file too large for the proxy, and the proxy had no graceful way to handle it.

This is the “config-limit cliff” failure pattern. A hard limit inside a system, unknown or unguarded, causes a complete crash instead of partial degradation when input exceeds it. The lesson is boundary testing: check what happens when configuration data approaches and exceeds its limits before that configuration ships to production.

What happened

In the weeks before 18 November, Cloudflare’s Bot Management feature used a configuration file to define its active detection rules and feature flags. A query against a database of enabled features generated that file.

A database permissions change altered how that query ran. The previous permission set had enforced a filter implicitly. Without it, the query returned duplicate rows. The generated configuration file ended up with more than twice its normal number of entries, well above 200 items.

The core proxy that loaded this file had a hard limit of 200 items. When it met a file over the limit, it did not fall back to the previous valid configuration. It did not log a warning and skip the oversized file. It panicked (an unhandled error in the proxy’s configuration-loading code) and crashed.

The crash happened in every Cloudflare data center, so the proxy was unavailable worldwide until the fix.

The timeline

  • 18 Nov, before 11:28 UTC: A database permissions change is deployed. The configuration query starts returning duplicate rows in the background.
  • 11:28 UTC: A scheduled configuration refresh sends the oversized Bot Management configuration file to proxy instances globally. Proxies begin crashing as they attempt to load the file.
  • ~11:30–14:30 UTC: The core proxy is unavailable globally (~3 hours of widespread disruption). Cloudflare engineers identify the oversized configuration file as the cause.
  • Mitigation: Engineers push a corrected configuration file that stays within the 200-item limit. Proxies reload and recover.
  • ~14:30 UTC: Service restored for most users. Full recovery confirmed by ~17:30 UTC.

Why it happened

The trigger was a database permissions change that removed an implicit query constraint, so duplicate entries came back. The cascade needed a second condition: the proxy’s configuration loader could not gracefully handle input over its hard size limit.

Cloudflare’s postmortem identifies both problems:

  1. Nobody checked the query change’s effect on output size or item count before deployment.
  2. The proxy’s configuration-loading code treated an oversized input as a fatal, unrecoverable error (a panic). It should have rejected the bad input and kept running on the last known-good configuration.

Neither condition alone would have caused a global outage. Together, they did. The database change was the “trigger.” The hard limit with no graceful degradation was the “cliff.”

The failure pattern

This is a config-limit cliff: a hard boundary inside a system that nobody tests at deployment time, and that causes a crash instead of a degradation when crossed.

Config-limit cliffs are especially dangerous in globally distributed infrastructure because:

  • A single configuration deployment reaches all instances at once. There is no canary window.
  • If the failure mode is a crash instead of a rejection, the entire fleet fails at once.
  • No instance can recover until someone generates and deploys a corrected configuration. Nothing heals itself.

The pattern applies to any system that accepts structured configuration input: request routers, feature-flag systems, WAF rule sets, API gateway configurations, serialization schemas. If a “configuration” entity has a hard limit on size, count or depth, and the system does not check inputs against that limit before accepting them, a config-limit cliff exists.

How it could have been prevented

Validate configuration output size before deployment. A pre-deployment check should have caught the query change that doubled the file: “does this query return more rows than the configuration consumer can handle?” Configuration generation pipelines need size assertions as well as syntax checks.

Fail open, not closed, on oversized configuration. When the proxy met a file over its limit, it should have rejected the new file, logged an error and kept running on the previous valid configuration. A panic that crashes the process is the worst possible response to bad input.

Apply the same limit in the generator. The system that generates and distributes configuration files should enforce the same 200-item limit as the consumer, and refuse to distribute a file that breaks it.

Canary configuration deployments. Rolling out configuration changes to a small subset of instances first would have shown the crash before it reached every data center.

Boundary tests in the pipeline. Automated tests that produce files at exactly the limit, one below and one above, and check the consumer’s behavior in each case, would have caught both the missing size check and the panic.

How to test for this with MaxoPerf

The config-limit cliff calls for two test types. A configuration test checks behavior at and around the hard limit. A volume test confirms the system handles large variation in input without silent failures.

Engine: k6 or Taurus.

Workload model: Closed model. You are testing correctness at specific input boundaries, not raw throughput.

Profile (example):

Phase 1: Boundary test. Send requests that produce payloads at 80%, 100%, and 110% of the documented limit. At 100%, the system should accept and respond correctly. At 110%, expect a clean 4xx rejection, not a crash or 5xx. Use 10–20 VUs, 2 minutes per phase.

Phase 2: Volume test under configuration churn. Simulate 1 configuration update per second for 10 minutes while baseline load continues at 200 VUs. Check that reloads, including large ones, do not spike the error rate.

Target: The configuration-loading, feature-flag evaluation or rule-set ingestion endpoint in your own staging environment.

Execution locations: A single managed MaxoPerf region for boundary testing; ≥2 regions for distributed configuration delivery volume tests.

What to read in results:

  • Error rate at 110% boundary: clean 4xx, not 5xx or timeout.
  • Error rate during churn: near zero. Any spike points to a gap in reload isolation.
  • p99 latency during reloads: a large spike suggests the reload blocks request handling.

Use the run artifacts tab to confirm that rejection messages are structured responses, not crash traces. Run the boundary test as a CI gate on any deployment that touches configuration generation or consumption.

Key takeaways

  • A configuration file that exceeded a hard internal limit, with no graceful fallback, caused the November 2025 Cloudflare outage. Traffic volume did not.
  • A hard limit with no graceful degradation is a cliff. Any system that panics or crashes on oversized input is one bad deployment away from a global outage.
  • Configuration generation and configuration consumption must share the same size constraints. If the consumer enforces a limit, the generator must enforce it too before distributing.
  • Boundary testing belongs in the deployment pipeline. Test configuration inputs at, below and above their limits before any deployment, and you catch these failures in staging instead of production.
  • Fail open on bad configuration. Reject the new file and continue with the last known-good state. A potential outage becomes a logged error.

Questions this article answers

What caused the Cloudflare outage in November 2025?

A database permissions change caused a configuration query to return duplicate rows, producing a Bot Management configuration file that was more than twice its normal size. The file exceeded a hard 200-item limit in Cloudflare's core proxy. The proxy panicked on the oversized input and crashed globally. There was no graceful degradation when the limit was hit.

How can teams prevent configuration-limit outages?

Run configuration tests and volume tests that push input sizes, item counts, and payload dimensions to and beyond their documented limits before deploying changes. Verify that the system degrades gracefully (returning an error or falling back to a safe default) instead of crashing when a limit is exceeded.

What is a config-limit cliff?

A config-limit cliff is when a system has a hard internal size or count limit that, once exceeded, causes an immediate failure rather than a graceful degradation. Systems that simply panic or crash on oversized input are the most dangerous kind because a single oversized deployment can take down the entire service.