Skip to content

Cache invalidation and purge testing

In CDN operations, cache invalidation is one of the highest-risk routine tasks. When you purge a cached resource, every edge node that held it discards it. The next request to each edge is a cache miss and goes straight to your origin. If your site has high traffic and you purge a widely cached asset during peak hours, the origin surge (the “thundering herd”) can take your origin down.

CDN purge testing checks three things: (1) purges work and stale content is gone, (2) origin absorbs the post-purge traffic spike without failing, and (3) the cache warms back up within an acceptable time window.

  • Have a test CDN environment or a low-traffic asset URL ready. Purge testing means you issue real purge API calls, which temporarily increase origin traffic. Do not purge production hot-path assets during peak hours without a tested runbook.
  • Know your CDN provider’s purge API. CloudFront, Fastly, Cloudflare, and Akamai each have different purge mechanisms (API calls, URL-based purge, tag-based purge). The MaxoPerf test is a load test. You trigger the purge outside MaxoPerf.
  • Know your origin’s baseline capacity from previous load tests, so you can tell when a post-purge spike is a real risk.

After a CDN purge, the thundering herd plays out like this:

  1. You purge /static/app.js across all edge nodes simultaneously.
  2. Every edge node marks the cached object as invalid.
  3. The next request to each edge node (possibly hundreds worldwide) is a cache miss and a synchronous request to your origin.
  4. If you serve 10,000 requests/second globally across 200 edge nodes, those 200 simultaneous origin fetches arrive within milliseconds of each other.
  5. Your origin must serve those 200 fetches before any edge node can re-cache the response and serve later requests from cache again.

The severity depends on:

  • Asset popularity: a home-page bundle that every request touches is far riskier to purge than a rarely used static resource.
  • TTL: with a 24-hour TTL, many more edge nodes hold the asset than with a 5-minute TTL. Shorter TTLs reduce thundering-herd risk.
  • Origin cache: if your origin has an in-memory or Redis-backed application cache, it may still serve the first 200 fetches quickly.
  • CDN origin shield: an origin shield tier collapses all edge misses into a single request to the shield, which fetches from origin once. See Origin shield and offload testing.

Test pattern 1: measure post-purge origin impact

Section titled “Test pattern 1: measure post-purge origin impact”

This pattern runs realistic sustained traffic, triggers a purge, and measures how long the origin spike lasts.

Run a load test at your production traffic rate for 5 minutes to warm the cache. Confirm the cache hit ratio is above your target (typically 80–95 % for static assets). Record baseline p95 TTFB.

Issue the purge through your CDN provider’s API, outside MaxoPerf:

Terminal window
# CloudFront example: invalidate a path
aws cloudfront create-invalidation \
--distribution-id E1234ABCDEF \
--paths "/static/app.js"
# Fastly example: purge a URL
curl -X PURGE https://your-cdn.example.com/static/app.js \
-H "Fastly-Key: ${FASTLY_API_KEY}"
# Cloudflare example: purge by URL
curl -X POST "https://api.cloudflare.com/client/v4/zones/${ZONE_ID}/purge_cache" \
-H "Authorization: Bearer ${CF_API_TOKEN}" \
-H "Content-Type: application/json" \
--data '{"files":["https://your-cdn.example.com/static/app.js"]}'

Step 3: continue the MaxoPerf load test through the purge

Section titled “Step 3: continue the MaxoPerf load test through the purge”

Keep the MaxoPerf run going through the purge. The run overview shows TTFB jump at the moment of the purge, as edge nodes stop serving from cache and start fetching from origin.

Measure how long TTFB takes to return to baseline. This is your cache re-warm time. During this window your origin carries extra load and users see slower responses.

Test pattern 2: staleness window validation

Section titled “Test pattern 2: staleness window validation”

This pattern checks that cached content stays stale no longer than expected after a content update, so users do not see outdated responses for too long.

The Taurus test asserts that a Cache-Control: max-age header value does not exceed your maximum acceptable staleness window:

execution:
- concurrency: 5
hold-for: 2m
scenario: staleness-check
scenarios:
staleness-check:
default-address: https://cdn.example.com
requests:
- label: check-max-age
url: /v1/catalog/featured-items
method: GET
assert:
- equals:
subject: http-code
value: '200'
- contains:
subject: headers
value: 'Cache-Control'
- not-contains:
subject: headers
value: 'no-cache' # this endpoint must be cached

After the test, compare the Cache-Control: max-age value in the response headers with your freshness requirements. If max-age is too long, users get stale data for too long after a content update.

A purge storm happens when a large batch purge (e.g. deploying a new version of your entire static asset bundle) invalidates many objects at once. This is the highest-risk purge scenario.

To simulate and measure it:

  1. Identify the set of assets affected by a typical deployment purge. List them in a CSV file.
  2. Write a Taurus or k6 test that sends concurrent requests for all of those asset URLs immediately after the purge.
  3. Run the load test at production traffic rate. Time the test to start immediately after your batch purge completes.
  4. Watch origin load during the re-warm window with your origin’s own observability, not only MaxoPerf. Origin errors or slow responses show up in the MaxoPerf Log tab as 5xx responses.

How to read purge test results in MaxoPerf

Section titled “How to read purge test results in MaxoPerf”

When you run a load test through a purge event, look for:

SignalWhat it means
Sudden TTFB spike in the run overviewPurge took effect. Edge nodes now fetch from origin.
TTFB returning to baseline within minutesCache is re-warming as expected. The spike duration is your risk window.
5xx errors in the Log tab immediately after purgeOrigin could not handle the post-purge burst. Scale origin or enable origin shield.
TTFB does not recover after purgeThe TTL is very short or origin returns Cache-Control: no-cache. The CDN does not re-cache responses after the origin fetch.
Assertion failures on X-Cache: HIT during recoveryNormal during the re-warm window. The assertion failure rate drops as the cache fills back up.

Do:

  • Test purge behaviour in a staging environment or against a low-traffic asset before purging production hot paths.
  • Keep the MaxoPerf load test running through the purge so the TTFB timeline shows the spike and the recovery.
  • Test both targeted purge (single URL) and batch purge (whole directory or tag). They hit origin very differently.
  • Know your origin’s capacity limit before purge testing so you can define a safe stop condition.

Don’t:

  • Trigger a purge test during peak production traffic hours without a rollback plan and someone watching the origin.
  • Assume a successful purge means re-caching worked. Verify it with a follow-up cache hit assertion test.
  • Use wildcard purge (/*) for purge testing in production without measuring origin impact first.