Cache invalidation and purge testing
In CDN operations, cache invalidation is one of the highest-risk routine tasks. When you purge a cached resource, every edge node that held it discards it. The next request to each edge is a cache miss and goes straight to your origin. If your site has high traffic and you purge a widely cached asset during peak hours, the origin surge (the “thundering herd”) can take your origin down.
CDN purge testing checks three things: (1) purges work and stale content is gone, (2) origin absorbs the post-purge traffic spike without failing, and (3) the cache warms back up within an acceptable time window.
Before you start
Section titled “Before you start”- Have a test CDN environment or a low-traffic asset URL ready. Purge testing means you issue real purge API calls, which temporarily increase origin traffic. Do not purge production hot-path assets during peak hours without a tested runbook.
- Know your CDN provider’s purge API. CloudFront, Fastly, Cloudflare, and Akamai each have different purge mechanisms (API calls, URL-based purge, tag-based purge). The MaxoPerf test is a load test. You trigger the purge outside MaxoPerf.
- Know your origin’s baseline capacity from previous load tests, so you can tell when a post-purge spike is a real risk.
Understanding the thundering herd
Section titled “Understanding the thundering herd”After a CDN purge, the thundering herd plays out like this:
- You purge
/static/app.jsacross all edge nodes simultaneously. - Every edge node marks the cached object as invalid.
- The next request to each edge node (possibly hundreds worldwide) is a cache miss and a synchronous request to your origin.
- If you serve 10,000 requests/second globally across 200 edge nodes, those 200 simultaneous origin fetches arrive within milliseconds of each other.
- Your origin must serve those 200 fetches before any edge node can re-cache the response and serve later requests from cache again.
The severity depends on:
- Asset popularity: a home-page bundle that every request touches is far riskier to purge than a rarely used static resource.
- TTL: with a 24-hour TTL, many more edge nodes hold the asset than with a 5-minute TTL. Shorter TTLs reduce thundering-herd risk.
- Origin cache: if your origin has an in-memory or Redis-backed application cache, it may still serve the first 200 fetches quickly.
- CDN origin shield: an origin shield tier collapses all edge misses into a single request to the shield, which fetches from origin once. See Origin shield and offload testing.
Test pattern 1: measure post-purge origin impact
Section titled “Test pattern 1: measure post-purge origin impact”This pattern runs realistic sustained traffic, triggers a purge, and measures how long the origin spike lasts.
Step 1: establish a warm-cache baseline
Section titled “Step 1: establish a warm-cache baseline”Run a load test at your production traffic rate for 5 minutes to warm the cache. Confirm the cache hit ratio is above your target (typically 80–95 % for static assets). Record baseline p95 TTFB.
Step 2: trigger the purge externally
Section titled “Step 2: trigger the purge externally”Issue the purge through your CDN provider’s API, outside MaxoPerf:
# CloudFront example: invalidate a pathaws cloudfront create-invalidation \ --distribution-id E1234ABCDEF \ --paths "/static/app.js"
# Fastly example: purge a URLcurl -X PURGE https://your-cdn.example.com/static/app.js \ -H "Fastly-Key: ${FASTLY_API_KEY}"
# Cloudflare example: purge by URLcurl -X POST "https://api.cloudflare.com/client/v4/zones/${ZONE_ID}/purge_cache" \ -H "Authorization: Bearer ${CF_API_TOKEN}" \ -H "Content-Type: application/json" \ --data '{"files":["https://your-cdn.example.com/static/app.js"]}'Step 3: continue the MaxoPerf load test through the purge
Section titled “Step 3: continue the MaxoPerf load test through the purge”Keep the MaxoPerf run going through the purge. The run overview shows TTFB jump at the moment of the purge, as edge nodes stop serving from cache and start fetching from origin.
Step 4: measure recovery time
Section titled “Step 4: measure recovery time”Measure how long TTFB takes to return to baseline. This is your cache re-warm time. During this window your origin carries extra load and users see slower responses.
Test pattern 2: staleness window validation
Section titled “Test pattern 2: staleness window validation”This pattern checks that cached content stays stale no longer than expected after a content update, so users do not see outdated responses for too long.
The Taurus test asserts that a Cache-Control: max-age header value does not exceed your maximum acceptable staleness window:
execution: - concurrency: 5 hold-for: 2m scenario: staleness-check
scenarios: staleness-check: default-address: https://cdn.example.com requests: - label: check-max-age url: /v1/catalog/featured-items method: GET assert: - equals: subject: http-code value: '200' - contains: subject: headers value: 'Cache-Control' - not-contains: subject: headers value: 'no-cache' # this endpoint must be cachedAfter the test, compare the Cache-Control: max-age value in the response headers with your freshness requirements. If max-age is too long, users get stale data for too long after a content update.
Test pattern 3: purge storm simulation
Section titled “Test pattern 3: purge storm simulation”A purge storm happens when a large batch purge (e.g. deploying a new version of your entire static asset bundle) invalidates many objects at once. This is the highest-risk purge scenario.
To simulate and measure it:
- Identify the set of assets affected by a typical deployment purge. List them in a CSV file.
- Write a Taurus or k6 test that sends concurrent requests for all of those asset URLs immediately after the purge.
- Run the load test at production traffic rate. Time the test to start immediately after your batch purge completes.
- Watch origin load during the re-warm window with your origin’s own observability, not only MaxoPerf. Origin errors or slow responses show up in the MaxoPerf Log tab as 5xx responses.
How to read purge test results in MaxoPerf
Section titled “How to read purge test results in MaxoPerf”When you run a load test through a purge event, look for:
| Signal | What it means |
|---|---|
| Sudden TTFB spike in the run overview | Purge took effect. Edge nodes now fetch from origin. |
| TTFB returning to baseline within minutes | Cache is re-warming as expected. The spike duration is your risk window. |
| 5xx errors in the Log tab immediately after purge | Origin could not handle the post-purge burst. Scale origin or enable origin shield. |
| TTFB does not recover after purge | The TTL is very short or origin returns Cache-Control: no-cache. The CDN does not re-cache responses after the origin fetch. |
Assertion failures on X-Cache: HIT during recovery | Normal during the re-warm window. The assertion failure rate drops as the cache fills back up. |
Do / don’t
Section titled “Do / don’t”Do:
- Test purge behaviour in a staging environment or against a low-traffic asset before purging production hot paths.
- Keep the MaxoPerf load test running through the purge so the TTFB timeline shows the spike and the recovery.
- Test both targeted purge (single URL) and batch purge (whole directory or tag). They hit origin very differently.
- Know your origin’s capacity limit before purge testing so you can define a safe stop condition.
Don’t:
- Trigger a purge test during peak production traffic hours without a rollback plan and someone watching the origin.
- Assume a successful purge means re-caching worked. Verify it with a follow-up cache hit assertion test.
- Use wildcard purge (
/*) for purge testing in production without measuring origin impact first.
Where to go next
Section titled “Where to go next”- Origin shield and offload testing: reduce purge storm risk with shield configuration.
- Cache hit/miss testing: verify the cache re-warms correctly after a purge.
- Cache key and Vary testing: how cache keys affect purge scope.
- Daily scenarios: the “purge storm” scenario in context.