Skip to content

Spike and stress for sales events

Before every BFCM, you need to know at what load our system breaks. If you do not, your customers will find out for you on the highest-revenue day of the year. Spike and stress tests answer the question on purpose, in a controlled environment, while you still have time to fix what they find.

  • The full-funnel load test from end-to-end journey load passes at your modeled peak VU count.
  • Failure criteria are configured on the test so the breaking point is captured automatically.
  • Stakeholders know this test will cause errors and may cause temporary unavailability on the target environment.
  • You are running against staging, not production.

A spike test and a stress test find different breaking points:

Test typeLoad shapeWhat it finds
SpikeNear-instant jump to 5×–10× normalConnection pool exhaustion, session store saturation, cold-start failure under sudden load
StressGradual ramp to 2×–5× normalThe concurrency ceiling, database lock contention, CPU/memory saturation

For BFCM you need both. A doorbuster is a spike problem. Sustained Cyber Monday traffic is a stress problem. Run only one and you get half the answer.

A flash-sale spike jumps from warm idle (50–100 VUs) to 5×–10× your modeled peak in under 30 seconds. The system has no time to autoscale. Any resource that is lazily initialized or limited by a connection pool fails at once.

execution:
- scenario: bfcm-spike
stages:
- duration: 2m
target: 100 # warm idle baseline
- duration: 20s
target: 6800 # 5× peak — doorbuster shock
- duration: 5m
target: 6800 # hold at shock level
- duration: 2m
target: 100 # drop and observe recovery
- duration: 5m
target: 100 # recovery window
scenarios:
bfcm-spike:
requests:
- url: https://api.staging.example.com/v1/products/flash-deal
label: GET /products/flash-deal
- url: https://api.staging.example.com/v1/cart/items
label: POST /cart/items
method: POST
headers:
Content-Type: application/json
body: '{"product_id": "flash-deal-001", "qty": 1}'
  1. Create a new test named bfcm-2025-doorbuster-spike in the MaxoPerf console.
  2. Upload the YAML above as the entrypoint.
  3. Add failure criteria: POST /cart/items p95 > 2000ms and error rate > 5 %. Past these thresholds, the spike counts as “broken.”
  4. Click Run and watch the Overview tab. The spike shows clearly in the VU chart.
  5. Note the exact VU count and timestamp when error rate first goes above 1 %. That is the practical breaking point.
  • Error rate during hold. Errors are acceptable if they stop as soon as VUs drop. Errors that continue after the drop mean the system has not recovered.
  • Connection errors vs. application errors. HTTP 5xx from your application means a logic failure. TCP connection errors mean the system stopped accepting new connections at all, which is much worse.
  • Recovery time. Measure how long the error rate takes to return to zero after the VU drop. Under 60 seconds is good. Over 5 minutes points to connection pool or state corruption.

Sustained Cyber Monday traffic is 3×–4× your average daily traffic, all day. The stress test finds the concurrency ceiling: the VU count above which performance degrades past what you accept. You then confirm your peak sits comfortably below it.

execution:
- scenario: bfcm-stress
stages:
- duration: 3m
target: 340 # 25 % of modeled peak
- duration: 10m
target: 340
- duration: 3m
target: 680 # 50 % of modeled peak
- duration: 10m
target: 680
- duration: 3m
target: 1360 # 100 % — modeled peak
- duration: 10m
target: 1360
- duration: 3m
target: 2040 # 150 % — safety margin target
- duration: 10m
target: 2040
- duration: 3m
target: 2720 # 200 % — push to find the break
- duration: 10m
target: 2720
- duration: 5m
target: 0
  1. Create a new test named bfcm-2025-stress-ramp in the MaxoPerf console.
  2. Upload the YAML above. Set failure criteria: p95 > 2000ms or error rate > 5 %.
  3. Run the test. Do not stop early. The inflection point often appears at the 150 %–200 % stage.
  4. When the run completes (or auto-fails on criteria), open the Overview tab.
  5. Find the inflection point: the stage where p95 latency bends sharply upward or error rate rises above 0.5 %.

You want the inflection point at at least 150 % of your modeled peak. That leaves 50 % headroom above the traffic you plan for. If the inflection is at 110 % of modeled peak, you have almost no margin and must fix bottlenecks before the event.

After both tests, record:

TestBreaking-point VUsBreaking-point metricHeadroom above modeled peak
Spike (doorbuster)4,200 VUsTCP connection errors3.1× modeled peak
Stress (gradual ramp)2,100 VUsp95 payment > 2000ms1.54× modeled peak

If headroom is below 1.5× on either test, prioritize fixing the bottleneck before the dress rehearsal at T-1 week.

  • Soak and stability: once the breaking point is above your peak, check multi-hour stability at your modeled load.
  • Stress test: the full stress test reference page.
  • Spike test: the full spike test reference page.
  • War-room and runbook: use the breaking-point VU count to define your game-day abort criteria.