Skip to content

Soak and stability for multi-hour sales

A system that survives a 10-minute spike test may still fall over at hour 6 of Cyber Monday. This failure is gradual resource depletion over time, not pool exhaustion under sudden load. Memory that is never freed piles up. File descriptors that are never closed run out. Database connection pools that do not evict stale connections fill up. A query plan that was fast when the table had 10,000 rows gets slow at 50 million.

Soak tests are the only way to find these problems before your customers do.

  • The full-funnel load test from end-to-end journey load passes cleanly at your modeled peak VU count.
  • The spike and stress tests from spike and stress for sales show at least 1.5× headroom above your modeled peak.
  • Nobody else needs the target environment during the soak window, which is typically 6–12 hours.
  • You capture application-level metrics (JVM heap, process RSS, open file descriptors, DB connection count) alongside the MaxoPerf run so you can correlate them.

BFCM lasts far longer than a 15-minute test. Sustained Black Friday traffic runs for 18+ hours, and Cyber Monday adds more. A system that passes every 15-minute load test but has a slow memory leak will OOM-crash at hour 8. A connection pool tuned for daily traffic will run dry under sustained 3× load, over hours rather than minutes.

The soak test adds something the load test does not cover. It uses the same load for much longer and looks for slow drift instead of peak values.

In a load or stress test you watch for peak values. In a soak test you watch the trend over time:

IndicatorHealthy patternConcerning pattern
p95 latencyFlat across the full durationSlow upward drift (>10 % over hours)
ThroughputFlat or very slight decline (<2 %)Gradual throughput decline at constant VUs
Error rateNear-zero throughoutErrors that appear only after 2–3 hours
Heap / RSS (app metric)Stable or GC-bounded oscillationMonotonically increasing, with no ceiling
DB connection countStable pool usageGradually reaching pool limit
Open file descriptorsStableGrowing without bound

A 10 % p95 drift over 8 hours is a signal. A 40 % drift, or errors that appear in the final 2 hours, needs investigation right away: something is leaking.

Use the same VU count as your modeled peak load. Do not raise VUs to “make it more interesting”. That turns it into a stress test. The soak test asks one question: does the system stay stable at its operating point for the length of a real BFCM sale?

execution:
- scenario: bfcm-soak
concurrency: 1360
ramp-up: 5m
hold-for: 8h
scenarios:
bfcm-soak:
think-time: 3s
requests:
- url: https://api.staging.example.com/v1/homepage
label: GET /homepage
think-time: 4s
- url: https://api.staging.example.com/v1/products/${PRODUCT_ID}
label: GET /product/:id
think-time: 3s
- url: https://api.staging.example.com/v1/cart/items
label: POST /cart/items
method: POST
headers:
Content-Type: application/json
Authorization: Bearer ${SESSION_TOKEN}
body: '{"product_id": "${PRODUCT_ID}", "qty": 1}'
think-time: 5s
- url: https://api.staging.example.com/v1/checkout/start
label: POST /checkout/start
method: POST
headers:
Authorization: Bearer ${SESSION_TOKEN}

Name the test bfcm-2025-soak-8h and schedule it to start in the evening so the result is ready for review the next morning. See Schedule recurring tests.

Set failure criteria with generous thresholds, so the soak test auto-fails only on serious degradation:

  • p95 latency (any label) > 3000ms → fail
  • error rate > 2 % → fail

Tighter criteria help with regression detection. They get in the way of an exploratory soak that looks for drift patterns.

  1. Upload the YAML and configure the test in the console as described above.
  2. Use the Schedules tab to fire the run at a set time, typically 22:00 the night before you need the result.
  3. Do not watch the run live for 8 hours. MaxoPerf streams metrics continuously, and the full 8-hour timeline is there for review when the run completes.
  4. The next morning, open the Overview tab and look at the full time axis, not the current window. Use the time range selector to show the whole run.

Correlating soak results with application metrics

Section titled “Correlating soak results with application metrics”

The soak run result tells you that something is leaking. Application metrics tell you what. When a soak result shows drift:

  1. Note the time at which drift becomes visible in the p95 chart (e.g., “drift begins at t+4h”).

  2. Open your APM or metrics dashboard and look at the same time window for: heap growth, GC frequency, DB connection pool usage, open file descriptor count.

  3. The metric that trends upward in the same time window as p95 drift is the root cause.

  4. Common causes and fixes:

    Root causeSymptomFix direction
    Memory leak in app codeHeap grows continuously; GC pauses lengthenFix object retention; heap dump analysis
    Connection pool not evictingDB connections approach pool limitSet idle eviction timeout and max-lifetime
    File descriptor leakFD count grows; eventually EMFILE errorsClose resources in finally blocks; audit stream handling
    Query plan regression under table growthp95 for DB-backed endpoints driftsAdd index; ANALYZE/VACUUM; pagination
    Log rotation bloatDisk fills; latency spikes from blocked writesRotate logs more frequently; async log writing

At T-1 week, your dress rehearsal soak should match the expected real-event duration:

  • Black Friday start to Cyber Monday end is 4 days. In practice, test the longest continuous high-traffic window, which is typically 18–24 hours.
  • A 24-hour soak at modeled peak VUs is the best BFCM dress rehearsal.
  • If runner budget rules out a 24-hour run, a 12-hour soak with clean results (no drift) is an acceptable stand-in.
DoDon’t
Use the same VU count as your modeled peakIncrease VUs during the soak, which changes what you’re testing
Schedule the soak to run overnight and review in the morningSit watching a live dashboard for 8 hours
Correlate soak drift with application-level metricsDismiss gradual drift as “noise”
Run the soak after the stress test passesRun a soak before establishing that the system handles peak load at all
Set failure criteria, even generous ones, so the run auto-fails on collapseRely on visual inspection of an 8-hour chart to detect a failure