Soak and stability for multi-hour sales
A system that survives a 10-minute spike test may still fall over at hour 6 of Cyber Monday. This failure is gradual resource depletion over time, not pool exhaustion under sudden load. Memory that is never freed piles up. File descriptors that are never closed run out. Database connection pools that do not evict stale connections fill up. A query plan that was fast when the table had 10,000 rows gets slow at 50 million.
Soak tests are the only way to find these problems before your customers do.
Before you start
Section titled “Before you start”- The full-funnel load test from end-to-end journey load passes cleanly at your modeled peak VU count.
- The spike and stress tests from spike and stress for sales show at least 1.5× headroom above your modeled peak.
- Nobody else needs the target environment during the soak window, which is typically 6–12 hours.
- You capture application-level metrics (JVM heap, process RSS, open file descriptors, DB connection count) alongside the MaxoPerf run so you can correlate them.
The BFCM stability problem
Section titled “The BFCM stability problem”BFCM lasts far longer than a 15-minute test. Sustained Black Friday traffic runs for 18+ hours, and Cyber Monday adds more. A system that passes every 15-minute load test but has a slow memory leak will OOM-crash at hour 8. A connection pool tuned for daily traffic will run dry under sustained 3× load, over hours rather than minutes.
The soak test adds something the load test does not cover. It uses the same load for much longer and looks for slow drift instead of peak values.
What to look for: the drift indicators
Section titled “What to look for: the drift indicators”In a load or stress test you watch for peak values. In a soak test you watch the trend over time:
| Indicator | Healthy pattern | Concerning pattern |
|---|---|---|
| p95 latency | Flat across the full duration | Slow upward drift (>10 % over hours) |
| Throughput | Flat or very slight decline (<2 %) | Gradual throughput decline at constant VUs |
| Error rate | Near-zero throughout | Errors that appear only after 2–3 hours |
| Heap / RSS (app metric) | Stable or GC-bounded oscillation | Monotonically increasing, with no ceiling |
| DB connection count | Stable pool usage | Gradually reaching pool limit |
| Open file descriptors | Stable | Growing without bound |
A 10 % p95 drift over 8 hours is a signal. A 40 % drift, or errors that appear in the final 2 hours, needs investigation right away: something is leaking.
Soak test configuration
Section titled “Soak test configuration”Use the same VU count as your modeled peak load. Do not raise VUs to “make it more interesting”. That turns it into a stress test. The soak test asks one question: does the system stay stable at its operating point for the length of a real BFCM sale?
execution: - scenario: bfcm-soak concurrency: 1360 ramp-up: 5m hold-for: 8h
scenarios: bfcm-soak: think-time: 3s requests: - url: https://api.staging.example.com/v1/homepage label: GET /homepage think-time: 4s - url: https://api.staging.example.com/v1/products/${PRODUCT_ID} label: GET /product/:id think-time: 3s - url: https://api.staging.example.com/v1/cart/items label: POST /cart/items method: POST headers: Content-Type: application/json Authorization: Bearer ${SESSION_TOKEN} body: '{"product_id": "${PRODUCT_ID}", "qty": 1}' think-time: 5s - url: https://api.staging.example.com/v1/checkout/start label: POST /checkout/start method: POST headers: Authorization: Bearer ${SESSION_TOKEN}Name the test bfcm-2025-soak-8h and schedule it to start in the evening so the result is ready for review the next morning. See Schedule recurring tests.
Set failure criteria with generous thresholds, so the soak test auto-fails only on serious degradation:
p95 latency (any label) > 3000ms→ failerror rate > 2 %→ fail
Tighter criteria help with regression detection. They get in the way of an exploratory soak that looks for drift patterns.
Running the soak in MaxoPerf
Section titled “Running the soak in MaxoPerf”- Upload the YAML and configure the test in the console as described above.
- Use the Schedules tab to fire the run at a set time, typically 22:00 the night before you need the result.
- Do not watch the run live for 8 hours. MaxoPerf streams metrics continuously, and the full 8-hour timeline is there for review when the run completes.
- The next morning, open the Overview tab and look at the full time axis, not the current window. Use the time range selector to show the whole run.
Correlating soak results with application metrics
Section titled “Correlating soak results with application metrics”The soak run result tells you that something is leaking. Application metrics tell you what. When a soak result shows drift:
-
Note the time at which drift becomes visible in the p95 chart (e.g., “drift begins at t+4h”).
-
Open your APM or metrics dashboard and look at the same time window for: heap growth, GC frequency, DB connection pool usage, open file descriptor count.
-
The metric that trends upward in the same time window as p95 drift is the root cause.
-
Common causes and fixes:
Root cause Symptom Fix direction Memory leak in app code Heap grows continuously; GC pauses lengthen Fix object retention; heap dump analysis Connection pool not evicting DB connections approach pool limit Set idle eviction timeout and max-lifetime File descriptor leak FD count grows; eventually EMFILEerrorsClose resources in finallyblocks; audit stream handlingQuery plan regression under table growth p95 for DB-backed endpoints drifts Add index; ANALYZE/VACUUM; pagination Log rotation bloat Disk fills; latency spikes from blocked writes Rotate logs more frequently; async log writing
Multi-hour soak for the dress rehearsal
Section titled “Multi-hour soak for the dress rehearsal”At T-1 week, your dress rehearsal soak should match the expected real-event duration:
- Black Friday start to Cyber Monday end is 4 days. In practice, test the longest continuous high-traffic window, which is typically 18–24 hours.
- A 24-hour soak at modeled peak VUs is the best BFCM dress rehearsal.
- If runner budget rules out a 24-hour run, a 12-hour soak with clean results (no drift) is an acceptable stand-in.
Do / don’t
Section titled “Do / don’t”| Do | Don’t |
|---|---|
| Use the same VU count as your modeled peak | Increase VUs during the soak, which changes what you’re testing |
| Schedule the soak to run overnight and review in the morning | Sit watching a live dashboard for 8 hours |
| Correlate soak drift with application-level metrics | Dismiss gradual drift as “noise” |
| Run the soak after the stress test passes | Run a soak before establishing that the system handles peak load at all |
| Set failure criteria, even generous ones, so the run auto-fails on collapse | Rely on visual inspection of an 8-hour chart to detect a failure |
Where to go next
Section titled “Where to go next”- Soak / endurance test: the full soak test reference.
- Third-party and payment dependencies: include dependency call patterns in the soak scenario.
- War-room and runbook: use soak results to set drift thresholds for game-day alerting.
- BFCM readiness checklist: the soak is a required checklist item before go/no-go.