Skip to content

Common pitfalls and anti-patterns

Experienced teams still make load testing mistakes that leave their results useless or misleading. This page lists the most common anti-patterns, explains why each one produces bad results, and shows what to do instead in a MaxoPerf workflow.

Anti-pattern 1: Testing in production without authorization or coordination

Section titled “Anti-pattern 1: Testing in production without authorization or coordination”

What it looks like: You run a load test straight against a production URL. You have not told your operations team, you have no maintenance window, and you have not confirmed you are authorized to generate that load.

Why it is harmful: To your system, a load test at scale looks the same as a DDoS attack. It can cause real outages for real users, trigger rate-limiter blocks that hit production traffic, exhaust connection pools, and run up bills (cloud auto-scaling, CDN egress, SMS/email notifications).

What to do instead: Test in a staging or dedicated performance environment. If you must test in production, coordinate with your SRE team, set a maintenance window, and get explicit authorization. See Safe and ethical load testing.


Anti-pattern 2: The “ramp to infinity” test mistaken for a load test

Section titled “Anti-pattern 2: The “ramp to infinity” test mistaken for a load test”

What it looks like: You set VUs to the maximum your account allows and call the result a “load test”. There is no target throughput, no steady-state window and no SLO to measure against.

Why it is harmful: This is a breakpoint/capacity test, not a load test. It tells you where your system breaks, not whether it meets your SLO at your expected peak. If you use it as a load test, you have no bar for “passing”.

What to do instead: Define a specific target load (e.g. 500 RPS, 200 concurrent VUs) from your production traffic. Run a load test at that target with a stability window. Use failure criteria to turn your SLO into a pass/fail gate.


Anti-pattern 3: Measuring only the ramp-up period

Section titled “Anti-pattern 3: Measuring only the ramp-up period”

What it looks like: A 5-minute test with a 5-minute ramp-up. You measure only while VUs are still climbing, with no steady-state window.

Why it is harmful: During ramp-up, JVM warm-up, DNS caching, connection pool growth and CDN priming all happen at once. Latency in this phase reads higher than at steady state, and throughput reads lower.

What to do instead: Build your test as ramp-up period + steady-state window + optional ramp-down. Measure only the steady-state window. A common pattern:

execution:
- ramp-up: 3m
hold-for: 10m # measure this window
ramp-down: 2m
concurrency: 200

Anti-pattern 4: Single-user, single-path testing

Section titled “Anti-pattern 4: Single-user, single-path testing”

What it looks like: 100 VUs all hitting GET /api/products/42 in an infinite loop.

Why it is harmful: Every layer (CDN, app cache, DB query cache) caches that one resource ID. Your results show cache-hit performance, not real-user performance. The product team ships the feature, the cache warms in production, and the first uncached request is 10× slower than your test showed.

What to do instead: Use a CSV data entity to vary the resource ID (and user session) per VU. Match your real traffic mix: several endpoints and several user paths. See Test data dos and don’ts and Designing realistic load.


Anti-pattern 5: Ignoring runner saturation

Section titled “Anti-pattern 5: Ignoring runner saturation”

What it looks like: You push 2,000 VUs across 2 runners and report the p99 latency as an application metric.

Why it is harmful: Each MaxoPerf runner has a recommended VU capacity. A saturated runner queues requests and adds latency in the load generator, before the request reaches your application. Your results then show a runner bottleneck, not an application bottleneck.

What to do instead: Check the Runners tab during every run. If any runner shows a degraded or warning state, or if p99 climbs with no matching rise in error rate, suspect runner saturation. Re-run with more runners (or fewer VUs per runner) and compare results.


Anti-pattern 6: The “passed last month, must still pass” assumption

Section titled “Anti-pattern 6: The “passed last month, must still pass” assumption”

What it looks like: A load test passed 6 months ago, so the team skips load testing for the next 10 releases.

Why it is harmful: Every release changes the system. Database schemas change, new middleware arrives, and third-party dependencies change how they perform. A 6-month-old load test result says little about how the system behaves today.

What to do instead: Run baseline regression tests in your CI pipeline with MaxoPerf scheduled tests and failure criteria. A regression test on every deploy takes minutes and catches performance regressions before they reach production.


Anti-pattern 7: Averaging percentiles across runs

Section titled “Anti-pattern 7: Averaging percentiles across runs”

What it looks like: You run three load tests and average their p99 latencies to get a “reliable” p99.

Why it is harmful: Averaging percentiles is mathematically wrong. The p99 of [p99_run1, p99_run2, p99_run3] is not the p99 of the combined distribution. The number you get means nothing and you cannot reproduce it.

What to do instead: Use the MaxoPerf comparison view to compare individual runs directly. To summarize several runs, list their p99 values one by one and note the range (min, max) instead of averaging them.


Anti-pattern 8: Load testing without failure criteria

Section titled “Anti-pattern 8: Load testing without failure criteria”

What it looks like: You run a test, glance at the charts, and decide it “looks fine”.

Why it is harmful: “Looks fine” varies between reviewers, gets sloppier under time pressure, and cannot run in a CI pipeline. Teams ship regressions because the chart “looked fine” next to the last run they remembered.

What to do instead: Define failure criteria for every test. At minimum:

  • p95(http_req_duration) < <your SLO threshold>
  • rate(errors) < 0.01 (1 % error budget)

MaxoPerf evaluates these after the run and sets the run status to Failed if any criterion is violated. Your CI pipeline gets a deterministic pass/fail signal to act on.


Anti-pattern 9: Not waiting for the system to stabilize before the test

Section titled “Anti-pattern 9: Not waiting for the system to stabilize before the test”

What it looks like: You deploy a new version and start a load test right away, before the application has filled its connection pool, loaded its caches or finished its startup routines.

Why it is harmful: Your first-minute results include cold-start effects that push latency up and throughput down. You may report a regression that does not exist at steady state.

What to do instead: Wait 2–5 minutes after deployment before you start a load test. Add a short warm-up script (a smoke test at low VUs) as the first step in your CI pipeline, before the full load run.


Anti-pattern 10: Treating a failed smoke test as a passing load test

Section titled “Anti-pattern 10: Treating a failed smoke test as a passing load test”

What it looks like: The smoke test (5 VUs, 60 seconds) shows 5 % errors, but the team runs the full load test anyway and reports its result.

Why it is harmful: If your system already errors at 5 VUs, the errors at 500 VUs are noise. They tell you nothing about failures caused by load. You have a functional bug to fix first.

What to do instead: Make a passing smoke test a hard gate before any load test. In MaxoPerf, use the pre-flight checklist and confirm the smoke test is green before you go on.


Anti-patternRoot causeFix
Testing prod without authorizationNo processDedicated perf environment + SRE coordination
Ramp-to-infinity as load testNo target definedDefine VU/RPS target from production traffic
Measuring only ramp-upTest designAdd steady-state hold period
Single path / single IDNo data varietyCSV data entity with per-VU variation
Runner saturation ignoredNo Runners tab checkMonitor Runners tab; adjust VU/runner ratio
Stale results trustedNo regression cadenceCI-gated scheduled baseline tests
Averaged percentilesMisunderstanding of statsCompare runs individually
No failure criteriaManual review onlySet p95 + error-rate criteria on every test
Cold-start artefactsPremature test startWarm up with smoke test before load run
Load test after failing smokeWrong gate orderSmoke gate → load test