Skip to content

Interpreting results dos and don'ts

A load test is only as useful as the conclusions you draw from it. Teams often misread results. They ship a “passing” build that hides real performance problems, or reject a healthy system because of a misleading metric. This page pairs concrete dos and don’ts with the MaxoPerf panels where each one matters.

The core principle: look at distributions, not point estimates

Section titled “The core principle: look at distributions, not point estimates”

A single number, whether the average, the max or a snapshot RPS, tells you almost nothing about how your system behaves at scale. Load testing is about distributions: how latency spreads across all requests, how errors spread across locations, and how throughput spreads across runners.

  • Do use p95 or p99 as the anchor for your latency SLO. The p95 is the worst experience 1 in 20 users has. The p99 covers the worst 1 in 100. These are the numbers that decide user-facing reliability. In MaxoPerf, the Overview tab’s latency chart shows all four percentile lines (p50, p90, p95, p99). Watch all four.

  • Do compare percentile lines over time, not only end-of-run averages. A rising p99 trend halfway through a soak test often predicts a failure that the post-run average would never show. Use the latency time-series chart, not only the summary table.

  • Do set failure criteria on a percentile threshold, not on the average. A failure criterion of p95(http_req_duration) < 500ms is precise. A criterion on the average can pass even when 10 % of users see 2-second responses.

  • Do check the error-rate panel next to latency. When latency drops during high load, slow requests are often being dropped (timing out or load-shed). Performance has not improved. Compare the latency trend with the error rate.

  • Do run long enough to get a stable steady-state window. The first 2–3 minutes of any run include JVM warm-up, connection pool growth and CDN cache priming. Do not measure only the first minute. Aim for at least 5–10 minutes of stable load after the ramp-up ends before you draw conclusions.

  • Do check that your request count is statistically meaningful. A p99 from 100 requests is meaningless, because any single outlier decides it. At 100 VUs with a 1-second think time, a 10-minute run produces ~60,000 requests. That is a meaningful sample. At 5 VUs for 60 seconds, you have ~300 requests, and the p99 from that run is noise.

  • Do compare runs against a known baseline. A result on its own is hard to judge. The comparing runs and baselines recipe shows how to set a baseline run in MaxoPerf and see delta indicators for p95 latency and error rate between the baseline and the current run.

  • Do check the Runners tab during and after every run. Runner saturation happens when a runner’s own CPU, memory or network becomes the bottleneck. The results look like a slow application but come from the load-generation infrastructure. Signs of runner saturation: p99 rises with no matching rise in error rate, throughput stays flat as VUs increase, and the runner health column shows CPU warnings.

  • Do confirm that all configured runners reported healthy. If a runner in one region reported errors or went into a degraded state, the results for that region are incomplete. Filter the Overview tab by location to spot problems in one region.

  • Do check runner count against your VU target. Each runner has a recommended VU capacity. With 2,000 VUs across 2 runners, each runner carries 1,000 VUs. Confirm that is within the healthy operating range for your configuration.

  • Don’t use the average (mean) latency as your headline SLO metric. The many fast requests pull the average down and hide the tail. A p50 of 80 ms average can sit next to a p99 of 3,000 ms. See Latency percentiles.

  • Don’t judge a test by the minimum latency. The minimum is almost always a cache hit or one lucky request. It tells you nothing about typical user experience.

  • Don’t compare two runs if their load profiles differ. A run at 100 VUs and a run at 500 VUs are not comparable. Use the same profile, ramp, duration and think time before you call something a regression or an improvement.

  • Don’t measure only the ramp-up period. During ramp-up, the VU count is climbing and results read low. Wait for the steady-state plateau, then measure.

  • Don’t run a soak test for less than an hour. Memory leaks, connection pool exhaustion and log file growth build up slowly. A 10-minute run will not show them. The overnight soak test recipe recommends a minimum 8-hour window.

  • Don’t declare a “pass” because the test finished without errors. A test with 0 errors at 10 VUs tells you nothing about behaviour at 500 VUs. Make sure your load level matches the target you are testing against.

  • Don’t report runner-saturated results as application results. If your Runners tab shows a runner in a degraded or warning state, re-run with fewer VUs per runner before you draw conclusions. Reduce the VU count or add more runners.

  • Don’t skip the Runners tab because the test finished. A runner can be saturated without raising explicit errors. Check the CPU and memory indicators for every runner, especially on high-VU or long-duration tests.

SignalWhere in MaxoPerfWhat to look for
p50 / p95 / p99 latencyOverview tab → Latency chartRising trend, diverging percentile lines
Error rateOverview tab → Error rate panelNon-zero rate, spikes coinciding with VU increases
Throughput (RPS)Overview tab → Throughput chartPlateau or drop that does not match VU count
Runner healthRunners tabDegraded/Error state, CPU/memory warnings
Per-location resultsOverview tab → Location filterRegional outliers
Request-level errorsLog tabSpecific status codes, messages, URLs