Latency percentiles deep dive
Latency is not a single number. Every request your system serves takes a slightly different amount of time. Network jitter slows some, garbage collection pauses slow others, and lock contention slows more. A useful performance measurement captures this whole distribution, not only its center.
Why average latency misleads
Section titled “Why average latency misleads”Suppose your API processes 1 000 requests in a minute. 990 of them complete in 50 ms. The remaining 10 hit a cache miss that causes a database round trip, and each takes 5 000 ms.
| Metric | Value |
|---|---|
| Average | ~98 ms |
| p50 (median) | 50 ms |
| p95 | 50 ms |
| p99 | 5 000 ms |
The average (98 ms) misleads in two directions. It makes the fast requests look slower than they are, and it hides the 1% of requests that are 100× slower. A user who lands in that 1% has your worst-case experience, and the average does not show it.
The p99 (5 000 ms) shows exactly that worst-case experience. At 1 000 requests per minute, 1% means 10 users per minute waiting five seconds.
What each percentile tells you
Section titled “What each percentile tells you”| Percentile | What it represents | When it matters |
|---|---|---|
| p50 (median) | Half of requests are faster, half are slower. The “typical” experience. | Good for understanding the center of the distribution. |
| p90 | 90% of requests complete within this time. | A common intermediate threshold. |
| p95 | 95% of requests complete within this time. | The standard SLO boundary for most web APIs. |
| p99 | 99% of requests complete within this time. | Critical for high-traffic APIs where 1% is still many users. |
| p99.9 | 99.9% of requests complete within this time. | Used in financial and safety-critical systems. |
MaxoPerf reports p50, p90, p95, and p99 on the Overview tab for every run. The glossary entry for p95 and p99 latency has the formal definitions.
Reading the percentile chart in MaxoPerf
Section titled “Reading the percentile chart in MaxoPerf”The latency chart on the run-detail Overview tab shows all four percentile lines over time. Read them like this:
Lines move together, similar gap throughout. The distribution shape is stable. The system responds consistently at all percentiles for the whole run. This is a healthy pattern.
Lines spread apart as the run progresses. Tail latency (p99) climbs faster than the median (p50). This often means a slow resource leak: a connection pool filling up, a memory cache hitting its limit, or a database seeing more contention over time.
p99 spikes, p50 stays flat. These are isolated slow requests. Possible causes: GC pause, cache miss storm, or an async task completing slowly behind one request.
All lines climb together. The system is approaching its capacity ceiling. Every request slows down, not only the outliers. Compare with the throughput chart. If throughput has plateaued, this is a clear capacity signal.
Percentiles in SLOs
Section titled “Percentiles in SLOs”A service level objective (SLO) for latency should name the percentile explicitly:
“p95 latency for the
POST /v1/ordersendpoint must be under 400 ms at 200 concurrent users.”
This is precise and testable. “Average latency under 200 ms” is neither. You can satisfy it with 990 fast requests and 10 very slow ones.
See Baselines, SLOs, and error budgets for how to turn percentile targets into MaxoPerf failure criteria.
Histogram vs percentiles
Section titled “Histogram vs percentiles”Some tools expose a latency histogram: the full distribution sorted into buckets (e.g., 0–50 ms, 51–100 ms, 101–250 ms). A histogram gives the most complete picture but is harder to scan quickly. Percentiles are a summary of the histogram: p95 is the response time that 95% of requests come in under.
MaxoPerf reports percentiles (the practical summary) instead of raw histograms. If you need the full histogram for detailed analysis, use the metrics export or a connected observability platform.
The impact of outliers on infrastructure planning
Section titled “The impact of outliers on infrastructure planning”At low traffic (100 RPM), a p99 of 5 000 ms means about 1 slow request per minute, which you can probably ignore. At high traffic (10 000 RPM), the same p99 means 100 slow requests per minute, each tying up a connection or thread for five seconds. Thread pools, connection limits, and request queues all feel the tail, not only the median.
Plan your infrastructure around the percentiles that matter at your real traffic level, not around what looks acceptable in a low-traffic smoke test.
Do / don’t
Section titled “Do / don’t”| Do | Don’t |
|---|---|
| Set SLOs on whichever of p95 or p99 your infrastructure and user experience require. | Set SLOs on average latency. |
| Look at percentile spread (gap between p50 and p99) to detect tail issues. | Celebrate a good p50 without checking p99. |
| Monitor all four percentile lines over the run duration to detect latency drift. | Only check the end-of-run summary number. The time-series shows patterns the summary hides. |
| Cross-reference latency climb with throughput. If both move, it’s capacity; if only p99 moves, it’s tail behavior. | Treat all latency climbs the same way. |
Where to go next
Section titled “Where to go next”- p95 and p99 latency: formal definition and quick reference.
- Core metrics explained: the full set of metrics on the Overview tab.
- Baselines, SLOs, and error budgets: how to turn percentile targets into testable SLOs.
- Anatomy of a load test result: a tour of every panel in a MaxoPerf run-detail page.
- Read run results and logs: practical guidance on reading a finished run in MaxoPerf.