Skip to content

AI load testing — do and don't

These do/don’t pairs come from running load tests against LLM inference endpoints, RAG pipelines and streaming AI APIs. Each pair is based on a real failure mode. The links point to the page that covers the topic in detail.


Do: measure TTFT and end-to-end latency as separate metrics

For streaming endpoints, time to first token (TTFT) and total generation time measure different things for the user. TTFT tells you when the user first sees a response. Total latency tells you when they can act on it. If you fold both into one http_req_duration metric, a long generation tail hides TTFT regressions.

Use a k6 custom Trend metric (llm_ttft_ms) to capture TTFT separately from http_req_duration. Set independent failure criteria thresholds for each.

See Streaming and token latency testing.

Don’t: treat non-streaming latency as a proxy for streaming UX

A non-streaming call waits for the full response before it returns. Users of a streaming UI feel TTFT (typically < 1 s), not total latency (typically 5–20 s). A non-streaming load test can show a healthy p95 of 4 s while the streaming TTFT p95 is 2.5 s, which is too slow for a chat UI. Test the streaming path with streaming turned on.

Do: set an explicit HTTP timeout of 30–60 seconds for LLM requests

The k6 default timeout is 10 seconds. An LLM generating 512 tokens at 30 tokens/s takes ~17 s. Without an explicit timeout, those long requests count as errors. Set timeout: '60s' in your k6 options, or the equivalent in Taurus.


Do: fix max_tokens in every load-test run

Output token count drives LLM response time more than anything else. If max_tokens varies across requests (or stays at the model default), your p95 latency measures a mix of short and long generations instead of your real workload. Set max_tokens to the value your production use case needs, and record it with the run.

Don’t: use a single repeated prompt for all virtual users

A single prompt takes the same inference path every time, and many inference servers answer it from the KV cache. The result is faster than production traffic will ever be. Use a CSV of realistic prompts: at least 50–100 distinct inputs for a basic test, and 200+ for a RAG test so the vector-DB cache does not skew it.

See RAG pipeline load testing.

Do: model realistic prompt sizes

A 20-token prompt and a 1 000-token RAG-context prompt hit the same endpoint, but their input token costs and TTFT values differ widely. Run your test at the prompt size your production workload uses. If your app uses long context windows (e.g., injecting retrieved documents), test with realistic context lengths, not the shortest happy-path prompt.

See Inference cost and token budget testing.


Do: start with conservative concurrency (5–10 VUs) and ramp slowly

LLM inference is GPU-bound. A 20-GPU node that handles 30 VUs well may saturate at 35 VUs, with a sharp latency jump and a burst of errors. Start at low concurrency, watch the steady-state metrics, then add VUs in steps. A 2-minute ramp-up for every 20 additional VUs gives GPU workers time to settle.

Don’t: jump straight to 100 VUs before knowing the saturation point

A test that starts at 100 VUs with no baseline gives you one data point: “the system is saturated.” It tells you nothing about the saturation knee, the highest concurrency you can hold before latency degrades. Find the knee first with a stepped stress test, then set production concurrency below it.

See AI API scalability: stress, spike, and soak.

Do: right-size concurrency to your production workload

Concurrency in a load test is different from user count. For LLM endpoints where each request takes 2–5 s, 20 concurrent users means 20 requests in flight at the same time. That is far more load than 20 interactive users, who spend most of their time reading. Set the VU count from how many inferences you expect at the same moment, not from your active-user count.


Do: test streaming at the cadence your app expects

If your application streams tokens to a chat UI, run your load test with "stream": true and parse the SSE events. A non-streaming test checks batch throughput but skips the streaming path. Streaming also holds HTTP connections open, a cost that non-streaming tests never show.

Don’t: use the default k6 10 s timeout for streaming responses

Streaming a 512-token response at 30 tokens/s takes ~17 s. Set timeout: '60s' (or longer for large max_tokens values). A timeout that fires mid-stream shows up as an error. It inflates the error rate and makes failure criteria checks fail for the wrong reason.


Do: track token consumption in every load-test run

For AI endpoints, latency data without token counts is only half the result. A run that cut p95 latency by 20% because average output tokens fell by 40% looks like an improvement. It may mean the model is truncating its output. Always record prompt_tokens, completion_tokens and the total from the usage field, and emit them as custom k6 metrics.

Do: project monthly cost from your load profile before launch

Multiply (requests/sec) × (seconds/month) × (cost/request) to get monthly spend. Do this before you launch. LLM inference cost grows with traffic faster than most teams expect: 8 RPS at $0.00035/request = $7,200/month. Run the calculation from your load-test token data.

See Inference cost and token budget testing.

Don’t: treat 429 responses as free

When a hosted LLM API rate-limits you (HTTP 429), the request failed. If your application retries automatically, those retries use quota and cost money. Under load, a rate-limit cascade can push real cost above your estimate, because every failed request creates more retry requests. Track the 429 rate as a separate failure criterion.


Do: watch 429 rate as a distinct failure criterion

429 rate-limit errors need a different fix from 503 server errors. A cluster of 429s means the quota ran out: raise your rate-limit tier or lower concurrency. A cluster of 503s means the inference server is overloaded: add capacity. One combined “error rate” metric hides which of the two you have.

See AI load test failure criteria.

Don’t: run high-concurrency load tests against production LLM APIs without quota headroom

A 100-VU load test against an LLM API with a 60-RPM quota uses up the quota at once. Check your provider’s TPM/RPM limits before you run. To test at production concurrency, request a quota increase or use a staging API key with a higher limit. Don’t find the limit by flooding production.


Do: account for model warm-up in your ramp-up design

Many inference servers, especially self-hosted ones on vLLM, TGI or Ollama, have GPU workers that build KV-cache structures on first use. With a freshly loaded model, the first 30–60 seconds of a test can show high TTFT that drops as the worker pool warms up. Keep the ramp-up at a 2-minute minimum and read steady-state metrics from the plateau, not the ramp.

Do: test the cold-start path deliberately with a spike test

If your system autoscales on GPU demand, test what happens when traffic jumps from zero to production concurrency. Autoscaling an inference cluster takes 3–10 minutes. During that window, users get 503 errors or 20–30 s TTFT while the new GPU worker starts. A spike test measures this cold-start window so you can decide whether to keep workers warm.

See AI API scalability: stress, spike, and soak.


Do: gate every AI load run with explicit failure criteria

A run that ends with status “Finished” has completed, nothing more. That does not mean it passed. Write your SLOs as failure criteria: p95 TTFT > 800 ms → fail, error rate > 2% → fail, throughput < 3 RPS → fail. The run then gets an automatic Passed / Failed verdict that CI pipelines can read without a person checking.

Don’t: set criteria tighter than your measured baseline on the first run

If you set p95 latency > 2 000 ms → fail before you have ever measured the model’s p95, you will either always pass (criteria too loose) or always fail (criteria too tight). Run three to five baseline runs without criteria first to learn the normal range. Then set thresholds at 120–150% of baseline p95.

See AI load test failure criteria.

Do: compare runs after every model upgrade, prompt change, or infrastructure change

Use MaxoPerf run comparison to get a measured Δ between the candidate and the baseline. A model upgrade that improves quality but adds 40% to p95 TTFT may break your latency SLO. Comparing every time catches slow, gradual regressions that a threshold-only gate misses.