LLM performance and load testing
A load test of a language model endpoint is an HTTP load test. The metrics you care about differ from an ordinary REST API, though. This page shows how to run a structured LLM load test in MaxoPerf, which AI metrics to collect, and how to read the run-detail tabs.
Before you start
Section titled “Before you start”- Read HTTP and REST load testing. The request mechanics are the same.
- Read Taurus fundamentals or k6 scripts on Maxoperf, depending on the engine you use.
- You need an LLM endpoint that accepts HTTP POST requests with a JSON body (OpenAI-compatible chat-completions format or a custom inference API). It does not have to be a commercial API. Self-hosted inference servers (vLLM, Ollama, TGI) work the same way.
AI-specific metrics
Section titled “AI-specific metrics”Standard load tests measure response time and throughput. LLM endpoints add metrics about tokens:
| Metric | Definition | Why it matters |
|---|---|---|
| TTFT (time to first token) | Latency from request sent to first byte of the response body received | Sets how responsive a streaming UI feels |
| TPOT (time per output token) | Time between successive tokens in a streaming response | Affects reading pace for streamed output |
| End-to-end latency | Time from request sent to full response received | Total user wait for non-streaming calls |
| Throughput (RPS) | Requests completed per second across all VUs | Capacity of the inference stack |
| Tokens/sec (output) | Output tokens generated per second across all VUs | GPU/inference throughput; correlates with cost |
| Concurrency | Simultaneous in-flight requests | Maps to number of parallel GPU contexts |
| p95 / p99 latency | 95th / 99th percentile of end-to-end latency | Tail behaviour under real concurrent load |
| Error rate | Fraction of non-2xx responses | Includes 429 rate-limit, 503 overload, 504 timeout |
| 429 rate | Fraction of rate-limit rejections specifically | Reveals provisioned quota vs demand gap |
| Cost per request | Token count × price per token for the model | Lets you project monthly cost from the load profile |
A real Taurus YAML for an LLM endpoint
Section titled “A real Taurus YAML for an LLM endpoint”execution: - concurrency: 20 # 20 simultaneous inference requests ramp-up: 2m # gentle ramp — GPU workers need warm-up time hold-for: 5m # sustained plateau to capture steady-state metrics scenario: llm-chat
scenarios: llm-chat: requests: - label: chat-completions url: https://inference.example.com/v1/chat/completions method: POST headers: Content-Type: application/json Authorization: "Bearer ${LLM_API_KEY}" # injected via MaxoPerf secrets body: > { "model": "llama-3-8b-instruct", "messages": [ {"role": "user", "content": "Summarise the benefits of load testing in three sentences."} ], "max_tokens": 256, "temperature": 0.7, "stream": false } assert: - contains: subject: http-code value: '200' - contains: subject: body value: '"choices"'What this YAML does:
concurrency: 20runs 20 users sending inference requests at the same time. Start low: LLM inference is GPU-bound and saturates quickly.ramp-up: 2mis a slow ramp. It gives the KV cache time to warm up and lets autoscaling react if you have it on.max_tokens: 256fixes the output budget so response times are comparable across runs. A varying output length makes latency comparisons unreliable.stream: falsereturns full responses for this first test. See Streaming and token latency testing for SSE variants.- The
Authorizationheader reads a MaxoPerf secret. Never hard-code API keys. See Manage test secrets.
A real k6 script for an LLM endpoint
Section titled “A real k6 script for an LLM endpoint”import http from 'k6/http';import { check, sleep } from 'k6';
export const options = { stages: [ { duration: '2m', target: 20 }, // ramp to 20 VUs { duration: '5m', target: 20 }, // hold plateau { duration: '1m', target: 0 }, // ramp down ], thresholds: { // p95 end-to-end latency under 3 s (adjust for your model/hardware) http_req_duration: ['p(95)<3000'], // error rate under 2% http_req_failed: ['rate<0.02'], },};
const LLM_ENDPOINT = 'https://inference.example.com/v1/chat/completions';const API_KEY = __ENV.LLM_API_KEY; // injected from MaxoPerf secrets
const PAYLOAD = JSON.stringify({ model: 'llama-3-8b-instruct', messages: [ { role: 'user', content: 'Summarise the benefits of load testing in three sentences.' }, ], max_tokens: 256, temperature: 0.7, stream: false,});
export default function () { const params = { headers: { 'Content-Type': 'application/json', 'Authorization': `Bearer ${API_KEY}`, }, timeout: '30s', // LLM inference can be slow; set an explicit timeout };
const res = http.post(LLM_ENDPOINT, PAYLOAD, params);
check(res, { 'status 200': (r) => r.status === 200, 'has choices field': (r) => r.json('choices') !== undefined, 'not rate-limited': (r) => r.status !== 429, });
// No sleep — LLM inference already has long natural think time}How to create and run the test in MaxoPerf
Section titled “How to create and run the test in MaxoPerf”-
In the MaxoPerf console, go to Tests → New test. Give the test a name such as
llm-chat-load-20vu. -
Open the Files tab. Upload your Taurus YAML (or k6
.jsfile) and mark it as the Entrypoint. MaxoPerf auto-detects the engine. -
Open Settings → Secrets and add
LLM_API_KEYwith your inference API key. MaxoPerf injects it as${LLM_API_KEY}at run time, so it never appears in the uploaded file. -
Open the Configuration tab. Set Virtual users to
20, Ramp-up to2m, and Duration to8m(2 m ramp + 5 m hold + 1 m ramp-down). Pick the runner location closest to the inference server so network time adds as little as possible to the measurements. -
In the Failure criteria section, add:
p95 latency > 3000 ms → failandError rate > 2% → fail. -
Click Run. The run detail page opens automatically.
How to read the results
Section titled “How to read the results”Overview tab
Section titled “Overview tab”- Throughput (RPS). For an LLM endpoint with
max_tokens: 256, expect 1–5 RPS per VU, depending on the model and hardware. If throughput levels off below the RPS you expect, the model is probably GPU-saturated. - p95 / p99 latency. LLM latency is orders of magnitude higher than a REST API. A p95 of 2–5 seconds is normal for a 7B–13B parameter model at moderate concurrency. Watch for a latency curve that keeps climbing and never levels off. It means the queue grows faster than the model can clear it.
- Error rate. 429s appear when the endpoint’s rate-limit quota is exceeded. 503/504s appear when the inference server is overloaded or out of memory.
Errors tab
Section titled “Errors tab”- A cluster of 429 errors early in the run points to a quota problem, not a performance problem. Contact your LLM API provider or adjust your rate limits.
- 504 timeouts under load mean the inference server is saturating. Reduce concurrency or add GPU capacity.
- 500 errors with
out of memoryin the body mean the KV cache is exhausted. Reducemax_tokensor add GPU nodes.
Do / don’t
Section titled “Do / don’t”| Do | Don’t |
|---|---|
Fix max_tokens so response lengths (and timings) are comparable across runs | Leave max_tokens open-ended; varying output length makes p95 meaningless |
Pass Authorization through MaxoPerf secrets and never hard-code API keys | Put API keys in uploaded test files |
| Start with 5–10 VUs and increase; LLM endpoints saturate fast | Jump to 100 VUs before you know the saturation point |
Label requests clearly (label: chat-completions) for a per-endpoint breakdown | Use unlabelled requests. The results panel then shows unnamed instead of a useful name |
| Set an explicit HTTP timeout (30–60 s) in k6 | Rely on the default 10 s timeout; inference often takes longer |
| Run against staging / non-production inference endpoints | Load-test production LLM APIs without provider approval and a rate-limit increase |
Where to go next
Section titled “Where to go next”- Streaming and token latency testing: measure TTFT precisely with SSE/chunked responses.
- Inference cost and token budget testing: project monthly cost from this load profile.
- RAG pipeline load testing: load-test retrieval + generation as one end-to-end endpoint.
- AI API scalability: stress, spike, and soak: push the endpoint past saturation.
- AI load test failure criteria: automate pass/fail gating on p95 TTFT and error rate.
- Stress test: stress testing in general, applied here.
- Cookbook: Failure criteria pass/fail gates: configure thresholds in the console.