Skip to content

LLM performance and load testing

A load test of a language model endpoint is an HTTP load test. The metrics you care about differ from an ordinary REST API, though. This page shows how to run a structured LLM load test in MaxoPerf, which AI metrics to collect, and how to read the run-detail tabs.

  • Read HTTP and REST load testing. The request mechanics are the same.
  • Read Taurus fundamentals or k6 scripts on Maxoperf, depending on the engine you use.
  • You need an LLM endpoint that accepts HTTP POST requests with a JSON body (OpenAI-compatible chat-completions format or a custom inference API). It does not have to be a commercial API. Self-hosted inference servers (vLLM, Ollama, TGI) work the same way.

Standard load tests measure response time and throughput. LLM endpoints add metrics about tokens:

MetricDefinitionWhy it matters
TTFT (time to first token)Latency from request sent to first byte of the response body receivedSets how responsive a streaming UI feels
TPOT (time per output token)Time between successive tokens in a streaming responseAffects reading pace for streamed output
End-to-end latencyTime from request sent to full response receivedTotal user wait for non-streaming calls
Throughput (RPS)Requests completed per second across all VUsCapacity of the inference stack
Tokens/sec (output)Output tokens generated per second across all VUsGPU/inference throughput; correlates with cost
ConcurrencySimultaneous in-flight requestsMaps to number of parallel GPU contexts
p95 / p99 latency95th / 99th percentile of end-to-end latencyTail behaviour under real concurrent load
Error rateFraction of non-2xx responsesIncludes 429 rate-limit, 503 overload, 504 timeout
429 rateFraction of rate-limit rejections specificallyReveals provisioned quota vs demand gap
Cost per requestToken count × price per token for the modelLets you project monthly cost from the load profile
execution:
- concurrency: 20 # 20 simultaneous inference requests
ramp-up: 2m # gentle ramp — GPU workers need warm-up time
hold-for: 5m # sustained plateau to capture steady-state metrics
scenario: llm-chat
scenarios:
llm-chat:
requests:
- label: chat-completions
url: https://inference.example.com/v1/chat/completions
method: POST
headers:
Content-Type: application/json
Authorization: "Bearer ${LLM_API_KEY}" # injected via MaxoPerf secrets
body: >
{
"model": "llama-3-8b-instruct",
"messages": [
{"role": "user", "content": "Summarise the benefits of load testing in three sentences."}
],
"max_tokens": 256,
"temperature": 0.7,
"stream": false
}
assert:
- contains:
subject: http-code
value: '200'
- contains:
subject: body
value: '"choices"'

What this YAML does:

  • concurrency: 20 runs 20 users sending inference requests at the same time. Start low: LLM inference is GPU-bound and saturates quickly.
  • ramp-up: 2m is a slow ramp. It gives the KV cache time to warm up and lets autoscaling react if you have it on.
  • max_tokens: 256 fixes the output budget so response times are comparable across runs. A varying output length makes latency comparisons unreliable.
  • stream: false returns full responses for this first test. See Streaming and token latency testing for SSE variants.
  • The Authorization header reads a MaxoPerf secret. Never hard-code API keys. See Manage test secrets.
import http from 'k6/http';
import { check, sleep } from 'k6';
export const options = {
stages: [
{ duration: '2m', target: 20 }, // ramp to 20 VUs
{ duration: '5m', target: 20 }, // hold plateau
{ duration: '1m', target: 0 }, // ramp down
],
thresholds: {
// p95 end-to-end latency under 3 s (adjust for your model/hardware)
http_req_duration: ['p(95)<3000'],
// error rate under 2%
http_req_failed: ['rate<0.02'],
},
};
const LLM_ENDPOINT = 'https://inference.example.com/v1/chat/completions';
const API_KEY = __ENV.LLM_API_KEY; // injected from MaxoPerf secrets
const PAYLOAD = JSON.stringify({
model: 'llama-3-8b-instruct',
messages: [
{ role: 'user', content: 'Summarise the benefits of load testing in three sentences.' },
],
max_tokens: 256,
temperature: 0.7,
stream: false,
});
export default function () {
const params = {
headers: {
'Content-Type': 'application/json',
'Authorization': `Bearer ${API_KEY}`,
},
timeout: '30s', // LLM inference can be slow; set an explicit timeout
};
const res = http.post(LLM_ENDPOINT, PAYLOAD, params);
check(res, {
'status 200': (r) => r.status === 200,
'has choices field': (r) => r.json('choices') !== undefined,
'not rate-limited': (r) => r.status !== 429,
});
// No sleep — LLM inference already has long natural think time
}

How to create and run the test in MaxoPerf

Section titled “How to create and run the test in MaxoPerf”
  1. In the MaxoPerf console, go to Tests → New test. Give the test a name such as llm-chat-load-20vu.

  2. Open the Files tab. Upload your Taurus YAML (or k6 .js file) and mark it as the Entrypoint. MaxoPerf auto-detects the engine.

  3. Open Settings → Secrets and add LLM_API_KEY with your inference API key. MaxoPerf injects it as ${LLM_API_KEY} at run time, so it never appears in the uploaded file.

  4. Open the Configuration tab. Set Virtual users to 20, Ramp-up to 2m, and Duration to 8m (2 m ramp + 5 m hold + 1 m ramp-down). Pick the runner location closest to the inference server so network time adds as little as possible to the measurements.

  5. In the Failure criteria section, add: p95 latency > 3000 ms → fail and Error rate > 2% → fail.

  6. Click Run. The run detail page opens automatically.

  • Throughput (RPS). For an LLM endpoint with max_tokens: 256, expect 1–5 RPS per VU, depending on the model and hardware. If throughput levels off below the RPS you expect, the model is probably GPU-saturated.
  • p95 / p99 latency. LLM latency is orders of magnitude higher than a REST API. A p95 of 2–5 seconds is normal for a 7B–13B parameter model at moderate concurrency. Watch for a latency curve that keeps climbing and never levels off. It means the queue grows faster than the model can clear it.
  • Error rate. 429s appear when the endpoint’s rate-limit quota is exceeded. 503/504s appear when the inference server is overloaded or out of memory.
  • A cluster of 429 errors early in the run points to a quota problem, not a performance problem. Contact your LLM API provider or adjust your rate limits.
  • 504 timeouts under load mean the inference server is saturating. Reduce concurrency or add GPU capacity.
  • 500 errors with out of memory in the body mean the KV cache is exhausted. Reduce max_tokens or add GPU nodes.
DoDon’t
Fix max_tokens so response lengths (and timings) are comparable across runsLeave max_tokens open-ended; varying output length makes p95 meaningless
Pass Authorization through MaxoPerf secrets and never hard-code API keysPut API keys in uploaded test files
Start with 5–10 VUs and increase; LLM endpoints saturate fastJump to 100 VUs before you know the saturation point
Label requests clearly (label: chat-completions) for a per-endpoint breakdownUse unlabelled requests. The results panel then shows unnamed instead of a useful name
Set an explicit HTTP timeout (30–60 s) in k6Rely on the default 10 s timeout; inference often takes longer
Run against staging / non-production inference endpointsLoad-test production LLM APIs without provider approval and a rate-limit increase