AI performance testing glossary
This glossary defines the performance and load-testing terms used in the AI Testing section. Each entry has a definition and a note on where the term shows up in MaxoPerf workflows or results. For short one-card definitions, see the SEO Glossary.
Term index
Section titled “Term index”B · C · Co · Co · Co · E · G · I · K · L · P · Q · R · S · T · Ti · Ti · To · To
Time to first token (TTFT)
Section titled “Time to first token (TTFT)”The time from sending an inference request to receiving the first token of the response. For streaming LLM APIs, this is the latency to the first data: chunk in the SSE stream. The prefill phase sets TTFT: the model processes all input tokens before it generates any output.
In MaxoPerf: TTFT is the latency to the first SSE event in a streaming response. In k6, emit TTFT as a custom Trend metric (llm_ttft_ms) by recording Date.now() at the first parsed data: line. Add a separate failure criterion on llm_ttft_ms p95, for example > 800 ms → fail for a chat UI.
See Streaming and token latency testing.
TPOT (time per output token)
Section titled “TPOT (time per output token)”The average time to generate each output token after the first. The decode phase sets TPOT: the model generates one token at a time, each from the context it has generated so far. TPOT sets the streaming “reading pace”. A high TPOT makes a chat UI visibly stutter.
In MaxoPerf: TPOT = (total generation time − TTFT) / (output tokens − 1). HTTP metrics cannot give you TPOT directly. Approximate it by recording the time from the first to the last SSE chunk and dividing by the output token count.
Inter-token latency (ITL)
Section titled “Inter-token latency (ITL)”The time between one token and the next in a streaming response. ITL and TPOT measure the same thing from two sides. ITL is the gap between each pair of tokens, observed at the client. TPOT is the average time per output token, computed afterwards. Inference literature uses the two terms interchangeably.
Time to last token (TTLT)
Section titled “Time to last token (TTLT)”The time from sending the request to receiving the final token of the response, which is the full streaming latency. For non-streaming calls, TTLT equals end-to-end latency. TTLT = TTFT + (output_tokens × TPOT).
In MaxoPerf: http_req_duration measures TTLT (last byte received). For streaming endpoints, this is the full generation time, not the user-visible response start.
End-to-end latency
Section titled “End-to-end latency”For non-streaming LLM calls, the total time from sending the request to receiving the full response body. This equals TTLT. For streaming calls, end-to-end latency is also TTLT, but the user sees the response start at TTFT.
In MaxoPerf: The standard http_req_duration metric. Set separate SLO thresholds for TTFT and end-to-end latency, because they answer different questions.
Tokens per second (tokens/sec)
Section titled “Tokens per second (tokens/sec)”The rate at which the inference system generates output tokens. You can measure it per request (generation speed) or across all concurrent requests (system throughput). Higher tokens/sec means the GPU is used more efficiently.
In MaxoPerf: Estimate aggregate tokens/sec by emitting a llm_output_tokens counter in your k6 script and dividing the total by the test duration in seconds. If tokens/sec falls as concurrency rises, the inference backend is saturating.
Throughput (RPS)
Section titled “Throughput (RPS)”Requests completed per second across all virtual users. For LLM endpoints with fixed max_tokens, RPS is the main capacity metric. It tells you how many inferences the system can sustain per unit of time.
In MaxoPerf: The Overview tab shows throughput in RPS over time. In a load test, throughput should level off at a sustainable value during the hold phase. If throughput levels off while latency climbs, the system has hit its processing ceiling.
Latency percentiles (p50, p95, p99)
Section titled “Latency percentiles (p50, p95, p99)”Points on the latency distribution of all requests in a run. p95 means 95% of requests complete within that time, and p99 means 99%. Mean latency misleads for LLM endpoints. Output lengths vary, so the distribution has a heavy tail, and a few long outliers pull the mean up.
In MaxoPerf: Use p95 (not the mean) as your primary SLO metric in failure criteria. p99 pulls away from p95 when some requests have unusually long generations. A high p99/p95 ratio points to a wide spread of output lengths or occasional GPU scheduling jitter.
Concurrency / virtual users
Section titled “Concurrency / virtual users”The number of requests in flight at the same time. In MaxoPerf (and k6/Taurus), this is the VU count. For LLM endpoints, each concurrent VU holds one open GPU context during generation. At high concurrency, GPU memory and worker pool limits run out before network bandwidth does.
In MaxoPerf: Start low (5–10 VUs) and add VUs in steps. The VU count where p95 latency starts to climb and keeps climbing is the saturation knee: the highest concurrency you can hold without degradation. See AI API scalability: stress, spike, and soak.
Queueing and saturation
Section titled “Queueing and saturation”When requests arrive faster than the inference server can process them, they queue. Queuing shows up as rising p95 latency while throughput stays flat or falls. The system accepts requests but does not complete them any faster. Queueing is the main sign that a system has passed its saturation point.
In MaxoPerf: Look for VU count climbing while RPS stays flat. That flat line is the saturation ceiling. When latency climbs and RPS does not, each extra VU adds to queue depth, not to throughput.
GPU saturation
Section titled “GPU saturation”The state where the GPU’s memory bandwidth, VRAM or compute is fully used, and the inference server cannot process more requests any faster. GPU saturation has a typical pattern: throughput levels off, p95 latency starts climbing linearly, and the error rate rises once KV-cache or worker pool limits are exceeded.
In MaxoPerf: Find the saturation knee in stress-test results. It is the VU count where the latency chart goes from flat to climbing, and it is the highest concurrency this model and hardware configuration supports. See AI API scalability: stress, spike, and soak.
KV cache
Section titled “KV cache”The key-value cache that an inference server keeps for each in-flight request. It stores intermediate attention results for the input tokens. The KV cache uses more GPU VRAM than anything else during inference. Each active request holds a KV cache sized by its context length. When VRAM runs out, new requests are rejected or degraded.
In MaxoPerf: KV-cache exhaustion appears as sudden 500 errors with out of memory bodies. Reduce max_tokens or the number of concurrent VUs. In soak tests, a slow KV-cache memory leak appears as upward drift in p95 TTFT over several hours.
Cold start / warm start
Section titled “Cold start / warm start”A cold start happens when an inference server must load model weights and start GPU workers before it can handle the first request. That adds 10–60 seconds of latency to the first requests after startup or after a scale-out event. A warm start is when GPU workers and model weights are already loaded in VRAM.
In MaxoPerf: Cold-start lag shows up as high TTFT in the first 30–60 seconds of a test, or in the first window of a spike test. Use a ramp-up phase so workers warm up before you read steady-state metrics. Test the cold-start path on purpose with a spike test to measure autoscaling lag.
Rate limit / HTTP 429
Section titled “Rate limit / HTTP 429”A rate-limit response (HTTP 429) from a hosted LLM API. It means the client went over its token-per-minute (TPM) or request-per-minute (RPM) quota. A 429 is quota enforcement, not an inference-server error. The fix differs from a 503/504 server error: raise the quota or shape the requests.
In MaxoPerf: Track the 429 rate as its own custom metric or failure criterion, separate from the overall error rate. A cluster of 429s early in the ramp-up means the quota is too small for the target concurrency. See AI load test failure criteria.
Prompt tokens / output tokens
Section titled “Prompt tokens / output tokens”The two billing buckets for hosted LLM APIs. Prompt tokens are the input tokens (system message + user message + any injected context). Output tokens are the model-generated response tokens. Most providers price output tokens 2–10× higher than input tokens.
In MaxoPerf: Extract both from the usage field in the response body and emit them as k6 Counter metrics. Use the totals to compute per-request cost and project monthly spend from the load profile. See Inference cost and token budget testing.
Cost per token / cost per request
Section titled “Cost per token / cost per request”The inference cost of a single token or a single complete request. For hosted APIs: cost = (prompt_tokens / 1000) × input_price + (output_tokens / 1000) × output_price. For self-hosted models: cost is compute time × hourly server cost / requests completed.
In MaxoPerf: Emit a k6 Trend metric (llm_cost_usd_per_req) computed from the usage field. The p50 and p95 of this metric, combined with RPS, give you the cost model for your load profile. See Inference cost and token budget testing.
Context window (as a latency factor)
Section titled “Context window (as a latency factor)”The maximum number of tokens an LLM can process in a single inference call, input and output combined. For load testing, requests that fill the context window (long system prompts + many retrieved RAG documents + long output) are the worst case for TTFT, TPOT and cost. The KV cache for a 128 K-token context is 64× larger than for a 2 K-token context.
In MaxoPerf: Test at the prompt lengths your application uses. A short-prompt test that shows p95 = 800 ms says nothing about what happens when your RAG system injects 4 000 tokens of retrieved context. Run separate load tests at realistic context lengths.
Server-sent events (SSE) / chunked transfer
Section titled “Server-sent events (SSE) / chunked transfer”Server-sent events (SSE) is the HTTP streaming mechanism most LLM APIs use to send tokens as they are generated. The server sends Content-Type: text/event-stream and emits data: {...} lines as tokens are generated. Chunked transfer encoding is the underlying HTTP/1.1 mechanism that delivers body bytes without a predetermined Content-Length.
In MaxoPerf: k6 buffers the full SSE response before the script function returns. Parse SSE chunks from res.body.split('\n') and record timestamps to extract TTFT. See Streaming and token latency testing.
Batching
Section titled “Batching”Grouping several inference requests into one GPU pass to raise throughput. Continuous batching (dynamic batching) is how modern inference servers (vLLM, TGI) pack requests of different lengths into the same GPU kernel. They start new requests without waiting for the current ones to finish.
In practice: Batching raises aggregate tokens/sec but adds to per-request TTFT, because a request may wait for a batch slot. In a load test you see this as a TTFT floor that rises a little as concurrency grows: each request waits for a batch slot before its prefill starts.
Streaming throughput
Section titled “Streaming throughput”The total rate at which the inference system delivers output tokens to all concurrent clients, in tokens/sec across all VUs. This GPU-level metric tells you whether the system keeps up with concurrent streaming demand. If streaming throughput falls below (VUs × expected_tokens/sec_per_user), clients see stutter.
In MaxoPerf: Compute streaming throughput as total_output_tokens / test_duration_seconds. If it falls while the VU count stays constant (for example during a soak test), VRAM or bandwidth is degrading.
Where to go next
Section titled “Where to go next”- AI testing do and don’t: do/don’t pairs that use these terms
- LLM performance and load testing: run your first LLM load test in MaxoPerf
- Streaming and token latency testing: measure TTFT and TPOT in practice
- Inference cost and token budget testing: cost metrics in depth
- Academy Glossary: general performance and load testing terms