Skip to content

AI API scalability — stress, spike, and soak

LLM and inference APIs fail in different ways from ordinary REST services. GPU memory is finite, worker pools are small, and autoscaling can lag by minutes, not seconds. A stress test exposes saturation. A spike test shows cold-start and autoscaling behaviour. A soak test catches KV-cache memory leaks and slow context-accumulation bugs. This page maps each test type to the AI failure mode it finds.

  • Read LLM performance and load testing and run a baseline load test first, so you know your steady-state p95 latency before you stress the endpoint.
  • Read Stress test, Spike test, and Soak / endurance test for how each test type works in general.
  • Run every stress, spike and soak test against staging, never production. An inference server under stress can crash out of memory and take other tenants down with it.
Failure modeSymptomsTest type that reveals it
GPU saturationp95 latency climbs linearly with VUs; throughput plateausStress
KV-cache exhaustionSudden spike in 500 errors with out of memory bodyStress (high max_tokens)
Worker pool saturationQueue depth grows; latency climbs but RPS stays flatStress
Cold-start lagFirst requests in a spike take 5–30 s while GPU worker initialisesSpike
Autoscaling lagErrors spike during the scale-out window before new instances are readySpike
Rate-limit cascade429 errors cluster at the start of a spikeSpike
Memory leak (KV cache)p95 latency drifts upward over hours; TPOT grows slowlySoak
Connection pool exhaustion503 errors appear only after several hours of sustained loadSoak

Stress test: finding the GPU saturation knee

Section titled “Stress test: finding the GPU saturation knee”

A stress test ramps VUs past the baseline load at a steady rate. You are looking for the point where latency bends upward and errors start.

execution:
- concurrency: 80 # target: 4× baseline load (baseline was 20 VUs)
ramp-up: 15m # slow ramp — 1 VU every ~11 s to observe each step
hold-for: 5m # brief hold at peak
scenario: llm-stress
scenarios:
llm-stress:
requests:
- label: chat-completions
url: https://inference.example.com/v1/chat/completions
method: POST
headers:
Content-Type: application/json
Authorization: "Bearer ${LLM_API_KEY}"
body: >
{
"model": "llama-3-8b-instruct",
"messages": [{"role": "user", "content": "Write a haiku about performance testing."}],
"max_tokens": 64,
"stream": false
}
  1. Duplicate your baseline load test. Rename it llm-stress-80vu. Update the Taurus YAML with concurrency: 80 and ramp-up: 15m.

  2. In the Configuration tab set Virtual users to 80, Ramp-up to 15m, Duration to 20m.

  3. Add a failure criterion: Error rate > 5% → fail. The run then records the breaking point for you.

  4. Click Run and watch the Overview tab live. The latency chart should bend at the point where p95 starts to climb and keeps climbing.

  5. When the run ends, note the VU count at that bend. That is your GPU saturation point for this model and hardware.

  • Latency inflection. The VU count where p95 goes from flat to climbing. This is the saturation knee: the most VUs you can hold before latency degrades.
  • Error rate step-change. 429s appear first (rate-limit quota), then 503/504s (the inference server is overloaded). A cluster of 503s at one VU count marks the worker pool ceiling.
  • Throughput plateau. RPS stops climbing while VUs keep rising. This is the server’s maximum token-generation throughput. Past this point, more VUs only lengthen the queue.

Spike test: cold-start and autoscaling lag

Section titled “Spike test: cold-start and autoscaling lag”

An inference server that autoscales on GPU demand can take 3–10 minutes to provision and warm a new GPU instance. A spike test with a near-instant VU surge shows you that cold-start window.

import http from 'k6/http';
import { check } from 'k6';
import { Rate } from 'k6/metrics';
const coldStartErrors = new Rate('cold_start_error_rate');
export const options = {
stages: [
{ duration: '2m', target: 5 }, // low baseline (warm state)
{ duration: '30s', target: 60 }, // near-instant surge (spike)
{ duration: '3m', target: 60 }, // hold at spike — autoscaling window
{ duration: '30s', target: 5 }, // drop back
{ duration: '2m', target: 5 }, // recovery observation
],
thresholds: {
http_req_duration: ['p(95)<10000'],
cold_start_error_rate: ['rate<0.10'], // tolerate up to 10% errors during cold-start window
},
};
const ENDPOINT = 'https://inference.example.com/v1/chat/completions';
const API_KEY = __ENV.LLM_API_KEY;
const PAYLOAD = JSON.stringify({
model: 'llama-3-8b-instruct',
messages: [{ role: 'user', content: 'What is 2 + 2?' }],
max_tokens: 32,
stream: false,
});
export default function () {
const res = http.post(ENDPOINT, PAYLOAD, {
headers: {
'Content-Type': 'application/json',
'Authorization': `Bearer ${API_KEY}`,
},
timeout: '30s',
});
const failed = res.status !== 200;
coldStartErrors.add(failed);
check(res, { 'status 200': (r) => r.status === 200 });
}

The cold_start_error_rate metric separates the short burst of cold-start errors from errors that continue after autoscaling should have finished. If cold_start_error_rate stays high for the whole 3-minute hold phase, autoscaling either never triggered or has not finished provisioning.

Soak test: KV-cache memory leak and connection exhaustion

Section titled “Soak test: KV-cache memory leak and connection exhaustion”

A soak test holds steady load for 4–8 hours to find resource leaks that grow slowly.

What AI soak tests reveal that short tests miss

Section titled “What AI soak tests reveal that short tests miss”
  • KV-cache accumulation. Some inference servers do not evict KV caches correctly, so memory grows with each request. p95 TTFT drifts up over hours as free GPU memory shrinks.
  • TPOT drift. Tokens/sec falls slowly as the model weights and the growing cached state compete for GPU VRAM bandwidth.
  • Connection pool exhaustion. HTTP connection pools to the vector DB or embedding service are not recycled correctly and run out after 4–6 hours.
execution:
- concurrency: 20 # same as baseline load test
ramp-up: 3m
hold-for: 6h # 6-hour soak
scenario: llm-soak
scenarios:
llm-soak:
requests:
- label: chat-completions
url: https://inference.example.com/v1/chat/completions
method: POST
headers:
Content-Type: application/json
Authorization: "Bearer ${LLM_API_KEY}"
body: >
{
"model": "llama-3-8b-instruct",
"messages": [{"role": "user", "content": "Describe a load testing best practice."}],
"max_tokens": 256,
"stream": false
}

Look for drift, not peaks:

  • p95 latency trend. The target is flat for 6 hours. Any upward slope, even 100 ms per hour, points to a resource leak.
  • Throughput decline. If RPS drops while VUs stay constant, the server is slowing down. Compare it with the inference server’s heap and VRAM metrics.
  • Error rate timing. Errors that start only after 3–4 hours, and not at run start, usually mean connection pool or file-descriptor exhaustion.
DoDon’t
Run a baseline load test before any stress / spike / soakStress an endpoint before you know its steady-state behaviour
Use short max_tokens for stress tests to see more VU steps before saturationUse large max_tokens for stress; you saturate at too few VUs for a useful curve
Observe the recovery window in spike tests (2 min minimum after VU drop)End the spike run at the VU peak, so you never see whether the system recovers
Run soak tests on staging with monitoring enabled for VRAM and heapRun AI soak tests on production shared with real users
Note the exact VU count at the saturation knee for capacity planningReport only the peak VU count; the knee is the number you act on