AI API scalability — stress, spike, and soak
LLM and inference APIs fail in different ways from ordinary REST services. GPU memory is finite, worker pools are small, and autoscaling can lag by minutes, not seconds. A stress test exposes saturation. A spike test shows cold-start and autoscaling behaviour. A soak test catches KV-cache memory leaks and slow context-accumulation bugs. This page maps each test type to the AI failure mode it finds.
Before you start
Section titled “Before you start”- Read LLM performance and load testing and run a baseline load test first, so you know your steady-state p95 latency before you stress the endpoint.
- Read Stress test, Spike test, and Soak / endurance test for how each test type works in general.
- Run every stress, spike and soak test against staging, never production. An inference server under stress can crash out of memory and take other tenants down with it.
AI-specific failure modes
Section titled “AI-specific failure modes”| Failure mode | Symptoms | Test type that reveals it |
|---|---|---|
| GPU saturation | p95 latency climbs linearly with VUs; throughput plateaus | Stress |
| KV-cache exhaustion | Sudden spike in 500 errors with out of memory body | Stress (high max_tokens) |
| Worker pool saturation | Queue depth grows; latency climbs but RPS stays flat | Stress |
| Cold-start lag | First requests in a spike take 5–30 s while GPU worker initialises | Spike |
| Autoscaling lag | Errors spike during the scale-out window before new instances are ready | Spike |
| Rate-limit cascade | 429 errors cluster at the start of a spike | Spike |
| Memory leak (KV cache) | p95 latency drifts upward over hours; TPOT grows slowly | Soak |
| Connection pool exhaustion | 503 errors appear only after several hours of sustained load | Soak |
Stress test: finding the GPU saturation knee
Section titled “Stress test: finding the GPU saturation knee”A stress test ramps VUs past the baseline load at a steady rate. You are looking for the point where latency bends upward and errors start.
Taurus YAML: stress profile
Section titled “Taurus YAML: stress profile”execution: - concurrency: 80 # target: 4× baseline load (baseline was 20 VUs) ramp-up: 15m # slow ramp — 1 VU every ~11 s to observe each step hold-for: 5m # brief hold at peak scenario: llm-stress
scenarios: llm-stress: requests: - label: chat-completions url: https://inference.example.com/v1/chat/completions method: POST headers: Content-Type: application/json Authorization: "Bearer ${LLM_API_KEY}" body: > { "model": "llama-3-8b-instruct", "messages": [{"role": "user", "content": "Write a haiku about performance testing."}], "max_tokens": 64, "stream": false }Console walk-through
Section titled “Console walk-through”-
Duplicate your baseline load test. Rename it
llm-stress-80vu. Update the Taurus YAML withconcurrency: 80andramp-up: 15m. -
In the Configuration tab set Virtual users to
80, Ramp-up to15m, Duration to20m. -
Add a failure criterion: Error rate > 5% → fail. The run then records the breaking point for you.
-
Click Run and watch the Overview tab live. The latency chart should bend at the point where p95 starts to climb and keeps climbing.
-
When the run ends, note the VU count at that bend. That is your GPU saturation point for this model and hardware.
Reading the stress result
Section titled “Reading the stress result”- Latency inflection. The VU count where p95 goes from flat to climbing. This is the saturation knee: the most VUs you can hold before latency degrades.
- Error rate step-change. 429s appear first (rate-limit quota), then 503/504s (the inference server is overloaded). A cluster of 503s at one VU count marks the worker pool ceiling.
- Throughput plateau. RPS stops climbing while VUs keep rising. This is the server’s maximum token-generation throughput. Past this point, more VUs only lengthen the queue.
Spike test: cold-start and autoscaling lag
Section titled “Spike test: cold-start and autoscaling lag”An inference server that autoscales on GPU demand can take 3–10 minutes to provision and warm a new GPU instance. A spike test with a near-instant VU surge shows you that cold-start window.
k6 spike script
Section titled “k6 spike script”import http from 'k6/http';import { check } from 'k6';import { Rate } from 'k6/metrics';
const coldStartErrors = new Rate('cold_start_error_rate');
export const options = { stages: [ { duration: '2m', target: 5 }, // low baseline (warm state) { duration: '30s', target: 60 }, // near-instant surge (spike) { duration: '3m', target: 60 }, // hold at spike — autoscaling window { duration: '30s', target: 5 }, // drop back { duration: '2m', target: 5 }, // recovery observation ], thresholds: { http_req_duration: ['p(95)<10000'], cold_start_error_rate: ['rate<0.10'], // tolerate up to 10% errors during cold-start window },};
const ENDPOINT = 'https://inference.example.com/v1/chat/completions';const API_KEY = __ENV.LLM_API_KEY;
const PAYLOAD = JSON.stringify({ model: 'llama-3-8b-instruct', messages: [{ role: 'user', content: 'What is 2 + 2?' }], max_tokens: 32, stream: false,});
export default function () { const res = http.post(ENDPOINT, PAYLOAD, { headers: { 'Content-Type': 'application/json', 'Authorization': `Bearer ${API_KEY}`, }, timeout: '30s', });
const failed = res.status !== 200; coldStartErrors.add(failed);
check(res, { 'status 200': (r) => r.status === 200 });}The cold_start_error_rate metric separates the short burst of cold-start errors from errors that continue after autoscaling should have finished. If cold_start_error_rate stays high for the whole 3-minute hold phase, autoscaling either never triggered or has not finished provisioning.
Soak test: KV-cache memory leak and connection exhaustion
Section titled “Soak test: KV-cache memory leak and connection exhaustion”A soak test holds steady load for 4–8 hours to find resource leaks that grow slowly.
What AI soak tests reveal that short tests miss
Section titled “What AI soak tests reveal that short tests miss”- KV-cache accumulation. Some inference servers do not evict KV caches correctly, so memory grows with each request. p95 TTFT drifts up over hours as free GPU memory shrinks.
- TPOT drift. Tokens/sec falls slowly as the model weights and the growing cached state compete for GPU VRAM bandwidth.
- Connection pool exhaustion. HTTP connection pools to the vector DB or embedding service are not recycled correctly and run out after 4–6 hours.
Taurus YAML: soak profile
Section titled “Taurus YAML: soak profile”execution: - concurrency: 20 # same as baseline load test ramp-up: 3m hold-for: 6h # 6-hour soak scenario: llm-soak
scenarios: llm-soak: requests: - label: chat-completions url: https://inference.example.com/v1/chat/completions method: POST headers: Content-Type: application/json Authorization: "Bearer ${LLM_API_KEY}" body: > { "model": "llama-3-8b-instruct", "messages": [{"role": "user", "content": "Describe a load testing best practice."}], "max_tokens": 256, "stream": false }Reading the AI soak result
Section titled “Reading the AI soak result”Look for drift, not peaks:
- p95 latency trend. The target is flat for 6 hours. Any upward slope, even 100 ms per hour, points to a resource leak.
- Throughput decline. If RPS drops while VUs stay constant, the server is slowing down. Compare it with the inference server’s heap and VRAM metrics.
- Error rate timing. Errors that start only after 3–4 hours, and not at run start, usually mean connection pool or file-descriptor exhaustion.
Do / don’t
Section titled “Do / don’t”| Do | Don’t |
|---|---|
| Run a baseline load test before any stress / spike / soak | Stress an endpoint before you know its steady-state behaviour |
Use short max_tokens for stress tests to see more VU steps before saturation | Use large max_tokens for stress; you saturate at too few VUs for a useful curve |
| Observe the recovery window in spike tests (2 min minimum after VU drop) | End the spike run at the VU peak, so you never see whether the system recovers |
| Run soak tests on staging with monitoring enabled for VRAM and heap | Run AI soak tests on production shared with real users |
| Note the exact VU count at the saturation knee for capacity planning | Report only the peak VU count; the knee is the number you act on |
Where to go next
Section titled “Where to go next”- Stress test: stress testing in general, with a console walk-through.
- Spike test: recovery window, autoscaling and circuit breakers.
- Soak / endurance test: drift, scheduling and how to read long runs.
- AI load test failure criteria: automate the pass/fail verdict for AI stress runs.
- LLM performance and load testing: run the baseline first.
- Streaming and token latency testing: measure TTFT drift during the soak hold phase.