AI load test failure criteria and run comparison
A MaxoPerf run that ends with status Finished has completed. That does not mean it passed. Without criteria, you have to open every chart by hand to decide whether an LLM or inference endpoint met its SLOs. Failure criteria write your AI SLOs into the test itself. The run then gets an automatic Passed / Failed verdict that CI pipelines, scheduled nightly runs and model-upgrade comparisons can read without a person checking.
Before you start
Section titled “Before you start”- Read Cookbook: Failure criteria pass/fail gates. This page adds AI thresholds to that general recipe.
- You have at least one completed baseline run for the LLM endpoint (from LLM performance and load testing).
- You know your SLO targets: p95 TTFT budget, acceptable error rate, minimum throughput floor.
AI-specific failure criteria
Section titled “AI-specific failure criteria”Standard load tests gate on p95 latency and error rate. AI endpoints need a few more thresholds:
| Criterion | Metric | Typical threshold | Why |
|---|---|---|---|
| TTFT budget | http_req_duration p95 (or llm_ttft_ms custom) | < 800 ms–3 000 ms | Governs UX for chat/copilot features |
| End-to-end latency | http_req_duration p99 | < 10 000 ms | Tail latency, which the slowest users feel |
| Error rate | Overall error rate | < 2–5 % | Includes 429 rate-limit and 503/504 overload |
| 429 rate-limit rate | Custom rate_limit_rate | < 1 % | Flags quota exhaustion independently of server errors |
| Throughput floor | Throughput (RPS) | > N RPS | Catches runs where the test itself under-loaded (startup errors) |
| Custom token metric | llm_ttft_ms p95 (from streaming k6 script) | < 800 ms | TTFT when measured precisely via SSE parsing |
Step by step: configuring AI failure criteria
Section titled “Step by step: configuring AI failure criteria”-
Open the test in the MaxoPerf console and go to the Configuration tab.
-
Scroll to Failure criteria and click Add criterion.
-
Add the TTFT / latency criterion:
- Metric: p95 latency
- Label: All labels (or
chat-completionsif you labelled your request) - Operator: greater than
- Value:
3000(milliseconds; adjust to your model’s expected p95)
-
Add the error-rate criterion:
- Metric: Error rate
- Operator: greater than
- Value:
2(percent)
-
Add the throughput floor:
- Metric: Throughput (RPS)
- Operator: less than
- Value:
3(the minimum RPS you expect at your VU count; this stops a misconfigured test from passing)
-
Click Save.
Adding custom metric thresholds
Section titled “Adding custom metric thresholds”If your k6 script emits a custom metric (e.g. llm_ttft_ms from the Streaming and token latency testing script), MaxoPerf lists it in the custom metrics section:
- In the Failure criteria section, click Add criterion.
- Select metric: Custom metric.
- Enter the metric name:
llm_ttft_ms. - Aggregation: p95.
- Operator: greater than.
- Value:
800. - Click Add.
This criterion fails the run if TTFT p95 goes above 800 ms. That is a direct latency SLO for streaming UIs.
Run comparison: model and endpoint versions
Section titled “Run comparison: model and endpoint versions”When you upgrade a model, change a prompt template or switch inference providers, compare the new endpoint with the baseline run. The comparison puts numbers on every regression and improvement.
How to trigger a comparable run
Section titled “How to trigger a comparable run”For a valid comparison, both runs must use:
- The same VU count, ramp-up, and hold duration.
- The same prompt body and
max_tokensvalue. - The same MaxoPerf runner location.
Create a new test named llm-chat-v2 (or duplicate the baseline test and update the endpoint URL / model name). Run it. Then compare.
Using run comparison in MaxoPerf
Section titled “Using run comparison in MaxoPerf”- Open the new run (the candidate).
- Click Compare in the run detail header. A run-selector dialog appears.
- Select the baseline run from the same test or a different test with the same scenario.
- The Overview tab switches to comparison mode. Latency and throughput charts show one line per run, and Δ chips show the difference between runs.
Key delta signals for AI endpoint comparisons:
| Δ metric | Interpretation |
|---|---|
| p95 latency Δ positive | Candidate is slower. Expected if the new model is larger |
| p95 latency Δ negative | Candidate is faster. Check that output quality has not dropped |
| Throughput Δ positive | Candidate completes more requests/sec, so inference is more efficient |
| Error rate Δ positive | Candidate has more errors, from a quota issue or instability |
| p99 / p95 ratio widens | Candidate has more tail variance: more scheduling jitter or longer outlier generations |
Reporting dashboard for AI runs
Section titled “Reporting dashboard for AI runs”A MaxoPerf custom dashboard puts the AI endpoint metrics in one view: TTFT percentiles, error rate breakdown and throughput. You stop switching between tabs.
To build an AI-focused dashboard:
- Open a completed run and click Dashboard → New dashboard.
- Add a latency percentile panel with p50, p95, p99 lines. Rename it “TTFT / end-to-end latency”.
- Add a throughput panel and rename it “Inference RPS”.
- Add an error rate panel segmented by status code (429 vs 5xx).
- If your script emits
llm_output_tokens, add a custom metric panel and set it tollm_output_tokens rate. This is your tokens/sec proxy. - Save the dashboard as “AI endpoint standard”. MaxoPerf reuses it on every later run of this test.
CI/CD integration
Section titled “CI/CD integration”Failure criteria let you run AI load tests in CI. When a run ends with status Failed, the MaxoPerf API returns the failed status and the criterion that was violated:
# Poll for run completion and surface the verdictcurl -s -H "Authorization: Bearer $MAXOPERF_API_KEY" \ "https://app.maxoperf.com/api/v1/runs/$RUN_ID" \ | jq '{status: .status, violations: .failure_criteria_violations}'A CI job that triggers a run and polls this endpoint can exit non-zero when status == "failed", which blocks the model upgrade deployment without manual steps. See Run a test from CI/CD for the full pipeline guide.
Do / don’t
Section titled “Do / don’t”| Do | Don’t |
|---|---|
| Start criteria loose; tighten over several runs as you learn the normal range | Set criteria tighter than your measured baseline; every run will fail |
| Use a throughput floor to catch misconfigured tests | Rely on latency criteria alone. A test that sent 1 RPS can pass on latency and prove nothing |
| Compare runs with identical load parameters for a fair model A/B | Compare runs with different VU counts; latency does not scale linearly with concurrency |
| Document the SLO thresholds in the test description so reviewers understand the pass bar | Leave criteria values undocumented. Later reviewers won’t know if 3000 ms was chosen on purpose |
| Use run comparison after every model upgrade, prompt change, or infrastructure change | Review comparisons only after a failure. Comparing every time catches slow regressions |
Where to go next
Section titled “Where to go next”- Cookbook: Failure criteria pass/fail gates: the general failure-criteria recipe.
- LLM performance and load testing: run the baseline this page compares against.
- AI API scalability: stress, spike, and soak: apply failure criteria to stress and soak runs to detect saturation automatically.
- Streaming and token latency testing: emit
llm_ttft_msfor precise TTFT criteria.