Failure criteria pass/fail gates
Problem: Without explicit thresholds, a MaxoPerf run always ends with status Finished. You have
to read the charts and decide yourself whether it passed. Automated tests (scheduled or
CI-triggered) need an objective, machine-readable signal. Failure criteria put your SLOs
into the test, so the run produces its own verdict.
Test type: Any. Failure criteria apply to all test types.
Prerequisites
Section titled “Prerequisites”- A MaxoPerf account and a test with at least one baseline run.
- Your SLO thresholds: p95 latency target, error rate budget, and optionally throughput floor.
How failure criteria work
Section titled “How failure criteria work”MaxoPerf evaluates failure criteria against the entire run (not per-second). When any criterion
is violated, the run status changes from Finished to Failed. Multiple criteria are combined with
AND: all must pass for the run to count as Finished.
The criteria you save on the test are the single source of truth, for every test type and engine.
MaxoPerf judges them in up to two places: inside the load engine while the run is live (so a stop
criterion can end the run early), and again over the complete run once it finishes. Most metrics
are judged in both places; metrics that only make sense for the whole run, such as throughput, are
judged over the complete run only.
- Any breach fails the run. A criterion with severity Error that is breached fails the run
whatever its action. The action only decides whether the run stops early (
stop) or keeps generating load until its planned end (continue). A Warning criterion is reported but never fails the run. - The verdict waits for the last results. A short run can finish in the same second its runners upload their final results. MaxoPerf holds the verdict until every runner’s final results have arrived (up to two minutes), so a breach in the last seconds is never missed. If results are still missing after that, the verdict is computed on what arrived and the run log says so.
- Uploaded Taurus YAML. When you upload a Taurus YAML with a
passfailreporting module, its criteria are imported into the test at upload time. From then on the console criteria are what every run enforces: the YAML’s ownpassfailblock is replaced at run time, and any row that was not imported is listed in the run log rather than silently dropped. Edit criteria in the console, not in the YAML. Native k6thresholdsinside a k6 script are left untouched. - Counts on multi-runner runs. Inside the load engine, each runner judges only its own share of
the load; the evaluation over the complete run judges the run-wide total. For a count such as
Hit count, or Failed responses measured as a count rather than a percentage (
> 100), a run with several runners can therefore stay under the threshold on every runner and still breach it run-wide: the run fails at the end, but it is not stopped early. Set a count threshold for the whole run, not per runner.
You can set criteria at:
- Request label level:
POST /checkout p95 < 300ms - Global level:
overall error rate < 1 % - Throughput floor:
RPS > 50(catches runs where the test itself did not generate enough load)
Step by step in MaxoPerf
Section titled “Step by step in MaxoPerf”1. Open the failure criteria section
Section titled “1. Open the failure criteria section”- Open the test in the MaxoPerf console and switch to the Configuration tab.
- Scroll to the Failure criteria section.
- Click Add criterion.
2. Add a latency criterion
Section titled “2. Add a latency criterion”- Select metric: p95 latency.
- Label: leave All labels, or pick a request label from the list for per-endpoint precision
(see
ANYlabel). - Operator: greater than.
- Value: your SLO threshold in milliseconds, for example
500. - Click Add.
3. Add an error-rate criterion
Section titled “3. Add an error-rate criterion”- Click Add criterion again.
- Select metric: Error rate.
- Operator: greater than.
- Value:
1(meaning 1 %). - Click Add.
4. (Optional) Add a throughput floor
Section titled “4. (Optional) Add a throughput floor”A throughput floor guards against a false “pass” when the test runner did not generate meaningful load (e.g. if a startup error limited traffic to 1 RPS):
- Click Add criterion.
- Select metric: Throughput (RPS).
- Operator: less than.
- Value: your expected minimum, for example
80(80 % of target RPS). - Click Add.
5. Save and run
Section titled “5. Save and run”- Review the criteria list and click Save.
- Click Run now (or let a scheduled run fire).
- When the run finishes, the status badge reflects the verdict.
6. Inspect a failed run
Section titled “6. Inspect a failed run”- Open a run with status
Failed. - The run’s Criteria tab lists every criterion with its label and status, and shows which one was violated and by how much.
- Use the metric charts to find when in the run the violation happened.
ANY label
Section titled “ANY label”Every criterion has a label, and the label is a filter: the criterion is evaluated only over the samples whose request (transaction) label matches it exactly.
- All labels is the default. It is stored as the literal value
ANYand evaluates the whole run across every label. Use it for run-wide gates such as “error rate below 1 %”. - The Label field suggests the labels your test really produces: the labels of its latest finished run, and, for tests built in the console, the request and step names of the test definition. You can also type a label that has not run yet.
- Labels are case-sensitive and must match exactly.
Checkoutandcheckoutare different labels, and a descriptive name such as “Any failed request” matches nothing. Put descriptions in the criterion’s message instead. ANYis reserved: a request literally namedANYcannot be targeted by a criterion, becauseANYalways means all labels.- Subjects that describe the whole run, such as concurrency, cannot be scoped to a label; the console resets their label to All labels.
- When you save a test’s criteria with a label that no known sample produces, the test’s Configuration tab shows a warning under the criterion. The save still succeeds (a new test has no labels until its first run), but check the spelling. A VarioTest member’s own criteria are checked against the labels of the test that member runs, and the warning appears under the criterion in that member’s card.
Over the API and MCP, send "label": "ANY" for all labels; an empty or missing label is stored as
ANY. GET /v1/tests/{testId}/sample-labels (MCP: list_test_labels) returns the labels you can
target.
No data
Section titled “No data”A criterion whose label matched no samples in a run cannot be evaluated. Instead of silently passing, it shows a No data badge in the run’s Criteria tab (and on the public share page). Hover or focus the badge for the explanation, including the labels the run did produce. The run log also records a warning naming the criterion and its label.
A No data criterion does not change the verdict by itself. It almost always means a mistyped or renamed label: fix the label, or switch the criterion to All labels.
Verify
Section titled “Verify”- A run that meets all criteria ends with status
Finished(green badge). - A run that violates one or more criteria ends with status
Failed(red badge). - The failure criteria panel on a failed run names the violated criterion, its label, and the measured value vs the threshold.
- No criterion shows No data. If one does, its label does not match any request label (see No data).
- CI pipelines that poll the run status read
failedand exit non-zero. See CI-gated performance test.
Variations
Section titled “Variations”- Per-label criteria: apply tighter criteria to critical endpoints (
POST /payment p95 < 200ms) and looser criteria to background endpoints (GET /analytics p95 < 2000ms). - Criteria-only failure, no threshold on VUs: leave VUs/throughput criteria off for exploratory runs where the goal is pure observation.
- Failure without stop: set the criterion’s action to
continueto keep the run going to its planned end after a breach. The run is still marked failed. Usestopto end the run as soon as the criterion is breached.
Where to go next
Section titled “Where to go next”- CI-gated performance test: use criteria to drive a CI pass/fail signal.
- Scheduled regression test: pair criteria with a nightly cron for automated regression detection.
- Comparing runs and baselines: after a failure, compare the failed run to the last good run to find the regression.
- Baselines, SLOs, and error budgets: how to choose the right threshold values.