Skip to content

Failure criteria pass/fail gates

Problem: Without explicit thresholds, a MaxoPerf run always ends with status Finished. You have to read the charts and decide yourself whether it passed. Automated tests (scheduled or CI-triggered) need an objective, machine-readable signal. Failure criteria put your SLOs into the test, so the run produces its own verdict.

Test type: Any. Failure criteria apply to all test types.

  • A MaxoPerf account and a test with at least one baseline run.
  • Your SLO thresholds: p95 latency target, error rate budget, and optionally throughput floor.

MaxoPerf evaluates failure criteria against the entire run (not per-second). When any criterion is violated, the run status changes from Finished to Failed. Multiple criteria are combined with AND: all must pass for the run to count as Finished.

The criteria you save on the test are the single source of truth, for every test type and engine. MaxoPerf judges them in up to two places: inside the load engine while the run is live (so a stop criterion can end the run early), and again over the complete run once it finishes. Most metrics are judged in both places; metrics that only make sense for the whole run, such as throughput, are judged over the complete run only.

  • Any breach fails the run. A criterion with severity Error that is breached fails the run whatever its action. The action only decides whether the run stops early (stop) or keeps generating load until its planned end (continue). A Warning criterion is reported but never fails the run.
  • The verdict waits for the last results. A short run can finish in the same second its runners upload their final results. MaxoPerf holds the verdict until every runner’s final results have arrived (up to two minutes), so a breach in the last seconds is never missed. If results are still missing after that, the verdict is computed on what arrived and the run log says so.
  • Uploaded Taurus YAML. When you upload a Taurus YAML with a passfail reporting module, its criteria are imported into the test at upload time. From then on the console criteria are what every run enforces: the YAML’s own passfail block is replaced at run time, and any row that was not imported is listed in the run log rather than silently dropped. Edit criteria in the console, not in the YAML. Native k6 thresholds inside a k6 script are left untouched.
  • Counts on multi-runner runs. Inside the load engine, each runner judges only its own share of the load; the evaluation over the complete run judges the run-wide total. For a count such as Hit count, or Failed responses measured as a count rather than a percentage (> 100), a run with several runners can therefore stay under the threshold on every runner and still breach it run-wide: the run fails at the end, but it is not stopped early. Set a count threshold for the whole run, not per runner.

You can set criteria at:

  • Request label level: POST /checkout p95 < 300ms
  • Global level: overall error rate < 1 %
  • Throughput floor: RPS > 50 (catches runs where the test itself did not generate enough load)
  1. Open the test in the MaxoPerf console and switch to the Configuration tab.
  2. Scroll to the Failure criteria section.
  3. Click Add criterion.
  1. Select metric: p95 latency.
  2. Label: leave All labels, or pick a request label from the list for per-endpoint precision (see ANY label).
  3. Operator: greater than.
  4. Value: your SLO threshold in milliseconds, for example 500.
  5. Click Add.
  1. Click Add criterion again.
  2. Select metric: Error rate.
  3. Operator: greater than.
  4. Value: 1 (meaning 1 %).
  5. Click Add.

A throughput floor guards against a false “pass” when the test runner did not generate meaningful load (e.g. if a startup error limited traffic to 1 RPS):

  1. Click Add criterion.
  2. Select metric: Throughput (RPS).
  3. Operator: less than.
  4. Value: your expected minimum, for example 80 (80 % of target RPS).
  5. Click Add.
  1. Review the criteria list and click Save.
  2. Click Run now (or let a scheduled run fire).
  3. When the run finishes, the status badge reflects the verdict.
  1. Open a run with status Failed.
  2. The run’s Criteria tab lists every criterion with its label and status, and shows which one was violated and by how much.
  3. Use the metric charts to find when in the run the violation happened.

Every criterion has a label, and the label is a filter: the criterion is evaluated only over the samples whose request (transaction) label matches it exactly.

  • All labels is the default. It is stored as the literal value ANY and evaluates the whole run across every label. Use it for run-wide gates such as “error rate below 1 %”.
  • The Label field suggests the labels your test really produces: the labels of its latest finished run, and, for tests built in the console, the request and step names of the test definition. You can also type a label that has not run yet.
  • Labels are case-sensitive and must match exactly. Checkout and checkout are different labels, and a descriptive name such as “Any failed request” matches nothing. Put descriptions in the criterion’s message instead.
  • ANY is reserved: a request literally named ANY cannot be targeted by a criterion, because ANY always means all labels.
  • Subjects that describe the whole run, such as concurrency, cannot be scoped to a label; the console resets their label to All labels.
  • When you save a test’s criteria with a label that no known sample produces, the test’s Configuration tab shows a warning under the criterion. The save still succeeds (a new test has no labels until its first run), but check the spelling. A VarioTest member’s own criteria are checked against the labels of the test that member runs, and the warning appears under the criterion in that member’s card.

Over the API and MCP, send "label": "ANY" for all labels; an empty or missing label is stored as ANY. GET /v1/tests/{testId}/sample-labels (MCP: list_test_labels) returns the labels you can target.

A criterion whose label matched no samples in a run cannot be evaluated. Instead of silently passing, it shows a No data badge in the run’s Criteria tab (and on the public share page). Hover or focus the badge for the explanation, including the labels the run did produce. The run log also records a warning naming the criterion and its label.

A No data criterion does not change the verdict by itself. It almost always means a mistyped or renamed label: fix the label, or switch the criterion to All labels.

  • A run that meets all criteria ends with status Finished (green badge).
  • A run that violates one or more criteria ends with status Failed (red badge).
  • The failure criteria panel on a failed run names the violated criterion, its label, and the measured value vs the threshold.
  • No criterion shows No data. If one does, its label does not match any request label (see No data).
  • CI pipelines that poll the run status read failed and exit non-zero. See CI-gated performance test.
  • Per-label criteria: apply tighter criteria to critical endpoints (POST /payment p95 < 200ms) and looser criteria to background endpoints (GET /analytics p95 < 2000ms).
  • Criteria-only failure, no threshold on VUs: leave VUs/throughput criteria off for exploratory runs where the goal is pure observation.
  • Failure without stop: set the criterion’s action to continue to keep the run going to its planned end after a breach. The run is still marked failed. Use stop to end the run as soon as the criterion is breached.