Skip to content

Comparing runs and baselines

Problem: A single run result does not tell you whether performance improved or got worse. Compare the current run against a designated baseline. The delta metrics then answer the question you care about: is it better or worse than last time?

Test type: Baseline regression test. Also useful after any test type change.

  • At least two completed runs of the same test (or two runs of tests with the same request labels).
  • A designated baseline run. This is usually the last run that met your SLO before the current change.

The MaxoPerf comparison view places two runs side by side and computes delta chips for:

  • p50, p90, p95, p99 latency per label and aggregate.
  • Throughput (RPS), aggregate and per label.
  • Error rate, aggregate and per label.
  • VU count (to verify the test conditions were the same).

Delta chips are green for improvement and red for regression, relative to the baseline run.

  1. Go to Runs in the console and click the run you want to evaluate.
  2. The run detail page opens on the Overview tab.
  1. In the run header, click the Compare button (or Compare to baseline if a default baseline is set).
  2. In the selector that appears, choose the baseline run. Recent runs come first. Use the search box to find a run by ID or date.
  3. The page reloads in split view with the current run on the left and the baseline on the right.

Each metric card shows three values:

  • Current run value, for example 245 ms p95.
  • Baseline value, for example 257 ms p95.
  • Delta: −12 ms (−5 %) in green, or +42 ms (+16 %) in red.

Work through the labels:

  • Green across all labels: the change is an improvement or neutral. Safe to ship (subject to other checks).
  • Red on one label: investigate that endpoint. Check whether the VU count and the data were the same.
  • Red on aggregate but mixed per-label: one slow endpoint is pulling the average up. Drill in.

The comparison view URL is stable because it encodes both run IDs. Copy it from the browser address bar and paste it into a pull request review, Slack message, or incident ticket. Teammates click the link and see the same view.

  • The comparison page shows both run IDs in the header.
  • Delta chips appear on all key metrics. If a label is missing from one run (because a scenario changed), the comparison shows N/A for that label.
  • The summary section shows the VU count and duration for both runs. Confirm they match before you read the delta chips.
  • Set a default baseline: in the test Configuration tab, pin a run as the “official” baseline. Every new run then shows a delta against it on the overview tab, and you do not need to open the compare view.
  • Cross-test comparison: compare a run from a new Taurus scenario against the equivalent k6 run to check that the migrated scenario produces the same throughput.
  • Post-incident comparison: after a production incident, compare the last pre-incident run to the first post-incident run to see whether performance changed along with the event.
  • VarioTest child run comparison: open a VarioTest child run and compare it to the same scenario from a prior VarioTest to isolate per-scenario regressions.