Skip to content

How to use test data

Six jobs, in the order you’ll usually do them: bring in data, shape it, attach it to a test, and confirm how it lands on the runners.

How to use test data.
  1. Open Test data in the left navigation and click New Dataset.
  2. Choose Import CSV and select a file from your machine. To reuse a file you uploaded before, choose the Shared CSV tab instead and pick it there.
  3. Check the preview: MaxoPerf detects the delimiter, the encoding, and the header row (First row is header), and shows the first five rows.
  4. Keep or change the Dataset Name, then click Create Dataset & Open Workbench. Each header becomes a column.
  1. From New Dataset, choose Blank dataset and enter a Dataset Name.
  2. Keep the Row count of 1,000 or change it, then click Open blank workbench. You can change the count later with the workbench’s Rows setting.
  3. Click Add Column, choose Fake Data, and pick a Synthetic generator: names, emails, addresses, numbers, and many more, searchable by name or category.
  4. Repeat for each column you need, then click Save Changes. A blank dataset can mix synthetic columns with columns you add later From CSV File.
  1. Select a column in the workbench to open its Column Inspector.
  2. Under Transformation Pipeline, pick a transform and click Add: trim, uppercase or lowercase, prefix or suffix, regex extract or replace, date or number formatting, a default for empty values, and more.
  3. Watch the before/after preview update against real rows as you fill in the transform’s fields, so you catch a wrong pattern before you save.
  4. For logic no built-in transform covers, add a JavaScript snippet transform and click Manage snippets. Use Search snippets to find one your team already wrote, or write a new one. The snippet you pick shows as Pinned: the dataset keeps that revision, so a later edit to the snippet never changes a dataset that’s already saved.
  5. Click Save Changes.
  1. Open the test and switch to its Data tab. The Test data tab is selected.
  2. Pick the dataset from Bound Dataset.
  3. Click Save binding. The tab shows Binding saved once it’s attached.

Split across runners and end-of-file behavior

Section titled “Split across runners and end-of-file behavior”
  1. On the same Test data tab, set Multi-runner distribution strategy: Disjoint slices (split) gives each runner its own range of rows; Full replicated gives every runner the whole dataset.
  2. Check Runner allocation to see the row range MaxoPerf plans to give each runner, based on the test’s resolved load configuration.
  3. Set End-of-file behavior: Recycle dataset loops over the rows, Stop at end stops the virtual user when they run out. If the control is disabled, MaxoPerf can’t patch this engine and script for end-of-file behavior yet. The reason is shown below the control, and a preference you saved earlier stays stored until you switch to a combination that supports it.
  4. Click Save binding.
  5. After a run, open the run’s Data tab to see the manifest MaxoPerf actually used: the bound dataset, its distribution, and each delivered file’s path, row count, and checksum (one slice per runner when split, one shared file when replicated).

Every route into a bound dataset shares the same three facts, shown on the test’s Data tab under Use in your script:

WhatValue
File/work/run/data/<file-name>.csv inside the runner (relative to the test bundle: data/<file-name>.csv)
Env var, one per datasetMAXOPERF_DATASET_<FILE_NAME>_CSV (the file name upper-cased, - turned into _)
Data directoryMAXOPERF_DATA_DIR = /work/run/data

The first row of the delivered file is the header, so the dataset’s own column names are the variable names your script uses. With split distribution each runner’s file holds only its own rows; the path and env var stay the same on every runner.

The File name is set when you save the binding (default: the dataset’s name, slugified). Renaming the dataset, or swapping in a different dataset on the same binding, keeps it. Change the File name on the Test data tab and MaxoPerf shows what moves before you save: the file, the env var, and, for k6 and Locust, the injected row variable. Update your script at the same time.

Runners export these variables; self-hosted outposts need an up-to-date runner image.

The examples below use a dataset bound with the File name checkout-skus and a sku_id column.

Point a CSV Data Set Config (CSVDataSet) element at checkout-skus.csv or data/checkout-skus.csv. MaxoPerf matches the element by that file name, rewrites its filename to the delivered file, and sets its variable names to the dataset’s columns (skipping the header row). Any variable names the element declared are replaced, so reference the dataset’s column names: ${sku_id}. Every bound dataset needs exactly one CSVDataSet that reads it.

This is JMeter’s only supported route: a .jmx with any CSVDataSet always gets it patched, so reading the file through the env var instead isn’t supported for JMeter.

On the Test data tab, End-of-file behavior turns on once MaxoPerf recognizes the CSVDataSet: choose Recycle dataset or Stop at end, and MaxoPerf writes that choice into the config element for the run. If the control stays disabled, the reason is shown below it.

Use one of these routes, not both.

Auto-inject. Put the marker comment // @maxoperf-data-bindings on its own line inside your default function. MaxoPerf loads the file and inserts the row lookup at the marker. The injected code may return once the dataset runs out, which is how Stop at end ends the iteration. Read columns from maxoperf_row_<name>, where <name> is the File name with - turned into _ (and a leading _ added when it starts with a digit). Here that’s maxoperf_row_checkout_skus.sku_id:

import http from 'k6/http';
export default function () {
// @maxoperf-data-bindings
http.get(`https://shop.example.com/api/products/${maxoperf_row_checkout_skus.sku_id}`);
}

Read it yourself. Load the file from the env var in init code. End-of-file behavior must stay Recycle dataset; Stop at end needs the marker.

import http from 'k6/http';
import { SharedArray } from 'k6/data';
import papaparse from 'https://jslib.k6.io/papaparse/5.1.1/index.js';
const skus = new SharedArray('checkout-skus', () =>
papaparse.parse(open(__ENV.MAXOPERF_DATASET_CHECKOUT_SKUS_CSV), { header: true, skipEmptyLines: true }).data,
);
export default function () {
const row = skus[__ITER % skus.length];
http.get(`https://shop.example.com/api/products/${row.sku_id}`);
}

Use one of these routes, not both.

Auto-inject. Put # @maxoperf-data-bindings on its own line inside the task method where MaxoPerf should pull a row, then read self.maxoperf_row_<name>["column"], with <name> the File name and - turned into _. Here that’s self.maxoperf_row_checkout_skus["sku_id"]. When the rows run out under Stop at end, MaxoPerf stops that user.

from locust import HttpUser, task
class Shopper(HttpUser):
host = "https://shop.example.com"
@task
def view_product(self):
# @maxoperf-data-bindings
sku = self.maxoperf_row_checkout_skus["sku_id"]
self.client.get(f"/api/products/{sku}")

Read it yourself. Open the file from the env var. Keep End-of-file behavior on Recycle dataset.

import csv
import itertools
import os
from locust import HttpUser, task
with open(os.environ["MAXOPERF_DATASET_CHECKOUT_SKUS_CSV"], newline="", encoding="utf-8") as f:
ROWS = list(csv.DictReader(f))
class Shopper(HttpUser):
host = "https://shop.example.com"
def on_start(self):
self.rows = itertools.cycle(ROWS)
@task
def view_product(self):
row = next(self.rows)
self.client.get(f"/api/products/{row['sku_id']}")

Use one of these routes, not both.

Auto-inject. Put // @maxoperf-data-bindings on its own line where the .feed(...) step belongs in your scenario chain. MaxoPerf defines the feeder at the top of your simulation class and inserts the .feed at the marker, so the script must declare exactly one class or object. Reference columns afterward as #{sku_id}. Gatling has no injected row variable: the columns are session attributes.

import io.gatling.core.Predef._
import io.gatling.http.Predef._
class CheckoutSimulation extends Simulation {
val httpProtocol = http.baseUrl("https://shop.example.com")
val scn = scenario("Checkout")
// @maxoperf-data-bindings
.exec(http("GET product").get("/api/products/#{sku_id}"))
setUp(scn.inject(atOnceUsers(10))).protocols(httpProtocol)
}

Read it yourself. Build your own feeder from the env var, for example csv(sys.env("MAXOPERF_DATASET_CHECKOUT_SKUS_CSV")).circular, and .feed it where you need it. Keep End-of-file behavior on Recycle dataset.

For a test built in the builder, and for an uploaded Taurus YAML whose scenario runs on apiritif, MaxoPerf adds or updates the scenario’s data-sources entry for you. It points at data/<file-name>.csv, carries your End-of-file behavior, and has no variable-names list: the header row is the only source of names that JMeter and apiritif read, and a declared name list would push the header into the data itself. MaxoPerf replaces the whole entry for that path, so extra keys you set on it don’t survive.

An uploaded Taurus YAML that runs on JMeter (the default when no executor is named) isn’t edited. Write the data-sources block yourself with the same shape (path: data/<file-name>.csv, no variable-names) and keep End-of-file behavior on Recycle dataset. A YAML that wraps a k6, Locust, Gatling, or .jmx script follows that engine’s route above.

Playwright, Selenium, and anything else MaxoPerf can’t patch read the CSV from the env var. This works as long as every binding on the test keeps End-of-file behavior on Recycle dataset. MaxoPerf refuses to start a run that pairs a script it can’t patch with Stop at end, rather than running with a setting it can’t honor.

The auto-inject marker belongs in exactly one script, once. A marker that appears twice, or in two scripts, also stops the run from starting, whatever the End-of-file behavior.