Skip to content

CSV data-driven test

Problem: Your test needs a pool of distinct users, products, or inputs so that each virtual user works on its own data. That avoids cache collisions and unique-constraint conflicts on writes, and it keeps read distributions realistic. MaxoPerf datasets are the built-in way to do this.

Feed real data into a load test.

Test type: Any (Load test, Stress test, Soak test).

  • A MaxoPerf account and a workspace.
  • A CSV file with a header row and at least one data row per intended virtual user.
  • A Taurus YAML test that will consume the CSV columns.

A well-formed data file:

email,password,user_id
user001@example.test,P@ssw0rd!,u-001
user002@example.test,P@ssw0rd!,u-002
user003@example.test,P@ssw0rd!,u-003

Guidelines:

  • Header row names become the dataset’s column names. The Taurus scenario reads them straight off the header, so don’t declare a separate name list.
  • One row per distinct entry. For 50 VUs, 50+ rows avoids recycling.
  • Never put real user credentials in version-controlled files. Generate synthetic ones.
  1. Open Test data in the left navigation and click New Dataset.
  2. Choose Import CSV and select the file. MaxoPerf detects the delimiter and the header row and previews the first five rows.
  3. Keep or change the Dataset Name, for example Login users, then click Create Dataset & Open Workbench. Each header becomes a column.
  1. Open the test and switch to its Data tab. The Test data tab is selected.
  2. Pick the dataset from Bound Dataset and click Save binding.
  3. Under Use in your script, note the File name (the dataset’s name, slugified, by default) and the env var MaxoPerf will export for it. Keep the default unless your script needs a specific one.

4. Write the Taurus scenario to use the columns

Section titled “4. Write the Taurus scenario to use the columns”
execution:
- scenario: data-driven-login
concurrency: 50
ramp-up: 2m
hold-for: 5m
scenarios:
data-driven-login:
data-sources:
- path: data/login-users.csv
random-order: false
requests:
- label: POST /auth/login
url: https://api.example.com/auth/login
method: POST
headers:
Content-Type: application/json
body: '{"email": "${email}", "password": "${password}"}'
- label: GET /users/${user_id}/profile
url: https://api.example.com/users/${user_id}/profile
method: GET

The path: data/login-users.csv matches the File name MaxoPerf shows on the Data tab (login-users for a dataset named Login users). The column headers supply ${email}, ${password}, and ${user_id}, so don’t declare a separate name list: the header row is the only source of names Taurus reads, and a declared list would turn the header into a data row.

This scenario names no executor, so it runs on JMeter, Taurus’s default. MaxoPerf delivers the file but leaves a JMeter-run YAML exactly as you wrote it: keep the data-sources block, and keep End-of-file behavior on Recycle dataset (the default). To have MaxoPerf write the data-sources entry for you and unlock Stop at end, add executor: apiritif to the execution entry. MaxoPerf then replaces the entry for that path with its own, so drop random-order. See Use the data in your script.

  1. Switch to the Files tab and upload the Taurus YAML as the entrypoint.
  2. Confirm the dataset is bound (Data tab).
  3. Click Run now.
  • The run overview shows requests against multiple distinct user IDs (no single user_id dominates the error log).
  • The POST /auth/login label has a near-zero error rate. If you see 401 errors, check that the CSV credentials are valid in the target environment.
  • The GET /users/${user_id}/profile label shows a realistic response-time distribution. A suspiciously uniform one points to caching on a single resource.
  • The run’s Data tab lists the file MaxoPerf delivered, its row count, and its checksum.
  • Random order: on the JMeter route above, set random-order: true in your data-source block to shuffle rows instead of handing them out in sequence. On apiritif MaxoPerf rewrites the entry, so the setting doesn’t survive.
  • Multiple datasets: bind more than one dataset per test for independent data pools (e.g. products + users), each with its own File name and env var.
  • Split across runners: set Multi-runner distribution strategy to Disjoint slices (split) so each runner gets its own row range instead of the full file.
  • Large datasets: MaxoPerf streams the CSV to runners instead of loading it all into memory. Files with hundreds of thousands of rows work without special configuration.