Skip to content

Test data dos and don'ts

Test data is a common and often overlooked cause of unrealistic load test results. If every virtual user sends the same hardcoded values, the test bypasses caches, creates database hotspots that production never has, and behaves nothing like real traffic. This page shows how to use MaxoPerf’s data features correctly, and what to avoid.

  • Read about MaxoPerf data entities and parameters before this page.
  • Know the difference between a data entity (a CSV file attached to a workspace) and a variant (a parameterized version of a test that selects a different row subset or column mapping from the same entity).
  • Have a CSV file with realistic data ready, or know how to generate synthetic data that covers your key paths.

When 200 virtual users all log in as user@example.com in parallel:

  • Your session table takes every write for a single user ID.
  • Your auth token cache serves the same token to all VUs, so tokens never refresh the way they would for real users.
  • Your rate-limiter sees one user making 200 requests/second and blocks them all.
  • Row-level locking in your database piles all the contention onto a single row.

None of this happens in production. Your test results describe an extreme edge case, not your system under real load.

  • Do use a MaxoPerf data entity (CSV) for user credentials, IDs, and any value that should vary per VU. Upload your CSV at workspace level, attach it to the test, and reference it by the parameter name in your script. Each VU gets a unique row.

    In Taurus, reference the attached CSV entity by the parameter name you mapped:

    scenarios:
    authenticated-browse:
    requests:
    - url: /api/products
    headers:
    Authorization: 'Bearer ${AUTH_TOKEN}'

    The attached data entity supplies the AUTH_TOKEN parameter, and each VU draws its own row.

  • Do use MaxoPerf Test Data variants when you want to run the same test against different data sets without duplicating test configurations. A variant selects a different row range or column mapping from the same entity. Use variants to run the same profile against dev, staging, and prod data sets.

  • Do generate enough rows for your peak VU count plus a buffer. If your test runs 500 VUs for 30 minutes with a 2-second think time, each VU makes roughly 900 requests. With 500 rows in your CSV and sequential allocation, rows will repeat. Use at least 1.5× your VU count in rows for non-repeating runs. For realistic variety, aim for 10× or more.

  • Do validate your CSV before uploading. With a malformed CSV (wrong delimiter, missing header, BOM characters), Taurus quietly falls back to empty parameter values. Check the file locally with a CSV linter, or run a smoke test first.

  • Do use MaxoPerf Secrets for credentials you do not want visible in plain text in your script files. Upload secrets at account or workspace level; reference them as environment variables in your script. The secret value is never stored in the test file.

    env:
    API_KEY: '${API_KEY_SECRET}'
  • Do prefer synthetic data over production data when you can. Synthetic data avoids GDPR/privacy concerns (no real PII in test artifacts), lets you create edge-case rows on purpose, and cannot delete or change real production records by accident.

  • Do scope your data to the test environment. Use test-environment user accounts, test product IDs, and test payment tokens. A load test against production data can create real orders, charge real payment methods, or trigger real notifications.

  • Don’t hardcode usernames, tokens, or IDs in your Taurus YAML or JMeter JMX. Hardcoded values end up in version control, where they go stale, and they cause the single-user contention problem described above.

  • Don’t put secrets (API keys, passwords, tokens) directly in script files. Even in a private repository, secrets in script files leak into build logs, test artifacts, and run history. Use MaxoPerf Secrets instead.

  • Don’t upload real production user data (names, emails, addresses) as a CSV data entity. Test artifacts, including attached data entities, may be visible to all workspace members. Replace PII with synthetic data before uploading.

  • Don’t use a CSV with fewer rows than your VU count for operations that need unique values. Uniqueness matters for login (session conflicts), resource creation (unique constraint violations) and rate-limiting (per-user limits). Fewer rows than VUs guarantees collisions.

  • Don’t share a single CSV across tests with incompatible column layouts. If two tests map a shared CSV’s columns differently, one of them quietly gets wrong parameter values. Give columns descriptive names and create separate entities for incompatible layouts.

  • Don’t use static correlation values (tokens, CSRF nonces) copied from a manual session. Static tokens expire. Once the token TTL runs out, your VUs start failing with 401/403 errors. Use dynamic token extraction (Taurus extract-jsonpath or JMeter Extractor) to fetch a fresh token in each VU’s session setup.

  • Don’t ignore the “attached to count” indicator on your data entity. If a data entity is attached to many tests, changing its column mapping or deleting rows can break all of them. Keep widely used entities stable, and create a new entity (or variant) for changes.

QuestionWhy it matters
Does every credential/ID vary per VU?Avoids single-user contention
Does the CSV have ≥ 1.5× the VU count in rows?Avoids row wrap-around collisions
Are secrets stored in MaxoPerf Secrets (not the script)?No credential leakage in artifacts
Is the data synthetic or de-identified?No PII in test artifacts
Are tokens fetched dynamically (not hardcoded)?Tokens stay valid for the full run
Does the data belong to the correct test environment?No accidental production mutations