Performance Test Requirements: 12 Questions to Answer Before You Script
Use 12 practical questions to define a performance test's decision, workload, data, observability, and pass criteria before scripting.
A performance script can be technically perfect and still answer the wrong question. Before you pick an engine or record a journey, write down the requirements that make a result actionable. The twelve questions below fit into a short review and expose most of the costly ambiguities.
Decision and scope
1. What decision will this test support? Name the release, event, capacity change, or incident question. “Prove the API is fast” is not a decision. “Can the order service sustain the forecast peak without breaking the checkout objective?” is one.
2. Which journeys and boundaries are in scope? List business flows as well as routes. Say where the test starts and ends, and whether each dependency is real, virtualized, or excluded. If checkout is in scope and payment authorization is not, say how that boundary responds.
3. Which user or system outcomes matter most? Rank them. A successful login, a completed write, a correct export, and a fast catalog read do not carry equal weight. The ranking should shape traffic weights and failure criteria.
Demand and environment
4. What demand shape should the test reproduce? Record peak, ramp, hold, seasonality, bursts, retries, and background work. Choose VUs or an arrival rate based on how demand enters the system. If the difference is unclear, read open and closed workload models.
5. Which numbers are measured and which are assumptions? Keep observed traffic apart from a planning multiplier and a hypothetical stress case. Give every assumption an owner and a reason. Do not bury uncertainty in an unlabeled “realistic” load value.
6. How production-like must the environment be? Specify versions, topology, caches, data volume, quotas, network path, feature flags, and dependency behavior. A smaller environment can still help, as long as you understand its bottlenecks and do not present the result as a direct claim about production capacity.
Flow, data, and safety
7. What state does each journey require? Document authentication, tokens, carts, records, correlation IDs, and asynchronous completion. A flow that skips setup may get a fast validation error back instead of doing the intended work.
8. Which data is shared, finite, or unique? Define allocation, reuse, collision behavior, refresh, and cleanup. Remove personal data. Decide what happens when the dataset runs out: stop, recycle with a known cache effect, or fail the scenario.
9. What makes the test safe to run? Record target authorization, maintenance windows, rate limits, sensitive operations, abort conditions, and a recovery owner. Include a maximum duration and a hard load ceiling. Safety belongs in the test contract, not in an informal chat message.
Evidence and judgment
10. What telemetry must be available? Require request distributions, status and assertion errors, queue time, achieved workload, generator health, and target saturation. Add database, cache, or dependency signals where they can explain the decision. Link the run to the version and environment.
11. What is a pass, fail, or invalid run? Set thresholds per journey and by consequence. Specify the percentile, window, error classification, and whether warm-up is excluded. A run is invalid if the generator missed its rate, the dataset ran out, or the target version changed unexpectedly.
12. What happens after the result? Name the triage owner, the artifact retention period, the comparison baseline, and the next experiment. A red result with no route to diagnosis is only a report. Nothing feeds back.
Turn answers into a one-page brief
Put the answers in a short brief with a table for journeys, weights, data, and criteria. Describe the path in prose: “arrival scheduler → edge → API → queue → worker → database.” Note where timing is measured and where a retry can multiply work. Ask someone outside the scripting effort to challenge the assumptions.
The brief should also say what the test cannot prove. A protocol test cannot prove visual rendering time. A synthetic dataset cannot prove production-specific hot partitions. A single region cannot prove global network behavior. Stating the limits makes the result easier to trust.
Only then choose implementation details. Whether the team runs Taurus, JMeter, k6, or another supported engine follows from the requirement. MaxoPerf helps when the artifact, execution context, and comparison result need to stay together, but no platform can fix a question nobody defined.
The pre-scripting review
Read the brief aloud and check that someone could answer these questions: what are we sending, to whom, from where, for how long, with which data, and what evidence changes the decision? If any answer needs a guess, keep scripting on hold. Fifteen minutes of precise requirements often saves days of debugging a script that was never valid.
Questions that expose hidden assumptions
Ask what happens when a response is slow, when a record is missing, and when the generator cannot keep up. The answers expose the intended backpressure, fixture policy, and validity rule. Ask where a request is timed, whether a retry counts as new demand, and whether a successful transport response can carry a failed business result. Write the answers next to the requirement instead of leaving them in an engineer’s head.
An export journey shows why. Its requirement may read “a requested export is acknowledged within two seconds and completed within five minutes for the planned mix.” That differs from timing only the POST response. The test needs a job ID, bounded polling, a completion assertion, timeout classification, and cleanup. Writing this down before scripting stops anyone from mistaking a fast acknowledgement for completed work.
Review the brief with the person who owns the user outcome and the person who owns the environment. The first can challenge priorities and thresholds. The second can challenge topology, quota, and dependency assumptions. Keep disagreements visible and settle them before the script creates false precision.
A requirement is a boundary contract
Write down the start and end event for every measurement. For an API, the start may be the request leaving the generator and the end a validated response. For an order, the end may be durable confirmation after an asynchronous worker. For a browser flow, the end may be a result the user can see. All of these boundaries can be valid, but you cannot compare them silently.
Add the expected invalid states: expired token, exhausted fixture, dependency timeout, wrong schema, and missed arrival schedule. Give each one a classification. The run can then tell “the product failed” apart from “the test was not valid.” Name a recovery owner for every destructive or long-lived state.
Once the requirements are approved, do not add a new metric just because it is easy to collect. Every signal should answer the decision or explain a failure. That keeps the report readable and stops a wall of telemetry from hiding the one criterion that matters.
Keep a decision log when a requirement changes. If the owner shifts from peak capacity to recovery time, update the load shape, telemetry, and threshold together. If you leave the old workload in place and change only the headline, the evidence no longer matches the test’s purpose.
Add a sign-off line for the scope owner, data owner, and environment owner. Their approval should cover the question and its limits. It does not certify an outcome in advance. The test may still fail. What matters is that its failure is safe and interpretable.
The brief should make that criterion impossible to miss.
Turn ambiguity into a test matrix
For each requirement, record the start event, end event, state, dependency, expected outcome, and invalid condition. An export can be acknowledged quickly and take minutes to finish, so its matrix needs both acknowledgement and completion criteria. A token refresh can return a valid status while issuing a token that does not work, so its matrix needs an authenticated follow-up call. These details stop a green result at the transport level from passing for a green result at the business level.
Ask the requirement owner which evidence would change the release decision. If the answer is “nothing yet,” do not add a threshold just to have one. Add instrumentation, or leave the question explicitly open with an owner. Requirements earn their keep when they tell the script what to prove and tell the reviewer what not to infer.