The Pre-Run Load Testing Checklist That Prevents Misleading Results
Check authorization, data, telemetry, generator capacity, abort conditions, and recovery before every meaningful load test.
Many misleading load tests fail before the first ramp step. The target is wrong, the data runs out, the generator misses its rate, or the dashboard cannot tell a test error from a service error. A short pre-run check protects both the system and the decision you will make from the result.
1. Confirm the question and target
Write down the decision, target origin, version, environment, and start window. Get authorization from the owner. Check that the hostname, region, tenant, and feature flags match the approved scope. A valid test against the wrong target is still a failed change-control process.
2. Prove one complete journey
Run one user through each critical flow. Check authentication, correlation, redirects, response shape, business assertions, asynchronous completion, and cleanup. Inspect the state changes. A fast 401 or a validation error is not a successful checkout.
Keep the one-user proof as an artifact. When a distributed run fails, you have a known-good run to compare against. The anatomy of a load-test result lists the evidence worth keeping.
3. Check data and secrets
Make sure the dataset holds enough records for the planned concurrency and duration. Confirm allocation, uniqueness, collision handling, reset, and expiry. Use synthetic identities and secret references. Never paste credentials into scripts, logs, screenshots, or result labels.
If the pool is finite, run a small exhaustion test. Decide what exhaustion does: invalidate the run, end the scenario, or trigger controlled recycling. Recycling can change cache behavior, so label it.
4. Validate workload achievement
Check the model, ramp, hold, pacing, request boundary, and retry policy. Estimate the concurrency and arrivals you expect. During warm-up, confirm that scheduled and achieved rates converge. Give the generator headroom on CPU, memory, socket count, bandwidth, and event-loop health.
If the worker cannot reach the requested arrival rate, stop. Fix it or add generators. A generator-capacity result says nothing about target capacity, so do not publish it as one.
5. Verify telemetry and clocks
Confirm you can see request distributions, status and assertion errors, queue time, active work, and target saturation. Add connection pools, databases, caches, dependencies, and autoscaling signals where they test the hypothesis. Check time synchronization, and carry a run identifier into logs and traces.
Open the dashboards before you start. If an alert has no owner, or nobody can read its signal during the window, the test is not ready. Observability added after a failure rarely recovers the missing evidence.
6. Write safety limits
Record the maximum arrival rate or concurrency, the duration, sensitive operations, rate limits, abort thresholds, and who can stop the run. Ramp in a controlled way instead of trusting a limit to protect an unfamiliar target. For production work, write down the customer-impact indicators and a communications path.
Test the stop path in a safe environment. A stop button that halts new arrivals while workers keep retrying, or queues drain with no end in sight, is incomplete. Name who owns cleanup and recovery.
7. Define pass, fail, and invalid
Set criteria per journey: p95 or p99 latency, error budget, business assertions, queue delay, completion time, and recovery. State the warm-up and steady-state windows, the minimum sample count, and how retries are counted. Mark a run invalid if the target version changed, data ran out, the workload was missed, telemetry broke, or an unplanned dependency joined in.
8. Review the plan aloud
Ask a second engineer to answer: what is sent, from where, with which identities, for how long, at what rate, and what stops it? If an answer rests on an assumption that is not in the plan, fix the plan before you run.
A printable final pass
- Target and authorization verified.
- One-user journeys and assertions pass.
- Data, secrets, cleanup, and collision rules verified.
- Load shape, retries, limits, and generator capacity checked.
- Dashboards, logs, clocks, and run identity ready.
- Stop, abort, recovery, and communication paths tested.
- Pass, fail, and invalid-run rules approved.
Keep the scenario and the evidence together so you can repeat the run. The checklist still matters for a local run: it protects the result before scale makes every ambiguity expensive.
Turn the checklist into a gate
Automate the first checks where you can. Resolve the target, verify required configuration, run a one-user journey, assert fixture capacity, and fail if telemetry is unreachable. Keep human approval for authorization, maintenance windows, and destructive scope. Record who approved and when the approval expires.
For a long run, set checkpoints at warm-up, steady state, ramp changes, and recovery. At each one, compare planned and achieved workload, errors, queues, target saturation, generator health, and remaining data. A pre-run check cannot catch every mid-run failure. It makes your response predictable when one appears.
A good checklist is short enough to use and specific enough to stop an unsafe run. Drop generic confirmations that never change a decision. Add a check each time a past failure exposes a missing assumption.
The handoff after a failed check
If authorization is missing, stop and send the plan back to its owner. If one-user assertions fail, fix the flow before you scale. If fixtures run short, enlarge or partition them and rerun the exhaustion check. If the generator misses its rate, classify the result as client capacity. If telemetry is down, do not read a quiet alert channel as a green outcome.
Give every unchecked item an owner and a safe next move. After the fix, record the rerun and keep the original failed check as evidence of what the control prevented.
After the run, turn a surprising failure into a checklist update when it is a repeatable risk. Add the exact condition, the evidence you expect, and an owner. That keeps the checklist in step with the system instead of drifting into a static page.
For event readiness, add a second checklist for recovery. Stop new arrivals, watch queued work, verify it completes, inspect synthetic state, and confirm alerts return to normal. A test that ends at the last ramp step skips the question operators care about most. Name the recovery owner in the run brief before the window opens.
Make the brief executable
Put the run brief where the operator can use it during the window. Name the exact build, environment, public origin, scenario version, fixture version, scheduled load, duration, and allowed traffic boundary. Include who can abort and who can restore the environment. “Run the checkout test” cannot be reproduced when a later engineer has to guess the inventory, account policy, or rate limit you meant.
Check the generator separately from the target. Verify DNS and TLS resolution, credentials or synthetic tokens, data availability, worker CPU and network headroom, and the expected worker count. Start a one-user control and assert the business outcome, not only a status code. Then run a low-rate control and compare scheduled with achieved arrivals. If either control fails, stop and repair the setup. Do not spend the full window collecting invalid latency.
Define the evidence before the first ramp step. Decide which latency percentile, error classes, queue signals, dependency metrics, and logs you will keep. Confirm timestamps come from comparable clocks and that labels hold no secrets or unbounded identifiers. Set the abort threshold for customer impact and the recovery observation period. A threshold you set after a failure is a post-hoc explanation. It does not make the run safe.
At the end, leave the environment usable for the next team. Drain asynchronous work, remove synthetic records, restore feature flags or routing, and check that alerts are back to normal. Record anything manual or incomplete. The checklist works when another operator can tell whether the system, the generator, and the evidence are all ready, not just that traffic started.