From Production Traffic to a Load-Test Plan: A Practical Workflow
Turn production evidence into a load-test plan with explicit journeys, demand assumptions, safety limits, and measurable acceptance criteria.
A good load-test plan is a model of demand you can defend: who arrives, what they do, how often they retry, and which outcome would make the release unsafe. A list of endpoints is not enough. Start with production evidence, then make every assumption visible enough for another engineer to challenge.
Start with a decision, not a tool
Write the decision the test must support in one sentence. “Can checkout sustain the holiday peak while keeping the purchase journey within its service objective?” is useful. “Run 10,000 users” is not. The first tells you what to measure and what a pass means. The second only names an input.
Name the audience for the result and the deadline. A release gate needs a repeatable, relatively small test. An event-readiness review may need a larger rehearsal, a recovery exercise, and an explicit abort plan. Those are different tests even if they exercise the same URL.
Turn evidence into journeys
Collect more than a request count. Look at access logs, tracing summaries, product analytics, queue depth, and the business calendar. Remove health checks and known automation where you can identify them. Keep a note of what the source cannot distinguish: a server request may represent a cache hit, a browser retry, or a person refreshing a page.
Group the evidence into journeys a person or service recognizes. Examples include browse, search, sign in, add to cart, checkout, export, and an asynchronous job poll. Keep a journey intact long enough to reveal state and dependencies. A test that calls POST /checkout without first setting up a cart ends up measuring a validation error.
For each journey, record its share of arrivals, requests per journey, data requirements, and business consequence. A five-percent administrative workflow may be more important than a sixty-percent catalog read if it is the path that commits money or changes inventory.
Choose the control variable
Decide whether the model controls concurrent actors or an arrival rate. A closed model lets each virtual user finish before it begins another iteration, so rising latency lowers the rate at which that user starts work. An open model schedules arrivals independently and exposes queueing pressure more directly. Read open and closed workload models before choosing by habit.
State the ramp, hold, and stop conditions in plain language. “Ramp from the quiet baseline to the expected peak over fifteen minutes, hold for thirty, then step up until the error budget is spent” is reproducible. Add the expected concurrency or arrival rate at each stage, and label it as a model, not a forecast you promise to hit.
Make assumptions calculable
Suppose an event produces 1,200 browse arrivals per minute, 180 searches, and 60 checkout starts. Those values are hypothetical examples, not a claim about your system. If browse creates four requests and search creates six, the rough request rate is (1,200 × 4 + 180 × 6 + 60 × 8) / 60, or 114 requests per second before retries and background calls. Write the arithmetic down so changes to the journey mix have a visible effect.
Express uncertainty as scenarios, not as an unexplained multiplier. A baseline might use the observed peak. A planning case might raise arrivals by a stated factor. A stress case may shorten the arrival window. Include cache state, authentication mix, payload size, and dependency behavior. Each scenario should say what question it answers.
Design data and observability together
Create data that behaves like the production shape without copying personal information. Decide which records can be shared, which must be unique, how they are reset, and what happens when the pool runs out. A duplicate account can turn a performance test into an authentication test. An empty catalog can turn a useful cache test into an artificial best case.
Instrument the target and the generator. Capture latency distributions, status and assertion failures, achieved arrivals, queue time, saturation, connection-pool waits, and resource throttling. Link the run to a release and record the environment version. Without its workload and environment, a result is a number nobody can reproduce.
Define the pass before the run
Use criteria tied to user and system consequences: p95 latency for checkout, an error budget for reads, no lost writes, bounded queue time, and a generator-health condition. Avoid one global average. Decide whether a brief warm-up is excluded, how retries count, and what invalidates the run.
End with a reviewable plan: decision, journeys, weights, model, scenarios, data, telemetry, thresholds, abort conditions, and owner. Revisit it when production evidence or the release boundary changes. Keep old assumptions beside new ones so the comparison stays readable and history is not quietly rewritten.
A compact review checklist
- Is every major journey represented by a stateful flow rather than a bare endpoint?
- Are traffic sources, bot filtering, retries, cache state, and seasonal assumptions recorded?
- Does the load model match the way demand enters the system?
- Can the generator prove it achieved the requested workload?
- Are data collisions and cleanup observable?
- Are pass, abort, and invalid-run rules written before execution?
If another engineer can answer those questions from the document and reproduce the scenario, the plan is ready to script.
Example review notes
For a subscription product, the evidence might show a weekday peak at 10:00, a release-day spike at 10:05, and a separate nightly export. Treat those as three scenarios. The first asks about steady customer work, the second asks whether admission and autoscaling react, and the third asks whether background work steals capacity. One blended “peak” number cannot answer all three.
Before implementation, write a stop rule such as “halt if checkout assertion failures exceed the approved budget for two consecutive windows, or if synthetic writes appear outside the test partition.” Add a recovery check: queues drain, created records are removed, and alerts return to their prior state. Safety then becomes a testable part of the plan instead of a hope.
Finally, have the service owner sign off on assumptions that are not measurable yet. Assumptions are fine. Unlabelled assumptions are the flaw. When new production evidence arrives, update the weights or scenario and preserve the older run for comparison. With that history, a load-test plan improves over time instead of becoming a one-off script.
Make the model reviewable
Before production observations become test data, remove credentials, customer identifiers, and payload fields that the exercise does not need. Preserve relationships that affect work, such as route mix, object size, cache state, and the order of dependent actions. Review the transformation with the people responsible for privacy and target safety, then record which dimensions were intentionally changed. A smaller sample you can explain beats a large export no one can authorize or reproduce. Compare the planned mix with the source window after conversion and flag any route that was dropped, merged, or overrepresented. The review also catches a familiar trap: treating a traffic chart as a complete workload model when it only measured successful requests. Include rejected, retried, and asynchronous work when those outcomes change system pressure, and state where the load plan deliberately stops.
Include a table with journey, arrival share, requests or transactions, state, data partition, and objective. Add a second table for assumptions: source, transformation, confidence, and the experiment that would reduce uncertainty. A search journey might be counted from server traces, divided by three requests per journey, and marked medium confidence because retries are not separated. Its follow-up experiment identifies first attempts and retries.
State the non-goals: the plan may not prove client rendering, a provider’s private quota, or another region’s capacity. A precise limitation belongs in a good result. It tells the next engineer what evidence to add and heads off an unsupported conclusion.
Version the brief with the test artifact. When a traffic estimate changes, update the scenario deliberately and keep the previous run as a baseline. The plan then keeps explaining demand as it changes, instead of becoming a script with an unexplained constant.
Keep the transformation reversible enough that a reviewer can trace each scenario assumption back to an approved observation without exposing sensitive production data.