Skip to content

Stress test

A stress test pushes your system past its expected operating range on purpose. It finds the point where performance degrades or errors start to appear. Breaking the system is a means, not the aim. You want to learn where the limit is and how the system behaves as it gets close.

  • A load test has established baseline behaviour at expected traffic levels.
  • You have authorisation to test the target, and stakeholders know a stress run may cause elevated error rates or temporary unavailability.
  • Preferably run against staging, not production.

A stress test ramps the virtual user count steadily upward, past the number used in a normal load test, until one of the following occurs:

  • Latency climbs past an unacceptable threshold (e.g., p95 > 2 s).
  • Error rate rises above an acceptable ceiling (e.g., > 1 %).
  • The system becomes unresponsive and the run itself starts to fail.

The inflection point, where errors begin to spike, marks the breaking point of the current configuration. Everything below it is the system’s safe operating range.

ParameterValue
Virtual users (VUs)Ramp from baseline to 2×–5× expected peak
Ramp-upSlow and steady: 1–2 VU/s or 5-min stages
Hold duration5–10 min at peak; stop when limits are hit
Stop modeDuration (or stop manually when errors spike)
LocationsSame as your load test for comparison
  1. Duplicate your load-test test file and rename it api-stress (or similar).

  2. Edit the Taurus YAML to use a higher VU target and a gradual ramp:

    execution:
    - executor: jmeter
    concurrency: 500
    ramp-up: 10m
    hold-for: 5m
    scenario: api-stress
    scenarios:
    api-stress:
    requests:
    - url: https://api.staging.example.com/v1/products
    label: list-products

    This ramps from 0 to 500 VUs over 10 minutes, then holds for 5 minutes. The peak is roughly 5× the baseline load-test VU count of 100 VUs.

  3. In Load profile, set Virtual users to 500, Ramp-up to 10m, and Duration to 15m (ramp + hold).

  4. Set a Failure criteria threshold to capture the breaking point automatically. For example, error rate > 5 % makes the run fail and records the moment.

  5. Click Run and watch the Overview tab live. The latency chart shows the inflection clearly.

Open the Overview tab after the run:

  • Latency chart. Find the inflection point where p95 climbs fast. The VU count at that point is your effective breaking point.
  • Error rate panel. A step from near-zero to a rising error rate shows where the system saturates.
  • Throughput. It often flattens or drops before the error rate climbs, which means requests are queuing in the application or database.

Record the VU count at the inflection point and compare it against your expected peak load. If the breaking point is below 2× expected peak, your system needs more capacity margin.

DoDon’t
Start from the load-test baseline and change only the VU countRun a stress test before a smoke and load test
Ramp gradually to observe the degradation curveJump instantly to a high VU count and miss the inflection
Document the breaking-point VU count for capacity planningDraw conclusions from the peak VU count alone
Run on staging with stakeholder awarenessRun on production without a maintenance window