Skip to content

Video streaming do and don't

Streaming load tests can fail in ways REST API tests cannot. A test can run, produce results and show plausible throughput, and still be completely wrong because of a missing sleep() or a misread metric. This page collects the do/don’t rules that matter most from across the streaming testing section.

Follow every segment fetch with a pause equal to the segment duration. No rule in streaming load testing matters more.

# Correct — real-time cadence
- label: segment
url: https://cdn.example.com/stream/720p/seg-0002.ts
think-time: 4s # 4-second segment = 4-second wait

Without think-time, a single VU downloads segments as fast as the network allows: potentially 50–200 per second instead of 0.25 per second. Your “1,000 viewer” test then generates the load of 200,000+ viewers. It maxes out your CDN or fails on connection limits, and the results say nothing useful about behavior at viewer scale.

Testing only the master manifest endpoint is the most common streaming test mistake. It covers one endpoint out of three (or more) in the player request chain. Segment delivery, the CDN object that has to scale with every viewer, goes completely untested.

A manifest-only test tells you nothing about:

  • CDN segment delivery capacity
  • Segment download time vs segment duration (the rebuffer gate)
  • Origin segment packaging throughput
  • Per-rendition delivery differences

Always implement the full player loop: master manifest → variant playlist → segment loop.

Label every request so MaxoPerf can show per-label latency in the results breakdown:

requests:
- label: master-manifest # one label for manifest
url: ...
- label: variant-playlist-720p # rendition-specific labels
url: ...
- label: segment-720p # all 720p segments share this label
url: ...

Without labels, all requests collapse into one average that hides which part of the delivery chain is failing. In an unlabeled test, you cannot tell a manifest server problem from a CDN segment problem.

Don’t: ignore the segment latency vs segment duration relationship

Section titled “Don’t: ignore the segment latency vs segment duration relationship”

In a streaming test, raw millisecond latency means nothing until you compare it to segment duration:

segment download time < segment duration → viewer does not buffer
segment download time ≥ segment duration → viewer buffers

A segment p95 of 3,200ms sounds fine. If the segment duration is 4,000ms, it is fine. If the segment duration is 2,000ms, 5% of your viewers are buffering on every segment.

Always configure failure criteria with segment duration as the threshold:

# In MaxoPerf test configuration
criteria:
- p95(segment) > 4000ms for 2m: stop as failed # 4s segment duration

Do: model the full viewer population with realistic rendition mix

Section titled “Do: model the full viewer population with realistic rendition mix”

Use multiple execution scenarios weighted to your actual viewer analytics data:

execution:
- concurrency: 350 # 70% on 720p
scenario: viewer-720p
- concurrency: 100 # 20% on 1080p
scenario: viewer-1080p
- concurrency: 50 # 10% on 480p
scenario: viewer-480p

If your test only uses the highest-quality rendition, each VU consumes 5–40× more bytes than an average viewer. Either you hit bandwidth limits at a fraction of your target viewer count, or your VU count underestimates the viewer concurrency it really takes to cause failure.

Don’t: use a slow ramp for live event tests

Section titled “Don’t: use a slow ramp for live event tests”

Live event kickoff spikes need a fast ramp of 30 to 90 seconds. A 5-minute ramp warms CDN caches gradually and hides the cold-cache failure that happens when 15,000 viewers all arrive in the first minute:

# Wrong for live events — hides the kickoff failure mode
- concurrency: 10000
ramp-up: 5m
# Correct for live events
- concurrency: 10000
ramp-up: 30s # model the actual kickoff spike

Do: include manifest re-fetch in live stream scenarios

Section titled “Do: include manifest re-fetch in live stream scenarios”

For live HLS/DASH, viewers re-fetch the variant playlist every segment interval to find new segments. At 10,000 concurrent viewers with 4-second segments, that adds 2,500 manifest requests per second. This is real load on the manifest server, and your test must include it.

# After every N segments in a live scenario, re-fetch the variant playlist
- label: variant-playlist-refresh
url: https://live.example.com/stream/720p/playlist.m3u8
method: GET

Don’t: hard-code tokens or credentials in test files

Section titled “Don’t: hard-code tokens or credentials in test files”

Use MaxoPerf secrets for all auth tokens, API keys, and signing credentials. Hard-coded tokens in test files are a security risk if the files are shared or committed. They also turn credential rotation into test file edits instead of a single secrets update.

Do: run multi-region tests to catch edge PoP differences

Section titled “Do: run multi-region tests to catch edge PoP differences”

MaxoPerf distributes VUs across runner regions. Use several regions so the test reaches different CDN edge nodes:

  • A segment that is fast from EU-WEST may be slow from AP-SOUTH because that edge PoP is cold or under-sized.
  • Multi-region tests catch geographic SLA differences before real viewers report them.

Don’t: treat error rates under 1% as acceptable without inspection

Section titled “Don’t: treat error rates under 1% as acceptable without inspection”

In REST API tests, a 0.5% error rate is usually worth investigating but might be acceptable. In streaming tests, a 0.5% segment error rate means 0.5% of every viewer’s segments fail. Every viewer then sees roughly one buffering event per 800 segments (about 53 minutes at 4-second segments). Whether that is acceptable depends on your SLA. Either way, find out why before you dismiss it.

Inspect the Errors tab after every streaming run. Look for:

Error patternLikely cause
Segment 403 at scaleToken expiry, CDN auth rule, or key rotation issue
Segment 404 at scaleSegment not yet packaged, storage misconfiguration
Segment 5xxOrigin overwhelmed: autoscaling insufficient
Segment timeoutsCDN connection capacity exhausted
Manifest 429Manifest server rate-limiting concurrent sessions

Do: run a QoE-gated soak test for 24/7 linear channels

Section titled “Do: run a QoE-gated soak test for 24/7 linear channels”

Linear streaming infrastructure must carry constant load for hours or days. Latency drift (a gradual rise over time) points to memory leaks or resource exhaustion in packaging services. A short 30-minute test never shows these patterns. Run at least a 4-hour soak, with failure criteria that watch for gradual p95 drift:

criteria:
- avg-rt of segment > 3000ms for 10m: stop as failed # gradual drift gate