ARC Raiders passed its pre-launch stress test — then overloaded anyway. Here's the soak lesson.
ARC Raiders ran clean for days after launch, then login queues and matchmaking broke as concurrent players kept climbing. Reporting attributes this to sustained growth beyond a one-time peak. Embark did not confirm a root cause.
ARC Raiders launched on 30 October 2025 and reached over 300,000 concurrent players on Steam, a strong debut for a new extraction shooter from Embark Studios. Before launch, the team ran a “server slam” stress test. By all accounts, it passed.
Several days into launch, things changed. Reporting attributes the timing to sustained concurrency growth. The player count kept climbing instead of tapering off, login queues grew, matchmaking slowed, and party and voice services broke. Players found themselves stuck in queues or unable to form groups.
Embark Studios did not publicly confirm the specific root cause of these issues. Reporting describes a system that held at launch and then degraded as concurrency kept growing. A short pre-launch stress test is built to catch a different failure mode.
What happened
ARC Raiders launched on 30 October 2025. The pre-launch stress test cleared the initial spike, and the system absorbed the opening surge. For a while after launch, players got in. Then, around 2 November 2025, reports of login queues and matchmaking failures began to circulate. Reporting at the time places concurrent player counts above 330,000 when these issues surfaced.
Player reports describe a gradual build-up of problems, not a sudden crash. Login queues grew longer over time. Matchmaking became less reliable. Party formation and voice chat degraded. A system under sustained load fails like this. A system that hits a hard limit all at once fails differently.
The timeline
- 2025-10-30: ARC Raiders launches. The pre-launch server slam test had passed. The initial player influx is absorbed.
- Days post-launch: The concurrent player count keeps growing instead of settling at a steady state.
- ~2025-11-02: Login queues, matchmaking failures, and party/voice issues emerge. Reporting ties the timing to sustained concurrency growth beyond the one-time peak the slam test modelled.
- Recovery: Embark works on server capacity, and the issues ease as the studio responds. No detailed public postmortem has been released.
Why it happened
Reporting attributes the failure to sustained concurrency beyond what a one-time stress test modelled. That load exposed degradation the shorter test could not observe. Embark did not confirm a root cause publicly, so this section follows the pattern reporting describes, not a confirmed technical cause.
A pre-launch server slam test, like any short stress test, checks whether a system survives a rapid load spike. It answers “Can we absorb the opening surge?” It does not answer “What happens if we stay at or above that load for the next 48 hours?”
The second answer is often different. Under long periods of high concurrency, degradation builds up slowly:
- Connection pools to databases, matchmaking services or session stores may hold steady at peak for 15 minutes. Then they saturate as long-lived connections pile up without being recycled.
- Memory use on stateful services (session servers, party services, voice routing) can creep upward until a host nears its limit and starts dropping requests.
- Queue depths grow when a downstream dependency, such as a session store or a matchmaking cluster, falls behind. Each player’s login flow gets slower.
A 10–15 minute slam test ends before any of these patterns show up. A multi-hour soak test at the same concurrency level will surface all of them.
The failure pattern
This is soak-growth-collapse. A system holds at launch and through early peak, then degrades as concurrency keeps climbing or as time at elevated load runs past what was tested. “Collapse” is relative. It can mean an error rate that creeps up until the player experience breaks, not an all-or-nothing outage.
The pattern is common in games with stateful session and matchmaking layers. Each active session holds resources on the server. A short test does not cycle enough sessions to show pool saturation or memory drift. You need a test that sustains load for hours, through many session opens, matches and closes, to learn whether resource management holds up over time.
The soak / endurance test type is designed for exactly this scenario.
How it could have been prevented
Running a pre-launch slam test was the right call. The gap was in the question it was asked to answer.
Pair the slam test with a soak test. Follow the slam test with a multi-hour soak at the projected sustained peak, not only the top of the spike. If you expect the game to hold 200,000 concurrent players for twelve hours on launch day, run the soak at 200,000 virtual users for at least six to eight hours in staging.
Add below-p50 resource monitoring. Error rates and p99 latency are lagging indicators. Memory per session on your session servers, connection-pool use on your matchmaking cluster and queue depth on stateful services are leading indicators. They move before players see errors. If your instrumentation shows them during the soak, you catch the build-up early.
Model the full session lifecycle, not only the login. A test that only sends login requests skips most of the server-side session: match formation, in-session keep-alive, disconnect and reconnect. Session servers need the full cycle before memory and connection patterns build up.
Stage capacity increases ahead of predicted growth. Player growth can outrun your forecast (a good problem to have). Pre-provisioned capacity tiers, plus a plan to switch them on at set concurrency thresholds, shorten the lag between demand and capacity.
How to test for this with MaxoPerf
Use a soak / endurance test that runs for hours, not a spike that ends in minutes.
Profile shape
Use k6 or Taurus with an open-model workload:
- Ramp-up: 30 minutes from zero to your projected sustained peak concurrency. At ARC Raiders’ scale, that’s 150,000–200,000 virtual users. For your game, size it to your realistic peak.
- Hold: 4–8 hours at that concurrency. The hold is the part that finds the problem, so don’t cut it short.
- Observation window: optionally ramp to 120–150% of projected peak in the final hour. This models the “kept growing” scenario that reporting links to ARC Raiders’ post-launch degradation.
Target the stateful path
Point the test at the full session lifecycle in your staging environment: login → matchmaking request → party formation → session join → periodic keep-alive → disconnect. Each step holds server-side resources. If you test only the login step, you miss everything that builds up in the matchmaking and session layers behind it.
What to watch in results
In MaxoPerf’s results view and metrics explorer:
- Error rate over time: flat for the first 30 minutes, then slowly climbing. That is the soak pattern.
- p95 and p99 latency trend: latency that grows while concurrency stays stable points to resource build-up.
- Throughput plateau: if requests per second stop growing while you are still adding virtual users, something upstream is saturating.
Set failure criteria that trip if the error rate crosses 1% or if p99 latency exceeds your threshold at any point in the hold phase. A clean spike result next to a failing soak result tells you to check resource management before players find the problem.
Scheduling
Use MaxoPerf’s scheduled runs to run the soak test automatically before every major patch that touches session, matchmaking or party services. A change that leaves peak performance alone can still add a resource leak that only appears after hours at load.
Key takeaways
- A pre-launch stress test checks whether a system survives a spike. It does not check whether the system stays healthy under sustained load for hours. Answer both questions before a large launch.
- Reporting attributes ARC Raiders’ post-launch issues to sustained concurrency growth beyond what a one-time test modelled. Embark did not confirm a root cause publicly. The lesson is the pattern, not a confirmed engineering failure at Embark.
- Soak-growth-collapse is gradual. Error rates creep, latencies drift and queues grow. Multi-hour load tests show these signals long before players notice.
- Model the session lifecycle in the test (login, match, keep-alive, disconnect, reconnect) to expose connection-pool and memory-drift patterns. Login-only tests underestimate stateful resource pressure.
- Run soak tests as a regular gate for patches that touch session and matchmaking infrastructure, not only once before launch.
If you are preparing a multiplayer game launch or a major patch to stateful backend services, MaxoPerf’s soak / endurance test type gives you the profile and observability to answer the question a slam test leaves open.
Questions this article answers
Why did ARC Raiders have server problems after a successful pre-launch test?
Reporting attributes the issues to sustained concurrent growth beyond what a short one-time stress test could reveal. A pre-launch slam test checks whether the system survives a sudden peak. It does not expose gradual degradation (leaked resources, exhausted connection pools, slow memory growth) that appears only after hours at elevated concurrency. Embark Studios did not publicly confirm the specific root cause.
What is a soak test and why does it catch failures that a spike test misses?
A soak test, also called an endurance test, holds a target at sustained concurrency for hours rather than minutes. Over that duration, resource leaks, connection-pool exhaustion and configuration limits build up under continuous load. A short spike test never sees them because it ends before the degradation accumulates.
How long should a soak test run for a game launch?
A useful soak test runs for at least the expected continuous-play window of a typical launch day, often four to eight hours. The goal is to see whether error rates and latency stay stable while the system holds at or above projected peak concurrency for as long as real players will use it.