SSA "My Social Security" portal crashes — the cost of skipping a load test before launch
In early 2025, the Social Security Administration portal crashed repeatedly after a software update. Reporting attributed it to staff not testing the software against high user volumes before launch.
In early 2025, the Social Security Administration’s “My Social Security” portal crashed roughly four to five times within a few weeks. The portal serves around 70 million beneficiaries. Each crash locked users out of benefit statements, payment histories, and direct deposit settings, services many Americans rely on for basic financial planning.
Load testers return to this case when they explain why you test before you ship, not after. It also has a named cause, cited in reporting: technology staff did not test the software against a high volume of users to see if the servers could handle the rush.
What happened
In March 2025, the SSA portal started crashing again and again. Newsweek reported that the Social Security Administration confirmed it was responding to the crashes and working to stabilise the site. Gizmodo’s reporting tied the crashes to a recently deployed software update. The site went down several times over a few weeks, and tens of millions of users could not reach their accounts.
The crashes followed a pattern. Each one lined up with a surge of users trying to log in: a predictable thundering herd against a newly modified authentication path.
The timeline
- Software update deployed: A new anti-fraud check is introduced earlier in the authentication flow, meaning a larger share of login attempts now reach the fraud-check step.
- Crashes begin (March 2025): The portal starts experiencing repeated outages; the SSA confirms it is responding to crashes.
- Mass login surge: Anxious beneficiaries, many checking on potential benefit changes, flood the login page at the same time and amplify load on the new auth path.
- Four to five crashes over weeks: Reporting documents multiple failures across the period.
- Attribution: The Washington Post, cited via Fortune, reported that technology staff had not tested the software against a high volume of users before launch.
Why it happened
The proximate cause, as attributed in reporting, was that nobody tested the software against high user volumes before it went live.
In more detail, the update put a new anti-fraud check into the login path. Before, this check ran later in the flow, or not at all for many sessions. Moving it earlier meant that every login attempt now triggered it. The fraud-check step was presumably sized for its original, narrower role. Suddenly it had to handle the full concurrent login volume of a site with tens of millions of active users.
Anxious beneficiaries, many reacting to news about possible changes to their benefits, piled onto the site at once. The result was a textbook thundering herd against an undersized authentication service. The system collapsed under high legitimate demand on an auth path nobody had validated at that concurrency.
The Gizmodo report tied the crashes directly to the software update. The Fortune piece gave the clearest public statement of the testing gap.
The failure pattern
This is a thundering herd at a new authentication path. It is a variant of the broader thundering-herd class, and an especially dangerous one because it builds up unseen during development. The auth path existed before the update. The update changed which users hit it and when. Without a load test of the modified path at realistic concurrency, the team could not see that change until production collapsed under it.
The general rule for any team: if a change moves load to an earlier point in a user flow, or puts a previously lightweight service under full concurrency, load test it as if it were new infrastructure. From a load perspective, it is.
How it could have been prevented
Load test the modified path before launch, not after. The missing step is on record: nobody tested the software against high user volumes. A load test of the anti-fraud-check path at the expected concurrent login volume, run before deployment, would have revealed the capacity gap.
Test the specific path that changed, not only the overall site. The auth and fraud-check step is a distinct backend flow. Give it its own load profile, separate from a general site load test. Target the login endpoint directly and ramp virtual users to your production concurrent session count.
Define a failure criterion before the test runs. Decide ahead of time what p95 response time and error rate you accept at peak concurrency. If the test exceeds those thresholds, the release waits until the fix is validated.
Pre-scale the new dependency. If the fraud-check service was sized for its previous role, resize it for full login-path concurrency before you promote it into that role. Load testing tells you the capacity you need. Provisioning it is a pre-launch task.
Run a spike to simulate the mass-login surge. Beneficiaries arrived in a concentrated burst. A spike test models that better than a gradual ramp, because you need to know what happens when login attempts arrive together, not one after another.
How to test for this with MaxoPerf
The test type that matches this failure is a load test on the auth path itself, combined with a spike test that models the mass simultaneous login.
Engine: k6 or JMeter, targeting the login and session-creation endpoints.
Profile, load test (baseline validation):
- Ramp from 0 to expected peak concurrent sessions over 3 minutes
- Hold at peak for 15 minutes
- Watch for the onset of errors and for p95 response time degradation
Profile, spike (thundering herd simulation):
- Baseline: 500 virtual users
- Near-instant ramp to 10,000–50,000 VUs over 30 seconds (model the simultaneous login wave)
- Hold at peak for 5 minutes
- Ramp down over 2 minutes
Set the VU count from the concurrent session data for the busiest period you have seen in production. For a site with tens of millions of registered users, the simultaneous login peak during a high-anxiety news event could be orders of magnitude above normal.
Target: a staging environment with the modified auth path deployed, including the new anti-fraud check. Do not test the old path and call it done. Test the changed path at production-like concurrency.
Execution locations: MaxoPerf managed regions. Run from several locations to spread login source IPs and avoid tripping rate limits on the test itself.
Signals to watch in results:
- Login endpoint error rate, with a target of zero at expected peak
- p95 and p99 response time for the auth round-trip
- Virtual user throughput at the spike peak: does it hold, or does the system start rejecting connections?
- Downstream dependency response time (if your fraud check calls an external or internal service, add that service to your monitoring)
Gate the release. If the load test shows degradation above your threshold at expected peak, treat it as a release blocker. A clean result at realistic concurrency is the criterion for going live. Developer confidence and a single manual test are not.
Key takeaways
- A software update that changes which users hit a backend service is a load event as well as a code change. Test it that way.
- The SSA case is the clearest public example of skipping load testing: the system works in development and collapses in production on the first day of real concurrent load.
- Test the changed path specifically. A general site load test would not have exercised the anti-fraud-check path at full concurrency before the update.
- Define failure criteria before testing, not after, or the results tell you nothing you can act on.
- The fix for a missing load test is simple: run the test against your own staging environment before the next change ships.
The load test guide, the spike test guide, and the load failures series cover how to build a pre-launch testing practice that catches this class of failure before it reaches production.
Questions this article answers
Why did the Social Security Administration website crash in 2025?
A software update moved a new anti-fraud check earlier in the login flow, causing many more users to hit the authentication path simultaneously. Reporting attributed the crashes to technology staff not testing the software against a high volume of users before launch.
How do you load test a new authentication flow before launching it?
Run a load test on the specific auth path at the concurrent user volume you expect in production, before the feature goes live. Measure where error rates and response times degrade and fix capacity gaps before users encounter them.
What is a thundering herd in the context of web systems?
A thundering herd occurs when a large number of users or clients all attempt the same operation simultaneously, overwhelming a shared resource such as an authentication service or database connection pool.