Slack outage May 2025 — when infrastructure growth outpaces static configuration
In May 2025, Slack's growth hit a silent limit in its routing configuration, cutting live connectivity for users for nearly two hours. This post explains what the pattern means for any growing team.
Most capacity failures give warning. Traffic climbs, latency rises, and engineers see some signal that the system is nearing a limit. A subtler class of failure puts the limit outside the application code, in a static configuration that was correct when someone set it and has quietly turned wrong as the system grew.
In May 2025, Slack hit this pattern. Postmortem aggregators report the incident clearly enough to show how infrastructure growth can create invisible ceilings, and why load testing at realistic scale is the most reliable way to find them.
Note: this incident occurred in May 2025, just before this review’s twelve-month window. It is included because the failure pattern applies to any team running a growing SaaS platform, and because the underlying dynamic is still in play today.
What happened
On May 12, 2025, Slack had a disruption that lasted about one hour and 58 minutes. According to postmortem reporting, the cause was: “Infrastructure growth had outpaced static configurations, preventing routing updates from reaching the web layer.” With routing updates unable to reach the web layer, clients lost live connectivity data. They worked from stale state instead of fresh routing information. A significant portion of users saw elevated errors and degraded connectivity.
The incident ended once the routing configuration was updated to match the current infrastructure’s scale. No data was lost. The impact was disrupted connections and messaging.
The timeline
- May 12, 2025, onset: Infrastructure operating at a scale that has grown beyond what static routing configuration can describe.
- Routing propagation fails: Routing updates cannot reach the web layer; the configuration ceiling is hit.
- Clients lose live routing data: Web clients receive stale or incomplete routing state; live connectivity degrades.
- Elevated errors: Users experience disconnections, message delivery failures, and degraded real-time features.
- ~1h58m later: Routing configuration is updated; propagation resumes; service recovers.
Why it happened
The direct cause was a static configuration that defined how routing information could be described and propagated. That configuration carried a size or count limit set when the infrastructure was smaller. The infrastructure grew (more hosts, more routes, more services), and the configuration stayed the same. At some point the infrastructure exceeded what the configuration could represent.
Confidence in this account is medium. Postmortem aggregators report the routing-configuration cause clearly, but the exact nature of the static limit is not public. The pattern is reliable even though the technical boundaries are not all known. The core dynamic holds: a hard static limit stays invisible during growth and becomes a ceiling nobody finds until it is crossed.
The failure pattern
This is a capacity-ceiling failure of a specific sub-type: a silent configuration limit. With an infrastructure capacity ceiling, adding servers buys headroom. Scaling compute cannot fix a silent configuration limit. You have to find the static value, understand why it exists, and update it, and you cannot do that in real time once users are already affected.
Silent configuration limits tend to share a few dangerous properties:
- They give no warning. No metric says “you are at 80% of this limit.”
- They are often set once and forgotten. Someone set them when the system was much smaller and the limit looked generous.
- When they are crossed, the failure is hard to read because it does not look like resource exhaustion. It looks like a routing or networking problem.
If your infrastructure has grown a lot since its initial configuration, treat this pattern as an active risk.
How it could have been prevented
Audit static configurations for hidden limits. Review routing tables, configuration files, and infrastructure manifests that define the shape or size of the system on a regular schedule as the infrastructure grows. A limit that made sense at 100 nodes may be dangerously close at 1,000.
Test at projected fleet sizes, not current sizes. A scalability test at 1.5× or 2× the current infrastructure size reveals compute limits, and it also exercises the paths where routing and configuration data must describe a larger system. Configuration limits surface there.
Add limit-proximity alerting. Where you can quantify a configuration limit (e.g., maximum route table size, maximum entry count), set monitoring to alert when the system reaches 70–80% of it. You want to discover the limit before you reach the ceiling, not when you hit it.
Define and track an SLO. An SLO turns abstract reliability goals into a measurable target. Once you know your target, say 99.9% monthly availability, you can calculate how many minutes of outage remain in your error budget. That number gives you a reason to find and fix hidden limits before they consume the budget.
How to test for this with MaxoPerf
You want to find configuration ceilings before they find you, by testing at the scale your infrastructure will reach instead of the scale it runs at today.
Run a scalability test at projected fleet sizes. If you expect your infrastructure to double in the next two quarters, set up a scalability test that steps through 25%, 50%, 75%, 100%, 150%, and 200% of your target scale. Use MaxoPerf with k6 or Taurus to simulate the web-layer load that would come with each fleet size. An example profile:
- 5 virtual users for 2 minutes (baseline).
- Step to 25, 50, 100, 200, 400, and 800 VUs, holding each for 3 minutes.
- Measure p95 latency, error rate, and throughput at each step.
Watch for non-linear degradation. A healthy, well-configured system degrades roughly linearly: throughput climbs as VUs climb, and latency holds steady until real saturation. A silent configuration limit looks different. The system is fine up to a specific VU count, and then latency or errors jump all at once. That jump is the limit.
Pair it with a breakpoint test at your traffic growth horizon. A breakpoint test that starts at today’s peak and steps up to 3× or 5× current traffic exercises routing and configuration paths at scales you have never run. Note any sudden change in behavior that does not line up with compute metrics.
Use several managed regions to simulate realistic geographic traffic. Slack’s routing disruption involved how routing updates reach the web layer. A test from one point of origin may miss a propagation failure that geography would reveal. MaxoPerf’s managed cloud locations let you generate load from several regions at once, which comes closer to the distribution a real production failure would need to survive.
After each step increase, review the run results for latency and error rate, and also for any change in behavior that looks abrupt instead of gradual.
Key takeaways
- Static configuration limits stay invisible until crossed. Compute capacity has metrics. A configuration ceiling often has no metric that warns you it is close.
- Growth creates the risk silently. A configuration that was safe at 100 nodes may be dangerous at 1,000. The limit stays the same while the system changes.
- Test at tomorrow’s scale, not today’s. Scalability and breakpoint tests at projected fleet sizes are the most reliable way to find configuration ceilings before users do.
- Non-linear degradation in a test is a signal. If throughput or error rate jumps at a specific load level instead of degrading gradually, suspect a hard limit, not only resource saturation.
- An SLO makes the risk concrete. Knowing your availability target turns a vague capacity worry into a quantified risk you can prioritize alongside feature work.
Questions this article answers
Why did Slack go down in May 2025?
Infrastructure growth had outpaced static routing configuration, preventing routing updates from reaching the web layer. Clients lost live connectivity data, causing elevated errors for approximately two hours.
What is a silent configuration limit and how do I find mine?
A silent configuration limit is a hard-coded or statically defined ceiling in infrastructure that has no warning before it is reached. Scalability and breakpoint tests at projected fleet sizes can surface these limits before they cause production outages.
How does an SLO help teams prioritize capacity headroom?
An SLO defines the acceptable availability or latency target. When you know your target, you can quantify how much headroom a configuration limit eats into your error budget, and prioritize fixing it before it is reached.