War-room and runbook
import Screenshot from ‘@components/Screenshot.astro’;
Load tests help little if your team has no clear, practiced plan for game day. The war-room is a literal coordination structure. It says who watches which metric, which thresholds trigger escalation, and exactly what to do when something breaks at 00:07 on Black Friday morning.
You write the war-room document before the event, rehearse it in the dress rehearsal, and follow it on game day. Writing it at midnight under pressure is too late.
The war-room structure
Section titled “The war-room structure”A BFCM war-room has three layers:
| Layer | Who | What they watch | What they act on |
|---|---|---|---|
| Frontline monitoring | On-call engineer(s) | MaxoPerf live dashboard, error rate, p95 latency | Escalate if thresholds are breached |
| Incident command | Engineering lead | Frontline escalations, business impact | Authorize abort, coordinate response |
| Business liaison | Product/ops lead | Revenue metrics, customer support volume | Communicate decisions to business stakeholders |
Setting up the MaxoPerf live dashboard
Section titled “Setting up the MaxoPerf live dashboard”On game day, the MaxoPerf Active runs view shows every running load test and its metrics live. For a BFCM war-room, set up the following before the event:
- Pin the active-runs workspace view. Open the workspace in the MaxoPerf console and bookmark the active runs page. Every engineer in the war-room keeps this URL open.
- Create the BFCM test definitions in advance. Every test you will run on game day (doorbuster monitoring, checkout journey validation, payment-path synthetic) already has failure criteria. Nobody configures anything from scratch under pressure.
- Set up the comparison baseline. Pin last year’s dress-rehearsal run (or the T-1 week dress rehearsal) as the comparison baseline. On game day, you can compare any new run to the baseline in one click.
The monitoring assignment matrix
Section titled “The monitoring assignment matrix”Assign each metric and threshold to a named person. When “everyone watches everything”, nobody catches a problem fast enough.
| Engineer | Primary metric | Amber threshold | Red threshold |
|---|---|---|---|
| SRE lead | Overall error rate | > 0.5 % | > 2 % |
| Backend engineer 1 | p95 latency (checkout flow) | > 600 ms | > 1200 ms |
| Backend engineer 2 | p95 latency (payment endpoint) | > 800 ms | > 2000 ms |
| Infra engineer | Database connection pool utilization | > 70 % | > 90 % |
| Frontend engineer | CDN cache hit rate | < 85 % | < 70 % |
| DBA | Slow query rate | > 5/min | > 20/min |
Each engineer checks their metric every 5 minutes and reports status in the war-room channel: green / amber / red.
Abort criteria
Section titled “Abort criteria”Decide before the event which conditions trigger an action: reduce traffic, roll back, or abort the sale. Under pressure, with revenue numbers in front of you, people make the wrong call.
| Condition | Duration | Action |
|---|---|---|
| Error rate > 2 % | Sustained 3 min | Amber: engineering escalation, investigate |
| Error rate > 5 % | Sustained 2 min | Red: incident command decision required |
| Error rate > 10 % | Any | Immediate rollback / traffic shed |
| p95 payment > 3000 ms | Sustained 5 min | Throttle checkout entry, open incident |
| Payment service 503s | Any | Circuit break to fallback; incident command |
| Database connection pool > 95 % | 1 min | Immediate action; possible pod restart |
| Complete checkout unavailability | 30 s | Activate maintenance page; full incident |
Write these criteria in the runbook. The engineering lead and the business lead review and sign them off at least one week before the event.
The runbook: action procedures
Section titled “The runbook: action procedures”For each abort condition above, the runbook lists the exact commands or console actions. Nobody should be searching documentation at midnight on Black Friday.
Example runbook entry: high error rate (amber)
Section titled “Example runbook entry: high error rate (amber)”Condition: Overall error rate > 2 % for 3 minutes
Immediate actions (frontline engineer):
- Open MaxoPerf run detail → Overview tab → Log/Errors tab.
- Note the specific error type (HTTP 5xx? timeout? specific endpoint?).
- Post in war-room channel:
[AMBER] Error rate 2.4%. Endpoint: POST /checkout/payment. Error: upstream timeout. Investigating. - Check payment gateway dashboard for upstream issues.
- Check infrastructure dashboard for pod health, CPU, memory.
Escalation: If error type is not identified within 5 minutes, escalate to incident command.
Resolution: When error rate drops below 1 % for 2 minutes, post: [GREEN] Error rate normalized. Cause: [description]. Monitoring continued.
Example runbook entry: complete checkout unavailability (red)
Section titled “Example runbook entry: complete checkout unavailability (red)”Condition: Checkout unavailable for > 30 seconds
Immediate actions (any engineer):
- Post in war-room channel:
[RED] Checkout DOWN. T=00:07. Activating maintenance page. - Execute maintenance page activation:
kubectl apply -f deploy/maintenance-mode/checkout-maintenance.yaml - Notify incident command.
- Begin incident response in dedicated incident channel.
Rehearsing the war-room
Section titled “Rehearsing the war-room”The dress rehearsal at T-1 week is also a war-room exercise. Run it with the full war-room structure in place:
- All engineers in their monitoring seats
- The war-room channel open
- The runbook in hand
- Incident command role staffed
Introduce amber conditions on purpose during the rehearsal (for example, raise the error rate on the mock payment server for 5 minutes) and practice the escalation flow. On game day, every action should be something the team has already done once.
Post-event debrief
Section titled “Post-event debrief”After the event, within 48 hours:
- Run the year-over-year comparison in MaxoPerf: this year’s peak run vs. last year’s.
- Document all amber and red events in the incident log.
- Update the runbook with any actions that were taken but not documented.
- Archive the game-day run in MaxoPerf with the tag
bfcm-2025-game-dayfor next year’s baseline.
Where to go next
Section titled “Where to go next”- BFCM readiness checklist: war-room setup is a required pre-event checklist item.
- Failure criteria pass/fail gates: the failure criteria that drive war-room alerts.
- Comparing runs and baselines: set up the year-over-year comparison during the debrief.
- Daily countdown calendar: when, in the T-6 week program, you write and rehearse the war-room runbook.