Incident Commander Simulator
Take the page as incident commander. Direct two responders, read the evidence they bring back, choose a mitigation, keep customers updated, and confirm the recovery holds before you resolve. Turn-based, with a review of every decision at the end.
Category: SRE
What You Will Learn
- Run an incident as the incident commander: coordinate, do not fix everything yourself
- Mitigate first with the smallest reversible change, then find the root cause
- Send investigations to responders and read the evidence they bring back
- Post status updates early and on a schedule, and keep them accurate
- Confirm the recovery holds before you resolve an incident
- Recognise database connection exhaustion and a bad feature flag rollout
Topics covered: incident-management, incident-response, sre, on-call, postgresql, feature-flags, interactive, educational
// simulator
Incident Commander Simulator
Take the page as incident commander. Direct two responders, read the evidence they bring back, choose a mitigation, keep customers updated, and confirm the recovery holds before you resolve. Turn-based, with a review of every decision at the end.
You are the incident commander
You do not fix things with your own hands. You direct a small team, decide what to try, keep customers informed, and decide when the incident is really over.
- Send work to your teamEach responder takes one task at a time. Investigations answer one question each.
- Reading is freeTime only moves when you advance the clock. Take the time to read the evidence.
- Verify before you closeRecovery has to hold for 8 minutes before you can resolve.
A product launch, a scaled-up API, and checkout failures within minutes.
Team: Priya (App on-call) and Marcus (Database on-call)
A deploy, a feature flag and a node upgrade in five minutes. Then checkout breaks.
Team: Priya (Checkout on-call) and Dev (Platform on-call)
About this incident commander simulator
What you'll practise
- Directing two responders instead of doing every task yourself
- Turning a vague alert into specific questions, and judging the answers
- Choosing a mitigation by scope, risk and how easily you can undo it
- Changing one thing at a time, so each step tells you something
- Writing status updates that are early, regular and true
- Telling "the graph turned green once" apart from a recovery that holds
How it works
- Turn-based: time only moves when you advance the clock, so reading the evidence costs nothing.
- Your team: each responder works on one task at a time. Investigations and changes take simulated minutes.
- The review: at the end you see the customer impact, how you communicated, any risky changes, and one decision worth revisiting.
Mitigate first, diagnose second
During an incident the first job is to stop the damage. A feature flag turned off, a rollback, or a smaller connection pool can restore service long before anyone understands the root cause. Pick the change that matches the evidence and is fastest and safest to undo, then investigate calmly once customers are fine.
Read more
The scenarios come from our post on connection pools and our guide to progressive delivery with feature flags. For the on-call side, read how to build an on-call rotation and escalation policy, and for a game about keeping services up under load, try Uptime Defender.
Try next
// game
Uptime Defender
Fast-paced SRE game where you defend your infrastructure uptime! Handle incoming incidents by adding nodes, rotating logs, failing over, restarting pods, and scaling databases.
// simulator
Preview Environment Simulator
See how a pull request becomes a temporary copy of your app, complete with a private URL, safe test data, review checks, and automatic cleanup.
// game
Heroku-Style Name Generator
Generate names the way Heroku, Docker, Kubernetes, and petname do, see how the word lists combine, and learn when the birthday problem makes random names collide.