Skip to main content

Incident Commander Simulator

Take the page as incident commander. Direct two responders, read the evidence they bring back, choose a mitigation, keep customers updated, and confirm the recovery holds before you resolve. Turn-based, with a review of every decision at the end.

Category: SRE

What You Will Learn

  • Run an incident as the incident commander: coordinate, do not fix everything yourself
  • Mitigate first with the smallest reversible change, then find the root cause
  • Send investigations to responders and read the evidence they bring back
  • Post status updates early and on a schedule, and keep them accurate
  • Confirm the recovery holds before you resolve an incident
  • Recognise database connection exhaustion and a bad feature flag rollout

Topics covered: incident-management, incident-response, sre, on-call, postgresql, feature-flags, interactive, educational

// simulator

Incident Commander Simulator

Take the page as incident commander. Direct two responders, read the evidence they bring back, choose a mitigation, keep customers updated, and confirm the recovery holds before you resolve. Turn-based, with a review of every decision at the end.

Supported byBecome a sponsor
SvixAtomsizedDigitalOceanDevDojoSMTPfastQuizAPI

You are the incident commander

You do not fix things with your own hands. You direct a small team, decide what to try, keep customers informed, and decide when the incident is really over.

  • Send work to your teamEach responder takes one task at a time. Investigations answer one question each.
  • Reading is freeTime only moves when you advance the clock. Take the time to read the evidence.
  • Verify before you closeRecovery has to hold for 8 minutes before you can resolve.
Launch day: too many clients
Easy

A product launch, a scaled-up API, and checkout failures within minutes.

Team: Priya (App on-call) and Marcus (Database on-call)

Three changes, one incident
Medium

A deploy, a feature flag and a node upgrade in five minutes. Then checkout breaks.

Team: Priya (Checkout on-call) and Dev (Platform on-call)

About this incident commander simulator

What you'll practise

  • Directing two responders instead of doing every task yourself
  • Turning a vague alert into specific questions, and judging the answers
  • Choosing a mitigation by scope, risk and how easily you can undo it
  • Changing one thing at a time, so each step tells you something
  • Writing status updates that are early, regular and true
  • Telling "the graph turned green once" apart from a recovery that holds

How it works

  • Turn-based: time only moves when you advance the clock, so reading the evidence costs nothing.
  • Your team: each responder works on one task at a time. Investigations and changes take simulated minutes.
  • The review: at the end you see the customer impact, how you communicated, any risky changes, and one decision worth revisiting.

Mitigate first, diagnose second

During an incident the first job is to stop the damage. A feature flag turned off, a rollback, or a smaller connection pool can restore service long before anyone understands the root cause. Pick the change that matches the evidence and is fastest and safest to undo, then investigate calmly once customers are fine.

Read more

The scenarios come from our post on connection pools and our guide to progressive delivery with feature flags. For the on-call side, read how to build an on-call rotation and escalation policy, and for a game about keeping services up under load, try Uptime Defender.

Try next

Sponsored
Carbon Ads
$ cd /games
// share