Incident Command / Field Kit

You run the incident; the incident doesn't run you. A calm, offline command flow: what to do at each phase, what to ask every domain when they join the bridge and what to look for, comms you can paste, and a timeline that becomes your status update and postmortem. Nothing here calls the network.

Be like water — flexible. Understand the tools; don't marry the methods.

One incident commander. The IC coordinates and decides — the IC does not put hands on the keyboard. Work the phases in order, but skip ahead if the incident tells you to. Stabilize before you diagnose.

🧭 Steady first. Take one breath before the first action. Say the three things out loud — what's the impact, what's the next action, who owns it — and the room gets calmer with you. You don't have to have the answer yet; you just have to run the next step. You've done hard things before. You've got this.

Size it first

SEV1 · Critical

Major outage, data-loss risk, or safety/regulatory exposure. All-hands, exec comms, 15-min cadence.

SEV2 · High

Significant degradation or partial outage. Dedicated bridge, 30-min cadence.

SEV3 · Moderate

Limited impact, workaround exists. Normal hours, hourly/as-needed updates.

SEV4 · Low

Negligible / cosmetic. Track and fix; no bridge needed.

01 Declare & assign
  • Decide: is this an incident? If in doubt, declare — you can always downgrade.
  • Set the severity and open the bridge + incident doc.
  • Name the Incident Commander out loud. One IC. Add a Scribe/Comms and Ops lead(s).
  • State the current impact and the time it started.

IC mantra: coordinate, don't fix. Your job is the whole board, not one square.

02 Assemble the right people
  • Page only the domains you need — a crowded bridge is noise. Add more as hypotheses point there.
  • Give each joiner a 10-second brief: impact, what you know, what you need from them.
  • Assign explicit owners: "You own the DB angle. You own comms."

Ask everyone: what changed recently in your area?

03 Stabilize — stop the bleeding
  • Ask the fastest-mitigation question first: can we restore service now, before we know why?
  • Reach for the reversible levers: rollback, failover, feature flag, scale out, drain, reroute.
  • Prefer mitigation over a perfect fix. Customers feel recovery, not root cause.

Rule: restore first, understand second. Root cause can wait; the outage can't.

Breathe — the reversible move now beats the perfect move in an hour. It's okay to mitigate before you fully understand. Pick the safe lever and pull it.

04 Coordinate the diagnosis
  • One conversation at a time. Track hypotheses and who's checking what, with a timebox.
  • Rule things in or out explicitly — "DB ruled out as of 14:20."
  • When a domain lead comes back, use the Domain check-ins tab to know what to ask.
  • Deep on the database? Hand them the DB Field Kit for triage.

Ask on repeat: what's the customer impact right now, and what's our fastest path out?

05 Communicate on a cadence
  • Send updates on a fixed cadence even when there's no change — silence reads as chaos.
  • Separate internal (technical) from stakeholder (impact + ETA) messaging.
  • Use the Comms tab for paste-ready templates.

Breathe — a quick "still working, next update at 2:45" keeps everyone calm, including you. Pause, post it, come back. Telling the team you're running checks is not a delay — it's the job.

06 Resolve & verify
  • Confirm recovery with data, not hope — the graphs and health checks, not one person's "looks fine."
  • Watch for secondary failures (queues draining, caches cold, retries stampeding).
  • Declare resolved, drop cadence, thank the team.
07 Hand off / close out
  • Long-running? Do a clean handoff: current state, what's been tried, next steps, open owners.
  • Capture the timeline (Log tab) while it's fresh.
  • Schedule a blameless review. The system failed, not the person.

Incident Commander Field Kit — a ReliabilityOps tool. Offline, single-file, read-only; adapt it to your organization's severity model, roles, and comms policy. Nothing on this page calls the network; your timeline never leaves the browser.