You run the incident; the incident doesn't run you. A calm, offline command flow: what to do at each phase, what to ask every domain when they join the bridge and what to look for, comms you can paste, and a timeline that becomes your status update and postmortem. Nothing here calls the network.
Be like water — flexible. Understand the tools; don't marry the methods.
One incident commander. The IC coordinates and decides — the IC does not put hands on the keyboard. Work the phases in order, but skip ahead if the incident tells you to. Stabilize before you diagnose.
🧭 Steady first. Take one breath before the first action. Say the three things out loud — what's the impact, what's the next action, who owns it — and the room gets calmer with you. You don't have to have the answer yet; you just have to run the next step. You've done hard things before. You've got this.
Size it first
SEV1 · Critical
Major outage, data-loss risk, or safety/regulatory exposure. All-hands, exec comms, 15-min cadence.
SEV2 · High
Significant degradation or partial outage. Dedicated bridge, 30-min cadence.
SEV3 · Moderate
Limited impact, workaround exists. Normal hours, hourly/as-needed updates.
SEV4 · Low
Negligible / cosmetic. Track and fix; no bridge needed.
01 Declare & assign
Decide: is this an incident? If in doubt, declare — you can always downgrade.
Set the severity and open the bridge + incident doc.
Name the Incident Commander out loud. One IC. Add a Scribe/Comms and Ops lead(s).
State the current impact and the time it started.
IC mantra: coordinate, don't fix. Your job is the whole board, not one square.
02 Assemble the right people
Page only the domains you need — a crowded bridge is noise. Add more as hypotheses point there.
Give each joiner a 10-second brief: impact, what you know, what you need from them.
Assign explicit owners: "You own the DB angle. You own comms."
Ask everyone: what changed recently in your area?
03 Stabilize — stop the bleeding
Ask the fastest-mitigation question first: can we restore service now, before we know why?
Reach for the reversible levers: rollback, failover, feature flag, scale out, drain, reroute.
Prefer mitigation over a perfect fix. Customers feel recovery, not root cause.
Rule: restore first, understand second. Root cause can wait; the outage can't.
Breathe — the reversible move now beats the perfect move in an hour. It's okay to mitigate before you fully understand. Pick the safe lever and pull it.
04 Coordinate the diagnosis
One conversation at a time. Track hypotheses and who's checking what, with a timebox.
Rule things in or out explicitly — "DB ruled out as of 14:20."
When a domain lead comes back, use the Domain check-ins tab to know what to ask.
Deep on the database? Hand them the DB Field Kit for triage.
Ask on repeat: what's the customer impact right now, and what's our fastest path out?
05 Communicate on a cadence
Send updates on a fixed cadence even when there's no change — silence reads as chaos.
Separate internal (technical) from stakeholder (impact + ETA) messaging.
Use the Comms tab for paste-ready templates.
Breathe — a quick "still working, next update at 2:45" keeps everyone calm, including you. Pause, post it, come back. Telling the team you're running checks is not a delay — it's the job.
06 Resolve & verify
Confirm recovery with data, not hope — the graphs and health checks, not one person's "looks fine."
Watch for secondary failures (queues draining, caches cold, retries stampeding).
Declare resolved, drop cadence, thank the team.
07 Hand off / close out
Long-running? Do a clean handoff: current state, what's been tried, next steps, open owners.
Capture the timeline (Log tab) while it's fresh.
Schedule a blameless review. The system failed, not the person.
When a domain rep joins — or comes back with findings — ask the same three things: what did you find, what did you change, is your area ruled in or out (and how confident)? Below: the specific questions and signals per domain.
🗄️ Database
Ask
Any deploy, migration, or schema change recently?
Replication / failover healthy? Did we fail over?
Connection saturation, blocking, or long-running queries?
Is the provider's status page green? Are we in their SLA window?
Do we have a fallback or cache we can lean on?
Look for
Upstream incident, timeouts/error codes to the vendor endpoint
📣 Product / Business
Ask
Exactly who and what is affected? Which customer journeys?
Who owns customer/stakeholder comms?
Any regulatory or revenue-critical clock ticking?
Look for
Support ticket spike, affected revenue paths, SLA/contract exposure
Calm, factual, on a cadence. State impact, what you're doing, and the next update time — always give a next update time. Cadence: SEV1 ~15 min · SEV2 ~30 min · SEV3 hourly.
🌊 Be like water. You don't need the perfect words — you need honest ones, on time. "Here's what we know, here's what we're doing, here's when I'll be back." That's enough. Send it, breathe, keep going.
Initial notification
We are investigating [degradation / outage] affecting [scope / customer journey], first observed at [time TZ]. Impact: [what users see]. A bridge is open and the team is engaged. Next update by [time].
Update (still working)
Update on [incident] — [time TZ]. Status: [investigating / mitigating / monitoring]. What we know: [1-2 lines]. Current impact: [scope]. Actions in progress: [mitigation]. Next update by [time].
Mitigated (service restored, watching)
[Incident] — service has been restored as of [time TZ] via [mitigation]. We are monitoring for stability and confirming full recovery. Root cause analysis to follow. Next update by [time] or on change.
Resolved
[Incident] is resolved as of [time TZ]. Duration: [start–end]. Impact: [summary]. Cause (preliminary): [brief]. A blameless review is scheduled and a full writeup will follow. Thank you for your patience.
Stakeholder / exec one-liner
[SEV#] [system] — [impact in business terms]. Team engaged since [time], [mitigating/monitoring]. Customer impact: [scope]. Next update [time].
Timestamped record of what you saw and did. Feeds your status updates and the postmortem. Stored only in this browser — nothing is uploaded.
Incident Commander Field Kit — a ReliabilityOps tool. Offline, single-file, read-only; adapt it to your organization's severity model, roles, and comms policy. Nothing on this page calls the network; your timeline never leaves the browser.