skip to content

Incident Management

Running an outage well: declaring incidents, severity levels, the Incident Commander system, and disciplined communication. 'Walk me through an outage you handled' appears in almost every senior interview loop — this is the vocabulary and structure it expects.

on this pageshow

questions

22

Why do incident response practices call for a dedicated channel or bridge per incident rather than discussing the outage in the team's usual chat channel, and what belongs on a voice bridge versus in writing?

level: juniorimportance: must knowfreq 62%

answer

  1. one incident, one authoritative place
  2. joiners read instead of asking status
  3. direct messages hide the decisions
  4. bridge to think, channel to record
  5. working channel apart from status channel

basics

~20 s

A per-incident channel gives one authoritative place for the response, so anyone joining reads the scrollback instead of interrupting responders, and it becomes the raw material for the postmortem timeline. A voice bridge is faster for fast-moving coordination but leaves no record, so decisions and actions still get typed into the channel.

solid answer

~50 s

One incident, one place. A dedicated channel means a person joining at minute forty catches up by reading rather than by asking "what's the status?" — which is the interruption that most damages a response. It keeps the outage traffic out of the team's routine channel where it would be interleaved with unrelated work, gives everything a natural boundary for the postmortem timeline, and stops the response fragmenting across direct messages where half the decisions become invisible. A voice bridge is genuinely better for rapid back-and-forth, screen sharing and fast decisions under time pressure, so most serious incidents run both. The rule that makes that safe is: anything said on the bridge that is an observation, an action taken, or a decision gets typed into the channel too. Otherwise the record has a hole exactly where the incident was moving fastest. It is also worth separating the responders' working channel from a status channel for people who only want updates.

go deeper

for a junior

Be ready to say that a dedicated channel gives one place to catch up from, keeps the outage out of routine traffic, and becomes the postmortem record — and that anything decided on a call must still be typed into it.

for a middle

Explain why direct messages and side conversations are the real damage, and describe the two-surface split between a responder working channel and a stakeholder status channel.

for a senior

Show the trade-off honestly: a bridge is faster for coordination but silent for the record, so name the discipline that keeps observations, actions and decisions written down while the incident is moving fastest.

for a principal

Own the mechanics that make it automatic — channel created on declaration with a consistent naming scheme, pinned impact and commander, change events posting themselves, archive rather than delete — so the practice does not depend on anyone remembering it at 03:00.

## One incident, one place The single organising rule of incident communication is that at any moment there is exactly one authoritative location for the response. Creating a channel per incident, named so it is unmistakable, achieves this cheaply. Everything follows from it. **Self-service catch-up.** People join an incident continuously — an escalated specialist, a support lead, someone waking up. If the state of the response lives in a channel, they read it. If it does not, every arrival costs a responder an interruption to explain, and interruptions land hardest on precisely the people whose attention is scarcest. "Read the channel, then ask what is unclear" is only possible if the channel is complete. **Separation from routine traffic.** In the team's normal channel the incident is interleaved with standup chatter, code review pings and someone's unrelated question. Nobody can reconstruct it later, and newcomers cannot tell which of the last two hundred messages are the outage. **A boundary for the record.** The channel's contents are the raw material for the postmortem timeline. A per-incident channel gives that record clean start and end boundaries with nothing to filter out. **No side channels.** The corrosive failure is direct messages. Two people solve part of the problem in a DM, the channel never learns it, someone else repeats the work, and the timeline is missing the piece that mattered. The same applies to a vendor's support portal or another team's channel: mirror the outcome back into the incident channel even when the conversation had to happen elsewhere. ## The bridge trade-off A voice or video bridge is not a competitor to the channel; it does something the channel cannot. Voice is far faster for high-bandwidth exchange — "what does that graph do at 14:02?", "try it now", "no, stop" — and screen sharing conveys in ten seconds what would take five minutes to describe. For a fast-moving incident with several people, a bridge is usually the right call. Its cost is that it is invisible and unsearchable. Nothing said on a call exists for someone joining later, for a stakeholder update, or for the postmortem. A bridge-only incident regularly produces a postmortem where nobody can say who decided to fail over or what they knew at the time. The resolution is a simple discipline: the bridge is where you think, the channel is where you record. Any observation, any action taken against production, and any decision gets typed into the channel as it happens, usually by whoever is coordinating rather than by the person with their hands on the system. One line is enough — "decision: failing over to the secondary, accepting up to 30 seconds of lost writes". Post the bridge link in the channel too, so joiners find it without asking. ## Working channel versus status channel As an incident grows, the responder channel attracts an audience: managers, account teams, curious colleagues. Their questions are legitimate but they are interruptions. The standard split is two surfaces — the working channel for responders, and a separate stakeholder or status channel that carries only the periodic updates. Say clearly in the working channel where updates will be posted so people can subscribe to the right one, and be prepared to move a stakeholder conversation out of the working channel politely and immediately. ## Practical mechanics A few things make this work rather than merely sound good. Create the channel automatically at declaration so nobody spends the first three minutes on naming, and use a consistent scheme so the channel is findable later. Pin the essentials at the top: current impact statement, incident commander, bridge link, and the location of any shared document. Post the current status periodically into the channel itself so a joiner sees it without scrolling. Have deployment, rollback and configuration systems post their events into the channel automatically, so the highest-value factual entries appear without anyone typing. And archive rather than delete when it is over, because that scrollback is the evidence the postmortem is built on. ## The decision behind the question Asked plainly: is the extra ceremony of a dedicated channel worth it for a small incident? For a five-minute blip handled by one person, it can be overhead. The reason mature teams do it anyway is that you cannot tell which incidents are small until afterwards, and you cannot retroactively create a record of the first twenty minutes of the one that turned out to be large. Creating the channel is cheap and reversible; discovering at hour two that the first hour was conducted in three separate direct-message threads is neither.

  • Two responders solve a piece of the problem in a direct message. What is the actual harm?
    Three harms. The rest of the response does not learn it and may duplicate the work; a joiner reading the channel gets a false picture of the state; and the postmortem timeline is missing the step that mattered. Anything that changes the shared understanding of the incident belongs in the channel, even when the conversation had to start elsewhere.
  • A manager keeps asking for status in the responder channel. How do you handle it?
    Point them to the stakeholder channel where updates are posted on a stated cadence, and answer once there rather than in the working channel. If no such channel exists, create it immediately — the request is legitimate, and the fix is a place to serve it that does not consume responder attention.
  • When is a dedicated incident channel genuinely unnecessary overhead?
    For a brief, low-severity issue handled by a single person with no customer impact and an obvious fix. The caveat is that you cannot tell in advance which incident becomes the large one, and the first twenty minutes cannot be reconstructed later — so teams that automate channel creation on declaration pay almost nothing and never face that gap.

saying these in an interview costs you the question

  • Running the incident in the team's normal chat channel
  • Coordinating the response through direct messages between responders
  • Holding the whole incident on a call with nothing written down
  • Letting stakeholders ask for status in the responders' working channel
  • Deleting the incident channel once the incident is closed

context

open as a page

You are on call. Five minutes after a routine deploy, your service's error rate jumps from 0.1% to 12% and users are seeing failures. What is your first action, and why is "open the logs and find the bug" the wrong one?

level: juniorimportance: must knowfreq 80%

basics

~20 s

Roll back to the last known-good release first, then investigate. Restoring users is the goal during an incident, the deploy timing is strong enough evidence to act on, and reading logs leaves users broken for however long the debugging takes.

open as a page

You are handling communications for an ongoing SEV1 outage where customer logins are failing. What goes into each stakeholder update, how often do you send one, and what should you never promise in it?

level: middleimportance: must knowfreq 70%

basics

~20 s

Incident updates go out on a fixed cadence tied to severity — commonly every 30 minutes at the top severity — and each one states user-visible impact, what the response is doing now, and the time of the next update. Never promise a restoration ETA.

open as a page

A major outage pulls a dozen engineers onto the bridge. Your incident process defines an Incident Commander, an Operations (Tech) Lead, a Communications Lead and a Scribe. What does each role own, and why is the Incident Commander explicitly kept out of the debugging?

level: middleimportance: must knowfreq 78%

basics

~20 s

The Incident Commander steers, the Operations Lead is the only one changing production, the Communications Lead handles everyone outside the response, and the Scribe timestamps decisions and actions. The IC stays out of debugging because attention is single-threaded — in a stack trace, nobody is steering.

open as a page

Your team is writing a SEV1–SEV4 severity matrix for its production services. Which dimensions should decide an incident's severity, and why must the matrix be written in terms of user impact rather than which component broke?

level: middleimportance: must knowfreq 72%

basics

~20 s

Grade on observable user impact: what share of users or requests is affected, how critical the blocked journey is, whether a workaround exists, and whether data or money is at risk. Which component failed predicts none of those.

open as a page

An incident is active and users are affected. You have several generic levers available: roll back the last release, fail over to another region or replica, flip a feature kill switch, shed load, or scale up. How do you choose between them in the first few minutes?

level: seniorimportance: must knowfreq 65%

basics

~20 s

Match the lever to the most likely change vector — code, config, a dependency, traffic, or capacity — then prefer whichever is fastest to take effect, smallest in blast radius, and easiest to undo. Apply one lever at a time so the user-facing signal tells you which one worked.

open as a page

In an incident-response process, what does the Incident Commander actually own during an active outage, and does the IC have to be the most technically knowledgeable person on the call?

level: juniorimportance: should knowfreq 62%

basics

~20 s

The Incident Commander owns the response, not the fix: they hold the current picture, assign every task to a named person, and make the calls. Deep technical knowledge is not the qualification — command is a coordination job a trained responder holds.

open as a page

In production operations, what is the difference between a monitoring alert firing and an incident being declared, and what happens in between?

level: juniorimportance: should knowfreq 55%

basics

~20 s

An alert is an automated signal that some condition was met. An incident is a human declaration that real or suspected impact warrants coordinated response, with a severity, an owner and a written record. Triage is the step between them.

open as a page

What has to be recorded while an incident is still in progress for the postmortem timeline to be usable, and why can't you just reconstruct it afterwards from the chat log?

level: middleimportance: should knowfreq 50%

basics

~20 s

Capture timestamped entries live for observations, actions and decisions, tagged as which of the three they are, plus the four anchor times: detection, declaration, mitigation applied, impact ended. Reconstruction fails because chat records when something was mentioned, not when it happened, and memory rewrites what people knew.

open as a page

Rolling back, restarting, or replacing instances during an incident usually destroys the state you would need to explain the failure later. What do you capture before you pull the mitigation lever, and how do you keep that capture from delaying the mitigation itself?

level: middleimportance: should knowfreq 45%

basics

~20 s

Capture only what dies with the process — thread stacks, heap state, local files, in-memory queues — and prefer quarantining one failing instance out of rotation over restarting them all. Anything already shipped to central logging, metrics or tracing survives the mitigation, so do not wait for it.

open as a page

You applied a mitigation five minutes ago and the dashboards look calmer. How do you decide the incident is actually mitigated, and what commonly makes an incident look recovered when it is not?

level: middleimportance: should knowfreq 50%

basics

~20 s

Verify against the user-facing SLI at the granularity that failed, not against a calmer dashboard or a cleared alert. The classic traps are an error ratio that fell because traffic fell, a metric window that has not turned over yet, and a backlog still draining behind a healthy-looking front door.

open as a page

Forty minutes into an incident your team still cannot explain the failure and a managed database service is one suspect. How do you decide when to escalate to another team or to the vendor, and what makes an escalation effective rather than a handoff?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Escalate on a clock, not on exhaustion: set a time box when the incident starts and pull in the specialist or vendor when it expires, even if you might have solved it yourself. Effective escalation keeps ownership local, runs in parallel with mitigation, and arrives with evidence attached.

open as a page

Your public API has been returning errors for twelve minutes and the cause is still unknown. How does the public status-page post differ from the internal stakeholder update, and what determines when you post publicly?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Public posts are triggered by confirmed customer-visible impact, not by a known cause. They state the symptom, scope and next update time in plain language with no internal names or speculation, while the internal update stays candid about hypotheses, ruled-out theories and mitigation options.

open as a page

Rollback is the usual first mitigation, but sometimes fixing forward is genuinely the less risky choice during a live incident. Give the concrete conditions under which you would fix forward, and how you would bound that decision.

level: seniorimportance: should knowfreq 55%

basics

~20 s

Fix forward when rollback is impossible, ineffective, or slower than the fix: an irreversible migration has already run, the previous build carries the same defect, or reverting takes forty minutes while a one-line change takes four. Bound it with a hard deadline and a cruder fallback mitigation.

open as a page

An incident has been running for six hours and the same person has held Incident Commander since it started. How do you hand off the IC role mid-incident without losing the response, and what does the handoff have to transfer?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Hand off explicitly, never by drift: brief the incoming IC on impact, timeline, active workstreams and their owners, and pending commitments; state "you are now IC" in the channel; then the outgoing IC shadows briefly and leaves. Fatigued judgment costs more than the context reload.

open as a page

Ninety minutes into a SEV1, your most senior engineer is running the whole response alone — reading logs, applying changes, and answering executives directly — and nobody else on the bridge has been assigned anything. What is wrong with this, and what do you do about it?

level: seniorimportance: should knowfreq 50%

basics

~20 s

One person holding every role means the response has no coordination, no timeline, and a bus factor of one: all state is in their head. Take command, get a spoken state dump, write it down, split the work among named owners, and put someone between them and the executives.

open as a page

It is 02:10 and you have one ambiguous signal: elevated 5xx errors on one of six API instances, no customer reports yet. Make the case for declaring an incident now on suspicion versus waiting for confirmation, and say how you would make that call repeatable for your team.

level: seniorimportance: should knowfreq 62%

basics

~20 s

Declare now. The costs are asymmetric: an over-declaration is undone with one message and a few interrupted minutes, while a late declaration adds unmitigated impact plus the response ramp-up you could have run in parallel. Declare low and adjust.

open as a page

You declared a SEV3 for elevated checkout errors. Forty minutes in, you discover the same bug also wrote incorrect balances to a subset of accounts. How do you handle severity mid-incident, and what are the rules for downgrading one?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Upgrade immediately on discovery, not at the next sync, and restate the grade explicitly with the time. Downgrade only after impact has actually stopped and recovery is verified — never to quiet the response — and record the peak severity, not the closing one.

open as a page

During a long revenue-affecting outage, executives and account managers start messaging responders directly for updates. How would you structure incident communications so leadership stays informed without interrupting the people fixing it?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Give each audience its own surface with one publishing point: a responder working channel, a broadcast status channel on a fixed cadence, and a named human who briefs leadership directly. Agree notification thresholds by severity in advance so nobody negotiates access during the incident.

open as a page

Across your organisation, incidents are detected in about three minutes but the median time to mitigate is around forty-five. As the engineering lead, what would you change so responders can stop user impact faster?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Attack the two things that fill those forty-two minutes: decision latency and lever latency. Give each service a short menu of mitigations with measured times, pre-authorise on-call engineers to pull them without approval, and treat a slow rollback or an undrilled failover as a defect to fix.

open as a page

You are establishing an incident-command practice for an engineering organisation of roughly 300 people. How do you decide who is allowed to act as Incident Commander, and what authority does the role need to be granted before it works?

level: principalimportance: nice to knowfreq 34%

basics

~20 s

Build a trained, cross-team pool of commanders rather than defaulting to service owners, qualify them by shadowing then leading under supervision, and have leadership grant the authority in writing beforehand: an IC can pull in anyone, suspend other priorities, and override objections for the incident's duration.

open as a page

Across forty engineering teams, incident severity grades are wildly inconsistent — one team's SEV1 is another's SEV3, and one team has not declared anything above SEV3 in a year. As the lead of the reliability practice, how would you calibrate severity across the organisation?

level: principalimportance: nice to knowfreq 34%

basics

~20 s

Treat it as an incentive problem, not a wording problem. Diagnose whether teams are inflating or deflating, calibrate with worked reference incidents rather than longer definitions, re-grade a sample of closed incidents regularly, and decouple punishing process from high tiers.

open as a page