skip to content

In an incident-response process, what does the Incident Commander actually own during an active outage, and does the IC have to be the most technically knowledgeable person on the call?

level: juniorimportance: should knowfreq 62%

answer

  1. owns the response, not the fix
  2. coordination role, not a seniority badge
  3. every task gets a named owner
  4. expertise drags you into the debugger
  5. announced out loud, never assumed

basics

~20 s

The Incident Commander owns the response, not the fix: they hold the current picture, assign every task to a named person, and make the calls. Deep technical knowledge is not the qualification — command is a coordination job a trained responder holds.

solid answer

~50 s

The Incident Commander owns the response as a whole. They hold the current picture of what is broken and what has been tried, decide what gets worked next, assign every task to a **named person** rather than to the room, and are the single place decisions land. They do not have to be the deepest expert on the failing system — often they should not be, because expertise drags you into a stack trace and you stop tracking everything else. What an IC needs is fluency in the process and the standing to direct people, including saying "stop that line, we are rolling back." On a small incident one engineer holds command implicitly; the moment there is more than one workstream, command becomes a real job that someone announces out loud so everyone knows who it is.

go deeper

for a junior

Be able to say plainly that the Incident Commander coordinates the response and does not do the fixing, and that any trained responder can hold the role regardless of tenure.

for a middle

Explain how command is exercised in practice: one holder of state, every task assigned to a named person, one decision-maker, and an explicit announcement of who has command.

for a senior

Show you have actually held it — how you keep the picture straight while people argue, timebox investigations, and choose the reversible option when the experts disagree.

for a principal

Own the argument that command authority comes from the process rather than from title, and describe how you make an IC's call stick when a principal engineer or an executive pushes back mid-incident.

## What "command" means Incident command is a role borrowed from emergency services: one person is designated as responsible for the *response*, and everyone else's work flows through them. In software incident response the Incident Commander (IC) is the single-threaded owner of four things: - **State.** What is broken, who is affected, how badly, what has been tried, what has been ruled out. If someone joins the call ten minutes in, the IC (or the IC's written summary) is what gets them current. - **Priority.** Which of the three plausible lines of investigation is being worked right now, and which are explicitly parked. - **Assignment.** Every task has a named owner. "Someone should check the database" is not an assignment; "Priya, check replica lag and report in five minutes" is. - **The decision.** When the room disagrees, the IC picks. Not because they know more, but because a response with two decision-makers has none. ## Why it is not a seniority badge The most common misconception is that the IC is whoever is most senior or knows the system best. That gets the causality backwards. The person who knows the system best is your most valuable *debugger*, and command is precisely the job that makes debugging impossible: the moment you are reading a stack trace, you have stopped tracking the other workstreams, the clock, and the people waiting for an answer. Attention is single-threaded. Handing command to your best expert typically produces an incident with a great debugger and no one steering. Command is also not a technical adjudication role. When two engineers propose different mitigations, the IC does not decide which is technically superior. They ask each one the same three questions — how long until users feel the effect, what is the blast radius if it is wrong, and can we undo it — and then pick, usually the reversible option, assign it to a named person, and put a timebox on it. That is a judgment about cost and risk, and it does not require being the better engineer. ## What the IC is doing minute to minute In practice a competent IC spends the incident doing unglamorous things: restating the current state out loud every few minutes so it stays synchronised; asking "who owns that?" every time a task appears; timeboxing investigations ("you have ten minutes on that theory, then we try the rollback"); protecting responders from interruption so the people with hands on the system are not answering questions from three directions; and repeatedly asking the standing question of incident response — is there something we can do that restores users *without* knowing the cause yet? The IC also decides that the incident is over. That means confirming the user-visible impact is actually gone rather than the alert merely having stopped firing, saying so explicitly in the channel, releasing responders so they are not idling on a bridge, and naming who owns the follow-up. ## What the IC does not own Just as importantly: the IC is not the person typing into production, not the person writing customer-facing updates, and not the person taking the executive's call. Those are separate jobs precisely so they do not eat the commander's attention. On a small incident one person may hold several of them at once — that is fine and normal — but they should know they are holding several roles, and know the trigger to hand one off: a second workstream appearing, or someone outside the response starting to demand attention. ## Making command explicit The failure mode that looks harmless is *implicit* command. Six people are on a call, everyone assumes someone is steering, nobody is, and forty minutes later the timeline shows three uncoordinated changes made to production. The fix costs one sentence: someone says "I am taking command," it goes in the incident channel, and it is visible to anyone who joins. If you take one thing from the role, it is that command must be **announced**, not assumed — and that the announcement is available to anyone, not just the most senior person present.

  • If the IC is not the deepest expert, how do they choose between two competing technical proposals?
    They do not adjudicate on technical depth. They ask each proposer the same questions — how fast will users feel it, what is the blast radius if it is wrong, and can we undo it — then pick, usually the reversible option, assign it to a named person, and timebox it. A decision made in two minutes and reversed beats a twenty-minute debate.
  • Who decides an incident is over, and what does that decision involve?
    The IC does. It means confirming the user-visible impact is actually gone rather than just the alert having stopped, saying so explicitly in the incident channel, standing responders down so nobody idles on a bridge, and naming who owns the follow-up before the channel disperses. Left implicit, incidents drift from active to abandoned.

saying these in an interview costs you the question

  • The IC should be the most senior engineer present
  • The IC is whoever happened to notice the alert
  • The IC debugs the problem and reports what they found
  • Command is implicit — nobody needs to state who has it
  • The IC's main job is briefing executives

context