skip to content

Incident Command & Roles

The ICS-derived role structure that keeps large incidents coordinated instead of chaotic. Interviewers ask 'who does what during a major outage?' to see if you know why the Incident Commander must not be the one debugging.

on this pageshow

questions

5

A major outage pulls a dozen engineers onto the bridge. Your incident process defines an Incident Commander, an Operations (Tech) Lead, a Communications Lead and a Scribe. What does each role own, and why is the Incident Commander explicitly kept out of the debugging?

level: middleimportance: must knowfreq 78%

answer

  1. four separable streams of attention
  2. one funnel for production changes
  3. comms lead is an interrupt shield
  4. attention is single-threaded
  5. span of control roughly five people

basics

~20 s

The Incident Commander steers, the Operations Lead is the only one changing production, the Communications Lead handles everyone outside the response, and the Scribe timestamps decisions and actions. The IC stays out of debugging because attention is single-threaded — in a stack trace, nobody is steering.

solid answer

~50 s

Four separable jobs, one person each. The **Incident Commander** holds state, sets priority, assigns work to named people and makes the calls. The **Operations or Tech Lead** is the only one with hands on production — they direct the responders actually running commands, so changes go through one funnel instead of three people editing config simultaneously. The **Communications Lead** absorbs everyone outside the response: stakeholders, support, execs, so the responders are not answering questions from four directions. The **Scribe** keeps a timestamped record of what was observed, decided and changed. The IC is kept out of debugging because attention is single-threaded: the moment they are reading logs they stop tracking the clock, the other workstreams and the people waiting — and nobody notices the response has lost its driver until much later. On a three-person incident one person legitimately wears several hats; you split when a second workstream or outside pressure appears.

go deeper

for a junior

Know the four roles by name and what each one owns, and be able to say why the person coordinating is not the person typing into production.

for a middle

Explain the mechanics: one funnel for production changes, comms as an interrupt shield, a running timestamped log, and attention as the reason command and debugging cannot share a head.

for a senior

Show the scaling judgment — which role you split off first on a growing incident, when four roles is theatre, and how you recover a response where three people are silently changing production.

for a principal

Argue for the role model as an organisational investment: who is trained for each seat, how you keep the process light enough that a 3am SEV3 still follows it, and how you measure whether roles are actually being held.

## The four jobs The role set is derived from the Incident Command System used in emergency services, and it exists to answer one question: when many people respond to one problem, how do you stop them colliding? Each role is a separable stream of attention. **Incident Commander.** Owns the response, not the system. Holds the current picture — impact, what has been tried, what is ruled out — sets what is being worked next, assigns each task to a named person, and decides when the room disagrees. The IC does not type into production and does not talk to executives. **Operations / Tech Lead.** The only role that changes the system, and the funnel every change goes through. On a large incident, several subject-matter experts are investigating in parallel; the Ops Lead directs them and serialises the actual production changes. This is the role that prevents the worst compound failure in incident response: two people independently "fixing" the same thing, so that when behaviour changes nobody knows which change did it, or whether they cancelled each other out. A useful discipline is that any production change is stated in the channel before it happens — "I am restarting the three EU pods now" — so the timeline is real and the IC can veto. **Communications Lead.** Owns everyone who is not responding: internal stakeholders, support and account teams, leadership. Their real function is to be an *interrupt shield*. Without one, the exec who wants an ETA asks the person with hands on the keyboard, which is both the slowest possible way to get an answer and a direct tax on time-to-recovery. **Scribe.** Keeps a timestamped log of observations, decisions and actions as they happen. This is not administrative overhead — during the incident it is the shared memory that lets a newly joined responder get current without interrupting anyone, and it is the only version of the timeline that is not reconstructed from memory afterwards. Role names vary between organisations. Google's SRE material describes incident command, operational work, communication and planning; other widely published processes use Incident Commander, Deputy, Scribe and subject-matter expert. The names matter less than the fact that the streams are separated and that everyone knows who holds which. ## Why the commander must not debug This is the part interviewers are actually probing, and the answer is about attention, not hierarchy. Debugging is deep, absorbing and narrow. Command is shallow, interruptible and broad. They are not merely different tasks — they are incompatible modes. An IC who opens a dashboard and starts forming a theory stops doing four things at once: tracking elapsed time, noticing that the second workstream has gone quiet, noticing that no comms update has gone out in twenty minutes, and remaining willing to abandon the current theory. The characteristic outcome is an incident where the technical work is excellent and the response is forty minutes longer than it needed to be, because nobody asked "can we mitigate without understanding this?" until late. There is a second, subtler reason. The person deep in a theory is the person least able to judge that the theory is wrong. Command is where you want someone uninvested enough to say "we have spent twenty minutes on this line, park it, roll back." If the IC is also the theory's author, that check disappears. ## The cost, and when not to split Splitting roles is not free: it costs people. Four roles on a two-person SEV3 at 3am is bureaucratic theatre, and a process that demands it will simply be ignored. The honest position in an interview is that role separation scales with the incident: - One responder holds all four roles implicitly, and should know that they are. - The first split, almost always, is comms — because outside pressure arrives before technical complexity does. - Command separates from operations once there is more than one workstream, or more than roughly five people to keep track of. The emergency-services heuristic is a span of control of three to seven, five being comfortable; beyond that, one person cannot hold who is doing what. - The Scribe is the role most often skipped and most often regretted, because the timeline cannot be recovered later at anything like the same fidelity. ## Making it work in practice Roles are announced, pinned in the incident channel, and re-announced when they change. Anyone joining should be able to read who holds what without asking. And the roles are *held*, not *assumed*: the failure mode is not usually a bad IC, it is four people who each believe someone else is commanding while three of them quietly make production changes.

  • Your incident has three engineers total. Do you still staff all four roles?
    No. Demanding four roles on a three-person incident is theatre and gets the process ignored. One person holds command and comms, another has hands on the system, and everyone knows the hats being worn. Split as the incident grows — comms first, because outside pressure usually arrives before technical complexity does.
  • Why insist that production changes funnel through one role rather than letting experts act directly?
    Because concurrent changes destroy the causal signal. If two people restart and reconfigure at the same moment and the error rate drops, you cannot tell which acted, whether they cancelled out, or what you must not repeat. Serialising through the Ops Lead — with each change announced before it happens — keeps the timeline honest and keeps the IC able to veto.
  • What does the Scribe give you during the incident, as opposed to afterwards?
    Shared memory. A timestamped running log lets a newly joined responder get current without interrupting anyone, keeps the IC from re-deriving what has been ruled out, and stops the room from re-testing a theory someone dismissed thirty minutes earlier. The postmortem value is real but secondary.

The IC is the air-traffic controller, not the pilot. The controller has never flown that particular aircraft and does not need to — their whole value is being the only person who can see everything in the sky at once, which is exactly what they lose the moment they climb into a cockpit.

saying these in an interview costs you the question

  • Roles are overhead; good engineers just self-organise
  • The IC should debug too — they are the best engineer available
  • Everyone can push fixes as long as they mention it later
  • The Scribe can reconstruct the timeline afterwards from Slack
  • Executives should get their ETA from whoever is fixing it

context

open as a page

In an incident-response process, what does the Incident Commander actually own during an active outage, and does the IC have to be the most technically knowledgeable person on the call?

level: juniorimportance: should knowfreq 62%

basics

~20 s

The Incident Commander owns the response, not the fix: they hold the current picture, assign every task to a named person, and make the calls. Deep technical knowledge is not the qualification — command is a coordination job a trained responder holds.

open as a page

An incident has been running for six hours and the same person has held Incident Commander since it started. How do you hand off the IC role mid-incident without losing the response, and what does the handoff have to transfer?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Hand off explicitly, never by drift: brief the incoming IC on impact, timeline, active workstreams and their owners, and pending commitments; state "you are now IC" in the channel; then the outgoing IC shadows briefly and leaves. Fatigued judgment costs more than the context reload.

open as a page

Ninety minutes into a SEV1, your most senior engineer is running the whole response alone — reading logs, applying changes, and answering executives directly — and nobody else on the bridge has been assigned anything. What is wrong with this, and what do you do about it?

level: seniorimportance: should knowfreq 50%

basics

~20 s

One person holding every role means the response has no coordination, no timeline, and a bus factor of one: all state is in their head. Take command, get a spoken state dump, write it down, split the work among named owners, and put someone between them and the executives.

open as a page

You are establishing an incident-command practice for an engineering organisation of roughly 300 people. How do you decide who is allowed to act as Incident Commander, and what authority does the role need to be granted before it works?

level: principalimportance: nice to knowfreq 34%

basics

~20 s

Build a trained, cross-team pool of commanders rather than defaulting to service owners, qualify them by shadowing then leading under supervision, and have leadership grant the authority in writing beforehand: an IC can pull in anyone, suspend other priorities, and override objections for the incident's duration.

open as a page