A major outage pulls a dozen engineers onto the bridge. Your incident process defines an Incident Commander, an Operations (Tech) Lead, a Communications Lead and a Scribe. What does each role own, and why is the Incident Commander explicitly kept out of the debugging?
answer
- four separable streams of attention
- one funnel for production changes
- comms lead is an interrupt shield
- attention is single-threaded
- span of control roughly five people
basics
~20 sThe Incident Commander steers, the Operations Lead is the only one changing production, the Communications Lead handles everyone outside the response, and the Scribe timestamps decisions and actions. The IC stays out of debugging because attention is single-threaded — in a stack trace, nobody is steering.
solid answer
~50 sFour separable jobs, one person each. The **Incident Commander** holds state, sets priority, assigns work to named people and makes the calls. The **Operations or Tech Lead** is the only one with hands on production — they direct the responders actually running commands, so changes go through one funnel instead of three people editing config simultaneously. The **Communications Lead** absorbs everyone outside the response: stakeholders, support, execs, so the responders are not answering questions from four directions. The **Scribe** keeps a timestamped record of what was observed, decided and changed. The IC is kept out of debugging because attention is single-threaded: the moment they are reading logs they stop tracking the clock, the other workstreams and the people waiting — and nobody notices the response has lost its driver until much later. On a three-person incident one person legitimately wears several hats; you split when a second workstream or outside pressure appears.
go deeper
Know the four roles by name and what each one owns, and be able to say why the person coordinating is not the person typing into production.
Explain the mechanics: one funnel for production changes, comms as an interrupt shield, a running timestamped log, and attention as the reason command and debugging cannot share a head.
Show the scaling judgment — which role you split off first on a growing incident, when four roles is theatre, and how you recover a response where three people are silently changing production.
Argue for the role model as an organisational investment: who is trained for each seat, how you keep the process light enough that a 3am SEV3 still follows it, and how you measure whether roles are actually being held.
## The four jobs The role set is derived from the Incident Command System used in emergency services, and it exists to answer one question: when many people respond to one problem, how do you stop them colliding? Each role is a separable stream of attention. **Incident Commander.** Owns the response, not the system. Holds the current picture — impact, what has been tried, what is ruled out — sets what is being worked next, assigns each task to a named person, and decides when the room disagrees. The IC does not type into production and does not talk to executives. **Operations / Tech Lead.** The only role that changes the system, and the funnel every change goes through. On a large incident, several subject-matter experts are investigating in parallel; the Ops Lead directs them and serialises the actual production changes. This is the role that prevents the worst compound failure in incident response: two people independently "fixing" the same thing, so that when behaviour changes nobody knows which change did it, or whether they cancelled each other out. A useful discipline is that any production change is stated in the channel before it happens — "I am restarting the three EU pods now" — so the timeline is real and the IC can veto. **Communications Lead.** Owns everyone who is not responding: internal stakeholders, support and account teams, leadership. Their real function is to be an *interrupt shield*. Without one, the exec who wants an ETA asks the person with hands on the keyboard, which is both the slowest possible way to get an answer and a direct tax on time-to-recovery. **Scribe.** Keeps a timestamped log of observations, decisions and actions as they happen. This is not administrative overhead — during the incident it is the shared memory that lets a newly joined responder get current without interrupting anyone, and it is the only version of the timeline that is not reconstructed from memory afterwards. Role names vary between organisations. Google's SRE material describes incident command, operational work, communication and planning; other widely published processes use Incident Commander, Deputy, Scribe and subject-matter expert. The names matter less than the fact that the streams are separated and that everyone knows who holds which. ## Why the commander must not debug This is the part interviewers are actually probing, and the answer is about attention, not hierarchy. Debugging is deep, absorbing and narrow. Command is shallow, interruptible and broad. They are not merely different tasks — they are incompatible modes. An IC who opens a dashboard and starts forming a theory stops doing four things at once: tracking elapsed time, noticing that the second workstream has gone quiet, noticing that no comms update has gone out in twenty minutes, and remaining willing to abandon the current theory. The characteristic outcome is an incident where the technical work is excellent and the response is forty minutes longer than it needed to be, because nobody asked "can we mitigate without understanding this?" until late. There is a second, subtler reason. The person deep in a theory is the person least able to judge that the theory is wrong. Command is where you want someone uninvested enough to say "we have spent twenty minutes on this line, park it, roll back." If the IC is also the theory's author, that check disappears. ## The cost, and when not to split Splitting roles is not free: it costs people. Four roles on a two-person SEV3 at 3am is bureaucratic theatre, and a process that demands it will simply be ignored. The honest position in an interview is that role separation scales with the incident: - One responder holds all four roles implicitly, and should know that they are. - The first split, almost always, is comms — because outside pressure arrives before technical complexity does. - Command separates from operations once there is more than one workstream, or more than roughly five people to keep track of. The emergency-services heuristic is a span of control of three to seven, five being comfortable; beyond that, one person cannot hold who is doing what. - The Scribe is the role most often skipped and most often regretted, because the timeline cannot be recovered later at anything like the same fidelity. ## Making it work in practice Roles are announced, pinned in the incident channel, and re-announced when they change. Anyone joining should be able to read who holds what without asking. And the roles are *held*, not *assumed*: the failure mode is not usually a bad IC, it is four people who each believe someone else is commanding while three of them quietly make production changes.
- Your incident has three engineers total. Do you still staff all four roles?No. Demanding four roles on a three-person incident is theatre and gets the process ignored. One person holds command and comms, another has hands on the system, and everyone knows the hats being worn. Split as the incident grows — comms first, because outside pressure usually arrives before technical complexity does.
- Why insist that production changes funnel through one role rather than letting experts act directly?Because concurrent changes destroy the causal signal. If two people restart and reconfigure at the same moment and the error rate drops, you cannot tell which acted, whether they cancelled out, or what you must not repeat. Serialising through the Ops Lead — with each change announced before it happens — keeps the timeline honest and keeps the IC able to veto.
- What does the Scribe give you during the incident, as opposed to afterwards?Shared memory. A timestamped running log lets a newly joined responder get current without interrupting anyone, keeps the IC from re-deriving what has been ruled out, and stops the room from re-testing a theory someone dismissed thirty minutes earlier. The postmortem value is real but secondary.
The IC is the air-traffic controller, not the pilot. The controller has never flown that particular aircraft and does not need to — their whole value is being the only person who can see everything in the sky at once, which is exactly what they lose the moment they climb into a cockpit.
saying these in an interview costs you the question
- Roles are overhead; good engineers just self-organise
- The IC should debug too — they are the best engineer available
- Everyone can push fixes as long as they mention it later
- The Scribe can reconstruct the timeline afterwards from Slack
- Executives should get their ETA from whoever is fixing it