You are establishing an incident-command practice for an engineering organisation of roughly 300 people. How do you decide who is allowed to act as Incident Commander, and what authority does the role need to be granted before it works?
answer
- trained cross-team pool, not the service owner
- command decays without practice
- authority granted in writing, in advance
- what happens when a VP joins the bridge
- distinct commanders per quarter
basics
~20 sBuild a trained, cross-team pool of commanders rather than defaulting to service owners, qualify them by shadowing then leading under supervision, and have leadership grant the authority in writing beforehand: an IC can pull in anyone, suspend other priorities, and override objections for the incident's duration.
solid answer
~50 sTwo decisions carry the practice. First, **who**: a deliberately cross-team pool of trained commanders, not the owner of whichever service broke — the owner is your best debugger and the worst candidate for command. Qualification is a short curriculum plus shadowing a real incident, then commanding one with an experienced IC watching. Aim for a pool wide enough that no individual is on the hook constantly; in a 300-person org that is tens of people, not five. Second, **authority**: it has to be granted by leadership up front and in writing, because during an incident there is no time to negotiate it. A commander must be able to page anyone regardless of team, declare that other work is suspended for the duration, and have their call stand when a senior engineer or a director disagrees. Then measure it: what fraction of incidents had a named IC, how long from declaration to command, and how many distinct people commanded last quarter.
go deeper
Understand that in larger organisations the Incident Commander is often someone trained for the role from another team, rather than whoever owns the failing service.
Be able to explain why commanders are trained and rostered separately from service on-call, and what a shadow-then-lead qualification path looks like.
Argue the trade concretely: lost service intuition and training cost against a commander whose attention is undivided, and describe how you would size and refresh the pool.
Own the authority question — get the grant in writing from leadership before it is tested, define what an IC may override, and instrument the practice with named-IC coverage, time-to-command and distinct commanders per quarter.
## The two decisions that matter Everything else about an incident-command practice is documentation. The decisions with real cost are who may hold command, and what that person is permitted to do. ## Who commands The cheap option is that the on-call owner of the broken service commands. It costs nothing to set up and it is what happens by default in every organisation that has not decided otherwise. It is also the design that produces hero responses: the service owner is simultaneously your most valuable debugger and the person now responsible for coordination, comms and the clock. Under pressure, the debugging wins and command silently vanishes. The alternative is a **trained pool of commanders drawn from across teams**, from which an IC is assigned to an incident regardless of which service failed. This costs real money — training time, and the pool member's time during incidents in someone else's system — and it gives up the intuition a service owner would have had. What you buy is a commander whose attention is not competing with debugging, and who is structurally uninvested in the current theory. The sizing question is concrete. Command is a skill that decays without practice, so the pool must be small enough that members command often enough to stay sharp, and large enough that no one is permanently on the hook. In an organisation of a few hundred engineers with a handful of significant incidents a month, a pool of tens is defensible; a pool of five means five people carry every major outage and burn out inside a year. A useful review metric is simply how many distinct people held command last quarter — if it is three, you do not have a pool, you have three heroes and a document. Qualification should be a ladder, not a course completion: read the process, shadow a real incident as an observer, act as Scribe or comms on a live one, then command with an experienced IC beside you, then command alone. Deliberately include non-obvious people — support leads, engineering managers, product-adjacent engineers who are calm and organised. Command is a coordination skill, and treating it as a seniority reward both narrows the pool and reinforces exactly the misconception you are trying to kill. ## What authority the role carries Authority granted during an incident is authority you do not have. The organisation must decide, in advance and visibly, that for the duration of a declared incident the IC may: - **Page anyone**, on any team, at any hour, without their manager's permission. - **Suspend other work** for people pulled into the response, with no argument about sprint commitments. - **Make the final call** on the response, including over the objection of someone more senior — an engineer with a better theory, or a director who wants a different mitigation. - **Keep leadership at arm's length**, by routing all executive questions to the comms role rather than to responders. That last two are the ones organisations quietly fail to grant, and the failure is diagnosable in a single question: what happens when a VP joins the bridge and starts directing engineers? If the answer is "the IC defers," the role is decorative. The written grant matters because it lets a mid-level engineer commanding a SEV1 say "I have command, please route that through comms" to a director and be backed afterwards — and because the first time it is tested is exactly when nobody has capacity to litigate it. ## What you should not build Be honest about scale. Below roughly twenty engineers, a formal commander pool is overhead you will not sustain; the practice there is much simpler — name a commander out loud on every incident, and make sure it is not the person debugging. Building a heavyweight process ahead of the incident volume that justifies it produces a document nobody follows, which is worse than an informal norm people actually keep, because it makes the process itself untrustworthy. ## Keeping it alive The practice decays quietly, so instrument it. Three measures are enough and all are cheap: the share of incidents above a given severity that had a named, announced IC; time from declaration to command being taken, in minutes; and the count of distinct commanders per quarter. Combine that with periodic exercises so that people command in a low-stakes setting rather than meeting the role for the first time at 3am, and with a standing rule that whoever holds command is never the person with hands on production. If those three numbers are healthy and the role survives contact with an executive, the practice is real.
- What is the cost of taking command away from the service owner?You lose the intuition of someone who knows the system, and you pay for training plus pool members' time spent on other teams' incidents. You buy a commander whose attention is not competing with debugging and who is not invested in the current theory. On a small, homogeneous team the owner-as-IC trade can still be right.
- How do you tell whether the IC's authority is actually real?Ask what happens when a director joins the bridge and starts directing engineers. If the IC routes them to comms and keeps command, the grant is real. If the IC defers, the role is decorative and every difficult incident will quietly revert to whoever outranks the room.
- Would you build a commander pool for a 20-person startup?No. At that size a formal pool is overhead nobody sustains, and an unfollowed process is worse than an informal norm because it makes the whole practice untrustworthy. The version that scales down is one rule: name a commander out loud on every incident, and make sure it is not the person typing into production.
saying these in an interview costs you the question
- The most senior person available should command
- Whoever owns the broken service always commands
- Authority can be sorted out during the incident
- A commander pool of three or four is enough for any org
- Every organisation needs the full formal role structure