skip to content

Ninety minutes into a SEV1, your most senior engineer is running the whole response alone — reading logs, applying changes, and answering executives directly — and nobody else on the bridge has been assigned anything. What is wrong with this, and what do you do about it?

level: seniorimportance: should knowfreq 50%

answer

  1. every role collapsed into one person
  2. state trapped in one head
  3. bus factor of one, mid-incident
  4. add structure without stopping the work
  5. sixty-second state dump, then write it

basics

~20 s

One person holding every role means the response has no coordination, no timeline, and a bus factor of one: all state is in their head. Take command, get a spoken state dump, write it down, split the work among named owners, and put someone between them and the executives.

solid answer

~50 s

This is hero debugging, and the problem is not the engineer — it is that every role has collapsed into one person. There is no shared state, so nobody can help without interrupting them; no timeline is being captured; the executives asking for an ETA are taxing the exact person who could shorten it; and if that laptop dies the response restarts from zero. Intervene without stopping their work: take command yourself, ask for sixty seconds of state — what is broken, what has been ruled out, what you are trying now — write it into the channel, then split the remaining threads to named owners and put a comms person between them and leadership. Keep them on the line they are closest to, timebox it, and ask the standing question they have stopped asking: is there something that restores users without knowing the cause yet?

go deeper

for a junior

Recognise the pattern: when one person is doing everything, nobody else can help and nothing is being recorded. Know that the fix is someone taking coordination off them.

for a middle

Explain the concrete failures — state trapped in one head, no timeline, executives interrupting the critical path, bus factor of one — and the ordered steps that add structure without halting the work.

for a senior

Show you can intervene in flight: bounded interruption, state written where others can read it, named workstreams, a comms shield, a timebox, and the question about mitigating without a diagnosis.

for a principal

Address why the organisation produced this — no trained IC pool, undrilled process, knowledge concentrated in one engineer, and incentives that celebrate the save — and what you would change so the next SEV1 does not depend on one person.

## What is actually broken here It is worth being precise, because the reflex answer — "they should have delegated" — is true but shallow. Four concrete failures are running simultaneously. **State exists only in one head.** Eleven other people are on the bridge and none of them can contribute, because contributing would require knowing what has been tried. Every offer of help costs the one productive person an interruption, so they stop accepting help, which deepens the problem. This is why the response has a room full of idle responders and still feels understaffed. **Nobody is steering.** Deep in a stack trace, the engineer is not tracking elapsed time, not weighing whether this line of investigation deserves another twenty minutes, and not asking whether a mitigation exists that does not depend on the diagnosis. The absorbed debugger is structurally the worst-placed person to abandon their own theory. **The interrupt shield is missing.** Executive questions are landing directly on the keyboard. Each one is a context switch on the critical path, and the answer they get — a guess made under pressure — then circulates as a commitment. **Bus factor of one, mid-incident.** No timeline, no shared state, no second person current on the problem. A dropped connection, a dead battery or simple exhaustion resets the response to zero ninety minutes in. That is an operational risk, not a theoretical one. ## Intervening without making it worse The genuine tension — and the reason the interviewer is asking — is that this person may be five minutes from the answer. Barging in with process ceremony can cost more than it saves. So intervene in a way that adds structure without stopping the work: 1. **Take command explicitly.** Say it: "I'm taking IC, keep doing what you're doing." You are removing the coordination load from them, not removing them from the problem. 2. **Ask for sixty seconds of state, once.** Three questions only: what is broken, what have you ruled out, what are you trying right now. One interruption, bounded, and it buys everything that follows. 3. **Write it where everyone can read it.** Post that state in the incident channel and keep it updated. The moment it is written, the other eleven people can act without touching the hero. 4. **Split into named workstreams.** The threads that were queued behind one person — checking whether the last deploy correlates, pulling the dependency's status, confirming the blast radius, preparing a rollback so it is ready if called for — go to named owners with a report-back time. 5. **Put someone in front of the executives.** A comms role, immediately. Leadership gets a real cadence from someone else, and the debugger gets silence. 6. **Start a scribe.** Even a rough timestamped log turns private progress into shared progress. 7. **Timebox the current line.** "You have fifteen minutes on that theory; at 14:20 we take the mitigation path." Said out loud, agreed, and revisited. ## The standing question The most valuable thing a commander adds to a heroic debugging session is the question the debugger has stopped asking: does anything restore users *without* the diagnosis? Someone chasing a root cause for ninety minutes has usually stopped separating "understand it" from "stop the bleeding," and those are different objectives with different clocks. Reintroducing that question is often worth more than any technical contribution you could make. ## Why it keeps happening Hero debugging is not usually a discipline failure by an individual; it is what an organisation gets when it rewards the story of the person who saved the day, when there is no trained pool of commanders so the strongest engineer is assumed to be in charge, when nobody has drilled the process so structure feels slower than just doing it, and when only one person actually knows the system — a knowledge concentration the incident merely exposed. Naming those causes, rather than the individual, is also what keeps the conversation on the systemic side when it reaches the review afterwards. And say the uncomfortable part plainly: the response probably still ended with the system fixed. Hero responses often work. They are unacceptable because they work unrepeatably, at the cost of one person, and leave nothing behind — no timeline, no shared understanding, and no second person who could do it next time.

  • Isn't interrupting the person closest to the answer exactly the wrong move?
    Which is why the interruption is bounded and paid for once: three questions, sixty seconds, and they carry on. You are not taking the problem away, you are taking the coordination, the executives and the note-taking away. What you buy is eleven people able to act, a timeline, and a decision-maker who is not invested in the current theory.
  • How would you frame this in the review afterwards without blaming the engineer?
    Point at the conditions, not the person. No trained commander pool, so the strongest engineer was assumed to be in charge; no drilled process, so structure felt slower than acting; and knowledge concentrated in one head, which the incident exposed rather than created. Any competent engineer in that position behaves the same way.
  • What if the hero resists — insists the handholding is slowing them down?
    Concede the point they are right about and hold the one they are not. They keep the technical line, uninterrupted; you take comms, assignment and the clock. Frame it as removing load rather than adding oversight. If they still refuse to share state, that is now an availability risk to the response and the IC decides, not the debugger.

saying these in an interview costs you the question

  • It worked, so nothing needs to change
  • The problem is that the engineer has a bad attitude
  • Everyone should just watch and stay out of the way
  • Assign the other responders to shadow the same investigation
  • Leadership deserves answers from whoever is closest to the fix

context