After a customer-visible outage, an executive asks who was responsible and wants consequences. How do you respond as the engineering lead?
answer
- they want assurance, not a name
- answer with controls and dates
- punishment exports the knowledge
- judge behaviour, not outcome
- pre-agree the line before the outage
basics
~20 sGive the executive what they actually want — confidence it will not recur — rather than a name. Present the missing controls, the fixes with owners and dates, and the cost of punishing: you lose your best-informed engineer and teach everyone else to hide the next one. Reserve genuine recklessness for a private management track.
solid answer
~50 sI do not refuse the question, because refusing sounds like protecting the guilty. I reframe it around what the executive is really buying: assurance this does not happen again. So I bring the second story — what the system permitted, what made the wrong action look right, and why nothing caught it — plus the specific guardrails now landing, with owners and dates. Then I make the cost of the alternative explicit: dismissing or publicly punishing the operator removes the person who now understands this failure better than anyone, and it teaches every remaining engineer that the safe move is to stay quiet, which lengthens the next outage. I also concede the line honestly. If someone had knowingly disregarded a substantial, understood risk, that is a management matter, it is handled privately by their manager on its own evidence, and it never runs through the learning document. Then I offer a follow-up date on the remediation, because that is real accountability.
go deeper
Know that this conversation is not yours to have alone — escalate to your lead, and be able to say why the postmortem itself stays focused on the system.
Explain what you would offer instead of a name: what the system permitted, the guardrail now landing, and who owns it. Avoid answering with a slogan about blamelessness.
Make the cost of punishment concrete — losing the best-informed engineer, buying silence that lengthens the next outage — while conceding honestly that reckless behaviour is handled privately by management.
Own the pre-agreement: negotiate with leadership before an expensive outage what blamelessness covers, where the recklessness line sits, what leadership gets instead of a name, and how external or regulatory records are kept separate from the internal learning document.
## Read the question underneath the question An executive asking "who was responsible?" is rarely enjoying the prospect of punishment. They are usually holding something else: a customer or regulator asking them a question they cannot answer, a board meeting on Thursday, or a plain fear that the organisation does not have this under control. "Nobody is to blame" answers none of that and sounds evasive. The move is to answer the underlying need — *what changed so this cannot recur, and how will you know* — with more specificity than they expected, and to do it immediately rather than after a defence of process philosophy. ## What to bring into that conversation 1. **A crisp account of what the system permitted.** Not "an engineer made a mistake", but "a destructive command was available on the default path with no confirmation, production and staging were visually identical, and nothing detected the deletion for eleven minutes." Executives generally find this *more* reassuring than a name, because it is something an organisation can fix. 2. **The controls, with owners and dates.** Concrete and few. A short list of shipped guardrails beats a long list of intentions. 3. **Detection.** What will catch it faster next time, and by how much. Leaders anchor on time-to-recovery even when they ask about fault. 4. **A follow-up commitment.** Offer to come back on a specific date with evidence the work landed. That is accountability with a name on it, which is what was being asked for in the first place. ## Making the cost of punishment explicit The argument that lands is economic rather than moral: - **You lose the best-informed person.** After an incident, the operator involved understands that failure better than anyone in the company. Removing them exports that knowledge to a competitor and leaves the replacement to rediscover it. - **You buy silence for years.** Every engineer watching learns what candour costs. The next incident gets reported later, described more vaguely, and takes longer to resolve — and none of that appears on any dashboard as a consequence of this decision. - **You get the wrong fix.** A punished team produces postmortems designed to be safe to publish rather than accurate, and the guardrails that would have prevented the recurrence never get identified. - **The individual is a poor predictor.** Whoever happened to be on call that night is largely a scheduling artefact. The condition that permitted the failure is the durable variable; the person is not. ## Concede the line honestly Credibility here depends on not overclaiming. Blamelessness is not a promise that behaviour never has consequences. Just-culture models separate honest error, at-risk behaviour where someone drifted into a shortcut without recognising the risk, and reckless behaviour where someone knowingly disregarded a substantial, understood risk. Only the first two are learning material. The third exists, is handled privately by the person's manager on its own evidence, and stays out of the postmortem — and saying so plainly is what stops an executive suspecting that "blameless" is a shield. The corollary matters just as much: the category is judged by the behaviour, not by how expensive the outcome was. If the same shortcut would have drawn no reaction on a quiet day, punishing it because it happened to be costly on a loud one is a lottery, and teams read lotteries as arbitrary. ## Repeat involvement If the same individual keeps appearing at the centre of incidents, that is data, not a verdict. The likely readings are structural: they own the least-defended part of the system, they are the only person who touches it, they are exhausted from carrying a disproportionate share of the pager, or the tooling they use is the worst in the estate. Any of those is an organisational finding. Genuine individual performance concerns exist too, and they belong to that person's manager, privately, with normal evidence — never as an output of a postmortem process people are supposed to trust. ## Do this before you need it The conversation goes far better if leadership agreed to the model in advance: what blameless does and does not cover, where the recklessness line sits, and what leadership will be given instead of a name. Negotiating that during a costly outage, with a customer on the phone, is the worst possible moment. Pre-agreement is the actual principal-level work; the conversation described here is just the moment it gets tested. Also decide in advance what goes outside the company. Regulated and contractual contexts may require a formal record with named individuals, and that record is kept separately from the internal learning document by design — conflating the two is how internal postmortems become sanitised into uselessness. ## What a strong interview answer sounds like Do not perform indignation on behalf of blamelessness. Show that you can hold a hard line while giving a senior stakeholder something real: the system finding, the shipped controls, the detection improvement, the explicit cost of the punitive alternative, an honest concession about where recklessness is handled, and a date to report back.
- What if the same engineer has been involved in three incidents this quarter?Treat it as data about the system first. The usual explanations are structural: they own the least-defended service, they are the only person who touches it, or they are carrying a disproportionate pager load. Each is an organisational finding. A genuine performance concern is possible, but it belongs to their manager privately with ordinary evidence, never as an output of the postmortem process.
- How do you handle a regulator or customer contract that requires naming an individual?Keep two records. The external, formal account satisfies the obligation on its own terms, and the internal learning document stays separate and candid. Conflating them is how postmortems get lawyered into uselessness. Where legal exposure is real, agree the boundary with legal in advance rather than deciding it mid-incident.
- How do you get leadership to accept this before an expensive outage tests it?Agree the model in writing while nothing is on fire: what blamelessness covers, where the recklessness line sits, and what leadership receives instead of a name — the system finding, the controls with owners and dates, and a follow-up review. Then honour the follow-up visibly. Delivering that report on time is what buys the credibility you will spend during the next outage.
saying these in an interview costs you the question
- Nobody is responsible, that is what blameless means
- Refusing to answer the question at all
- Naming the on-call engineer to close the conversation
- Punishing the shortcut only because it was expensive
- Handling performance concerns inside the postmortem