A draft postmortem lists the root cause as "human error — the operator ran the wrong command." Why would an SRE reject that, and what should the document say instead?
answer
- carefulness is not a deployable control
- read it forwards, not backwards
- strike should have and failed to
- reachable, plausible, undetected
- training sits lowest in the hierarchy
basics
~20 s"Human error" is where an investigation stopped, not an explanation. Treat it as a symptom and record what made the wrong command reachable, plausible and undetected — no confirmation prompt, identical staging and production shells, a stale runbook, time pressure — and fix those.
solid answer
~50 sI would send it back, because "human error" ends the investigation exactly where the useful findings start. The operator's action is the last event in the chain, not the reason the chain existed. So I ask three questions and write the answers down: what made the wrong command *reachable* — was the destructive path the default, with no confirmation and no dry run? What made it *plausible* — did production and staging look identical, was the runbook stale, was this the third page of the shift? And what made it *undetected* — why did nothing catch it for eleven minutes? That is the second story: the world as it looked from inside the situation, before anyone knew the outcome. I also strike counterfactual phrasing — "should have checked", "failed to notice" — because it describes a world that did not happen and quietly reassigns the fault the document claims to avoid.
code
bash · 18 lines#!/usr/bin/env bash
set -euo pipefail
# Unguarded: silent about its target, no confirmation, destructive.
delete_data() {
rm -rf "${DATA_DIR:?DATA_DIR must be set}"
}
# Guarded: states the target and host, requires a typed match before acting.
delete_data_guarded() {
local target="${DATA_DIR:?DATA_DIR must be set}"
printf 'DELETE %s on host %s\n' "$target" "$(hostname)"
read -r -p 'Type the directory path to confirm: ' confirm
[[ "$confirm" == "$target" ]] || { echo 'aborted'; return 1; }
rm -rf "$target"
}
delete_data_guardedgo deeper
Know that "human error" is never accepted as a final answer, and that the follow-up question is what the system made easy to do wrong.
Walk through the replacement analysis out loud: what made the action reachable, plausible and undetected, and which concrete guardrail addresses each. Name hindsight bias and counterfactual phrasing.
Demonstrate that you edit real drafts — rejecting a first-story finding in review, rewriting counterfactual sentences, and pushing remediation up the hierarchy of controls instead of accepting a training item.
Argue about where the organisation's controls budget goes: which hazards get engineered away, which stay procedural because the engineering cost exceeds the exposure, and how you keep that decision explicit rather than accidental.
## Why the finding is unusable A postmortem exists to change something. "The operator ran the wrong command" changes nothing: the only remediation it suggests is that the operator be more careful, and carefulness is not a control you can deploy, monitor or verify. Worse, the finding is almost always true and almost always irrelevant — a human did indeed act, and saying so tells the next team nothing they can use. The rule of thumb is blunt: **if the only lever your root cause offers is a person's future attentiveness, you have not finished the analysis.** ## Hindsight bias, and the second story After an outage you know which of fifteen signals mattered. The operator did not. Reading the incident backwards from the outcome makes every missed cue look obvious and every decision look negligent, which is why safety literature insists on reconstructing the *second story*: the situation as it appeared from inside, with only the information available at that moment. The first story is "they ran the wrong command." The second story is "the runbook's copy-paste block still referenced the old namespace, the prompt in both shells is green, and the command has no dry-run flag." A reliable tell that a draft is still stuck in the first story is **counterfactual language**: *should have*, *could have*, *failed to*, *didn't bother to*. Each of these describes a world that did not happen and implies a choice that was never consciously available. Replace them with what the person actually perceived and did: - Before: "The on-call engineer failed to check which cluster the shell was pointed at." - After: "The shell prompt does not display the target cluster; the engineer had switched context nine minutes earlier while handling a separate alert." The second version is more specific, not vaguer — blamelessness raises the precision bar rather than lowering it. ## The three questions that replace the finding **Reachable.** Was the destructive action available on the default path, without friction proportional to its blast radius? Dangerous operations should be harder to perform than safe ones: an explicit typed confirmation, a dry-run default, a required second approver, a scoped credential that simply cannot address production. **Plausible.** What made the wrong action look right at that moment? Production and staging that are visually indistinguishable. A runbook whose steps drifted from reality two releases ago. An alert whose text names the wrong service. Pressure — the outage was already active, and speed felt cheaper than verification. **Undetected.** Why did the system not catch it? Was there no guard on the action, no confirmation of scale before deletion, no alert on a sudden drop in row counts, no soft-delete window that would have made recovery cheap? These three questions convert a dead-end into a list of concrete, ownable changes. ## Prefer stronger controls than training Safety practice ranks controls roughly as: eliminate the hazard, substitute something safer, engineer a guardrail, then fall back on procedure and training. "More training" and "be careful" sit at the bottom because they decay, do not transfer to new hires, and fail exactly when the system is under stress. If a postmortem's remediation list is dominated by documentation and training items, that is a signal the analysis stopped one layer too early. ```bash #!/usr/bin/env bash set -euo pipefail # The action a postmortem would label "human error": # destructive, silent about its target, no confirmation. delete_data() { rm -rf "${DATA_DIR:?DATA_DIR must be set}"; } # The system fix: the dangerous action states its target and demands a typed match. delete_data_guarded() { local target="${DATA_DIR:?DATA_DIR must be set}" printf 'DELETE %s on host %s\n' "$target" "$(hostname)" read -r -p 'Type the directory path to confirm: ' confirm [[ "$confirm" == "$target" ]] || { echo 'aborted'; return 1; } rm -rf "$target" } ``` ## What about genuinely poor judgement? It happens, and pretending otherwise damages credibility. The response is still not to write "human error" in the document. If someone knowingly disregarded a substantial, understood risk, that is a management conversation with their manager, held privately and on its own evidence. The postmortem still records what the system permitted, because the next person to be pressed for time deserves the guardrail regardless of what the last person's intentions were. ## How to answer this in an interview Interviewers use this prompt to separate people who have read about blameless postmortems from people who have edited one. Show the edit: quote the bad line, name hindsight bias and counterfactual phrasing, then produce the replacement findings and the corresponding controls. Mentioning that you would *reject the draft in review* — that this is a normal editorial act, not an accusation against the author — lands particularly well.
- Which words would you strike from a postmortem draft, and why?Counterfactuals: should have, could have, failed to, neglected to. Each describes a world that did not happen and implies a choice the person never consciously faced, which smuggles blame back into a supposedly blameless document. Replace them with what was actually perceived and done — "the prompt showed no cluster name" rather than "failed to verify the cluster".
- Is "the operator was undertrained" an acceptable system finding?It can be a real one, but it is the weakest useful control. Training decays, does not transfer to the next hire, and fails under exactly the pressure that caused the incident. Prefer, in order: remove the hazard, substitute a safer operation, engineer a guardrail, and only then adjust procedure and training. A remediation list that is all documentation and training usually means the analysis stopped early.
- What if the engineer genuinely ignored a documented runbook step?Then ask why the step was ignorable — was it slow, wrong, unfindable, or routinely skipped by everyone? Widespread skipping is drift and a systems finding. A single person knowingly disregarding an understood, substantial risk is possible but rare, and it is handled by their manager privately; the postmortem still records what the system permitted.
saying these in an interview costs you the question
- Human error is a legitimate root cause
- The fix is to tell people to be more careful
- The operator should have double-checked the target
- Add a warning to the runbook and close it out
- Blameless means we cannot mention what the operator did