skip to content

During a postmortem your 5 Whys chain ends at "the config parser didn't validate the field." Why do experienced reviewers push back on stopping there, and what analysis do you do instead?

level: middleimportance: must knowfreq 65%

answer

  1. chain versus conjunction
  2. one line, many conditions
  3. what else had to be true
  4. group by phase: reach, detect, mitigate
  5. stops at the first patchable diff

basics

~20 s

5 Whys follows one causal line and stops at the first fixable defect, so it names the bug but not the review, test, rollout and detection gaps that let the bug reach production. Ask "what else" at each step and record multiple contributing factors.

solid answer

~50 s

5 Whys is a linear technique: each answer feeds one next question, so the chain converges on a single convenient cause — usually a code or config defect someone can patch this week. But the bug is only one of the conditions that had to hold. Something let it merge, something let it ship to 100% of traffic, something delayed detection for 40 minutes, and something made the blast radius global instead of one shard. Those are independent contributing factors, and fixing only the parser leaves all of them in place for the next defect. In practice I keep the questioning but branch it: at each step ask "what else had to be true?" and group the answers by phase — what allowed the defect in, what delayed detection, what delayed mitigation, what amplified impact. The output is a short list of factors, not a chain with one terminus.

go deeper

for a junior

Know that 5 Whys asks why repeatedly to get past the surface symptom, and be able to say that most outages have more than one cause working together.

for a middle

Explain why the technique's linear shape is the limitation, and demonstrate branching by naming what else had to be true besides the defect — test gaps, rollout policy, detection delay.

for a senior

Show that you grade factors by which control failed at which phase, and argue concretely why the detection and mitigation factors usually pay back more than the defect fix.

for a principal

Own the stop rule and the standard. Decide how deep analysis goes before it becomes unactionable philosophy, and how the org's postmortem template forces the detection and blast-radius questions to be answered every time.

## What 5 Whys actually is 5 Whys comes from Toyota's production system: ask "why?" about a defect, then ask "why?" about that answer, roughly five times, until you reach something worth fixing. It is genuinely useful as a habit — it stops teams from filing "the service went down" as an explanation. The problem is not the questioning. The problem is the shape it forces on the answer. ## The shape problem Each "why" takes exactly one input and produces exactly one output, so the technique can only ever produce a chain: A because B because C because D. Real production incidents are not chains. They are conjunctions — several conditions that all had to hold at the same time, most of which had been sitting there harmlessly for months. Take the config-parser example. For a validation gap to become an outage, all of the following also had to be true: - the change went through review and nobody read the parser path - there was no test that fed a malformed field to the parser - the config pushed to every region at once rather than to one region first - the failure mode was a crash loop rather than a rejected config with the previous value retained - the alert that fired was on a downstream symptom and took several minutes to route to a human Every one of these is a control that could have stopped the incident, and every one of them failed independently. A chain that terminates at "the parser didn't validate" names one of them and silently accepts the other five as normal. ## Why the chain lands where it lands The chain terminates wherever the person writing it feels the discomfort stop. That is almost always a concrete artifact with an owner and a diff — a missing null check, a wrong constant. It rarely stops at "our rollout policy pushes config globally in one step," because that is a policy decision someone made deliberately and defending it is uncomfortable. The search terminates at the first satisfying answer, not at the most useful one. ## Branching instead of chaining The fix is small and mechanical. Keep asking why, but at every node also ask **"what else had to be true for this to cause an outage?"** — and let the answer be a list. You end up with a tree rather than a line, and then you prune it. A useful way to organize the branches is by incident phase, because each phase maps to a different class of control: - **Allowed the defect to exist** — design, review, test coverage, type or schema gaps. - **Allowed it to reach production** — rollout policy, staging fidelity, canary population and duration. - **Delayed detection** — what was the time from first user impact to the page firing, and why. - **Delayed mitigation** — was the rollback path known, tested, and fast; did the responder have the access. - **Amplified impact** — shared fate, no bulkheads, retries from clients turning a partial failure into a total one. A postmortem that has at least one honest entry under detection and mitigation is almost always more valuable than one with a perfect account of the defect, because the defect is unique and the detection gap will apply to the next twenty incidents. ## Where to stop Branching has the opposite failure mode: infinite regress into "because we hired too fast" and "because the industry rewards velocity." Two practical stop rules: 1. Stop when the factor is no longer something your organization has a lever on within a quarter. "The upstream vendor's API has no idempotency key" is a factor you can act on (guard it, cache it, escalate it). "Distributed systems are hard" is not. 2. Stop when a factor stops being *specific*. A factor you could paste into any other postmortem unchanged is not a finding. ## What to say in an interview Say that 5 Whys is fine as a prompt and wrong as a model, name the conjunction versus chain distinction, and give one concrete example where two independent factors both had to hold. Then say what you would change: for the parser incident, validating config in CI *and* staging the config push regionally are separate fixes with separate costs, and a chain-shaped analysis would have bought only the first one.

  • Is 5 Whys ever the right tool, or would you drop it entirely?
    It is a fine prompt and a bad model. For a small, genuinely narrow event — one team, one component, minutes of impact — a short chain gets you to something actionable quickly and the ceremony of a full factor analysis is not worth it. What I would not do is let the chain define the document's shape for a serious incident, because a linear tool cannot express the conjunction of controls that a real outage requires.
  • How do you keep branching from producing an unfundable pile of findings?
    By pruning as part of the analysis rather than afterwards. A branch survives if changing that system property would plausibly have prevented or shortened the impact, if it generalizes past this one incident, and if someone has a lever on it this quarter. A well-pruned single-incident analysis usually lands at three to six factors spread across at least two phases.
  • What does an empty detection row in the phase grouping tell you?
    Almost always that the analysis stopped once the bug was found, not that detection was perfect. Detection is where the cheapest recurring wins live, because the next incident will have a different defect and the same monitoring. I go back and ask two questions explicitly: when did user impact actually begin, and what was the first signal that could have fired earlier than the one that did.

saying these in an interview costs you the question

  • Treats the chain's last node as the root cause
  • Believes every incident has exactly one root cause
  • Stops at the code defect and ignores detection and rollout gaps
  • Extends the chain to "the industry moves too fast" instead of branching
  • Assumes more whys means better analysis

context