How do you choose between a why-chain, a fishbone diagram and a fault tree for one tracked product defect?
answer
- Match the shape to the failure
- One line, many categories, many routes
- Chain assumes a single dominant path
- Tree assumes any one route suffices
- A diagram widens the search, proves nothing
basics
~20 sMatch the technique to what you already know. A why-chain fits a single causal line you can walk back, a fishbone diagram fits an unknown cause spread across categories, and a fault tree fits a failure several routes can each produce.
solid answer
~50 sEach technique encodes an assumption about the failure, so pick the one whose assumption matches what the evidence already supports. A **why-chain** asks *why* of each answer in turn and assumes one dominant causal line; it is the cheapest and it is honest when the failing sequence is already known end to end. A **fishbone diagram** hangs candidate causes off labelled branches — code, data, configuration, environment, interface, process — and assumes you do not yet know where to look, so it widens the search and proves nothing. A **fault tree** decomposes the failure into routes marked *any one of these is enough* or *all of these together*, and assumes several routes reach it. In practice they compose: widen with categories, then chain down the branch the evidence supports. And none of them helps until the defect reproduces.
code
pseudocode · 18 linesGIVEN one tracked product defect with a confirmed reproduction
IF the failing sequence is known end to end
USE why_chain(failure) # one line; spend effort on evidence per link
ELSE IF nobody can say where the cause sits
USE cause_and_effect_diagram(
failure,
categories = [code, data, configuration,
environment, interface, process])
THEN why_chain(surviving_branch) # the diagram only widens the search
ELSE IF the same failure has been reached by different routes
USE fault_tree(top_event = failure_stated_precisely)
# each node marked: any-one-of | all-of
IF NOT reproduces(defect):
STOP - reproduce first, or record that the analysis rests on one observationgo deeper
Be ready to describe what each of the three looks like on paper: a chain of why-questions, a diagram sorting candidate causes into categories, and a failure broken into the routes that reach it. Knowing the shapes and having sat in a session is enough here.
Explain what each shape assumes about the failure and choose one out loud for a defect you actually worked on. An interviewer wants the assumption named — one dominant line, an unknown location, or several routes — not three definitions recited in order.
Show judgement about cost. Say when a chain is the cheapest honest answer, when widening the search first saves a wasted day, and when you would refuse to run any technique at all because the defect does not yet reproduce reliably.
Own how much of this a team standardises. Argue one default shape everyone can run unaided against a menu that fits more failures but gets applied inconsistently, and say which recurring classes of defect justify the heavier shape.
## Why the shape is the choice A root-cause technique is not a procedure that finds a cause. It is a **shape you impose on the explanation**, and each shape encodes an assumption about how the failure came about. Pick a shape whose assumption is false for your defect and the analysis still terminates — with a confident, wrong-shaped answer that everyone signs. That is why the choice matters at all, and why the honest first move is to ask what you already know about the failure rather than which technique the team reaches for by habit. Three shapes cover almost everything a team meets on a tracked product defect from a test cycle or the field. ## The three shapes and what each assumes **Iterative why-questioning** takes the observed failure and asks *why* of each answer in turn, producing a single line of links. Its assumption is that **one dominant causal line** runs back from the failure and that each step has a single predecessor worth naming. It is the cheapest of the three — a whiteboard and twenty minutes — and it is honest when the failing sequence is already known end to end and you are hunting for the earliest point that was under someone's control. **A category-based cause-and-effect diagram**, drawn as a fishbone, puts the failure at the head and hangs candidate causes off labelled branches. Its assumption is the opposite one: **you do not yet know where the cause sits**, so you need a prompt to look where the first guess did not. The classic category set comes from manufacturing — machine, method, material, measurement, people, environment — and software teams re-cut it into the places a defect actually originates: code, data, configuration, environment, the interface to another team's service, and the process around the change. Crucially the diagram **widens a search and proves nothing**. A branch is a candidate, never a finding. **A fault tree** starts from the failure as a single precisely stated top event and decomposes it downward into the routes that reach it, marking at each node whether *any one* of the branches below is enough on its own or whether *all of them* have to coincide. Its assumption is **multiplicity**: the same failure is reachable more than one way. It is the most expensive of the three and it demands a sharp top-event statement, because a vague failure produces a tree of vague branches. ## Side by side | Technique | Shape on paper | Assumes | Cheapest when | Breaks down when | | --- | --- | --- | --- | --- | | Why-questioning | one line of links | a single dominant path back | the failing sequence is already known | several routes contribute equally | | Cause-and-effect diagram | one head, many labelled branches | the cause could be anywhere | nobody can say where to look | it is treated as proof rather than search | | Fault tree | a top event decomposed into routes | several routes reach the same failure | the failure recurs by different routes | the top event is stated vaguely | ## Choosing, in practice 1. **Does it reproduce reliably?** If not, no shape helps yet. Get a reproduction first, or state plainly that the analysis rests on one unrepeatable observation. 2. **Do you know the failing sequence end to end?** Walk it as a chain and spend your effort on the evidence for each link rather than on the drawing. 3. **Can nobody say where to look?** Widen with categories first, kill the branches the evidence rules out, then chain down whichever branch survives. 4. **Have different reports reached the same failure by different routes?** Decompose into alternatives, or you will fix the route you happened to be shown. 5. **Is this a class of related defects rather than one defect?** Start wide. A chain drawn on the most recent occurrence explains the most recent occurrence and nothing else. ## They compose, and switching is not indecision Most real analyses use two shapes. Categories generate candidates and a chain is walked down the one branch the evidence supports. Or a chain reaches a link with two possible predecessors and quietly becomes a small tree at that node. Changing shape mid-analysis is the analysis telling you the assumption you started with was wrong, and continuing in the original shape to keep the drawing tidy is how a wrong-shaped answer ships. ## Where each one goes wrong - **A chain on a multi-route failure** picks one route silently. The fix lands, the failure recurs by the other route, and the second report looks like a duplicate. - **A diagram used as proof.** Filling every category makes the page look complete and settles nothing; the diagram's own output is a list of things to go and check. - **A tree with a soft top event.** "The product was slow" cannot be decomposed; "checkout failed to respond within the agreed limit for a signed-in customer" can. - **Counting instead of reasoning.** The number in the popular name of the why-technique is a mnemonic for keeping going, not a target to hit or a place to stop. - **Running any of them on an unreproduced defect** produces a shape full of guesses, and the shape makes the guesses look like structure.
- When would you draw a fishbone diagram first and a why-chain second?When nobody can yet say where the cause sits. The diagram widens the search: you list candidates under categories such as code, data, configuration, environment, interface and process, then kill the branches the evidence rules out. Once one branch survives, switch to a chain and walk it link by link. Going straight to a chain there just formalises the first guess.
- How do the fishbone categories change for a software defect?The original categories come from manufacturing and are usually re-cut into the places a software defect actually originates: code, data, configuration, environment, the interface to another team's service, and the process around the change. The categories exist only to prompt a look somewhere you would not have looked; a category that has yielded nothing in a year should be dropped rather than filled for tidiness.
- What do you do when a chain reaches a link with two possible predecessors?Stop treating it as a chain at that node. Record both predecessors as branches, mark whether either alone is enough or whether both must coincide, and gather evidence for each before choosing. Forcing the drawing back into one line to keep it tidy is how the route you did not investigate survives the fix.
Choosing between them is like choosing a search pattern: walking back along a trail you can already see, sweeping the areas nobody has looked at yet, or mapping every route that reaches the same clearing.
saying these in an interview costs you the question
- Treats five as a required number of whys
- Picks the technique by habit, not by what is known
- Uses a cause-and-effect diagram to prove a cause
- Assumes every defect has exactly one root cause
- Draws one chain when reports differ in reproduction
- Starts before the defect reproduces reliably