skip to content

Which fixed defects earn a root-cause investigation, and why do blanket always-or-never rules fail?

level: middleimportance: must knowfreq 58%

answer

  1. Not every defect carries a lesson
  2. What did missing this one cost?
  3. Escape route, repetition, blind detection
  4. Both extremes replace a decision with a constant
  5. Written conditions, checked at close

basics

~20 s

Escaped, repeating, costly and undetectable defects earn an investigation; a typo caught in review does not. Blanket rules fail alike: investigating everything turns the work into a form nobody reads, investigating nothing pays for the same weakness again.

solid answer

~40 s

Selection is a decision, not a reflex. **The facts that earn an investigation**: the defect reached a customer, so every internal stage passed it; the same shape has now closed three times; the escape damaged data, money or trust; or nobody can name the check that should have caught it. Fix size is not a signal, and a one-line fix can hide the most instructive escape. **Both blanket rules fail for the same reason** - they decouple effort from the lesson available. Investigating everything spreads a fixed pool of attention evenly, so the defect that carried a lesson gets the same shallow pass as a typo. Investigating nothing spends nothing and re-learns the weakness by paying for it again. Write the conditions down, check them when the defect closes, and allow a recorded override.

code

pseudocode · 13 lines
pseudocode
# checked once, when the defect record is closed
function earnsInvestigation(defect):
    if defect.foundBy == CUSTOMER and defect.harm in {DATA, MONEY, SAFETY}:
        return true
    if countClosedWithSameShape(defect.shape, window = 90 days) >= 2:
        return true
    if defect.checkThatShouldHaveCaughtIt == UNKNOWN:
        return true                    # the output would be a check that does not exist
    return false                       # fixed, closed, no further effort

# an override is allowed in both directions, with a reason on the record
if reviewer.overrides:
    record(defect, decision = reviewer.decision, reason = reviewer.reason)

go deeper

for a junior

Be ready to say that fixing a defect and explaining it are two different jobs, and that only some defects get the second one. Knowing one or two selection signals - it reached a customer, it has happened before - is enough at this level.

for a middle

Expect to list the selection signals and explain why fix size is not among them. Be able to say when the rule is applied - at closing time, when the escape route is finally known - and what the output of an investigation is.

for a senior

An interviewer expects you to attack both extremes and show they fail identically, and to describe a rule cheap enough that a busy team actually applies it: written conditions, a recorded override, a cap on concurrent investigations.

for a principal

Own the economics. Attention is a fixed budget, so selection is really allocation: which defects buy the most durable change per hour spent, and how you notice when the conditions have drifted too tight or too loose.

## The decision that comes after the fix A defect is reproduced, fixed, verified and closed. A separate decision follows, and it is the one this question is about: does anyone now spend further effort explaining **how the defect came to exist and how it survived every check between the change and the customer**? That effort is not free. It costs an hour or more from the people who hold the evidence, and it usually ends in a change that costs more again. Everything has a cause, so *is there a cause to find* is never the question. The question is whether **this** cause is worth the hour, and for most defects the honest answer is no. Teams get this wrong in two symmetrical ways, and both are wrong for the same reason - which is the second half of the question. ## The facts that select a defect No single fact decides it. A short list, checked when the defect record is closed, does: - **How far it escaped.** A defect a customer met passed every stage that was meant to stop it. That escape route is a fact about the team's own checks rather than about the defect, and it is the most reusable thing an investigation can produce. - **Repetition.** The third defect of the same shape is evidence of a standing weakness. One is an event; three is a pattern with an address. - **What the escape cost.** Corrupted data, money moved wrongly, a safety-relevant behaviour, a large affected population. The economics of finding defects late are a subject of their own; here the cost only supplies the motive for spending the hour. - **Severity, strictly as an input.** A recorded severity is a signal, not the decision. A crash on a path almost nobody reaches may teach less than a quietly wrong figure in a report thousands of people trust. - **Detection blindness.** Nobody can name the check that should have caught it. This is usually the most valuable investigation available, because its output is a check that does not yet exist. | Fact about the defect | Selects it? | Why | | --- | --- | --- | | The fix was a single line | No, by itself | Fix size measures the code, not the lesson | | A customer found it first | Yes | Every internal stage passed it unnoticed | | Third of this shape this quarter | Yes | A repeated weakness has one address | | It was hard to reproduce | Weak signal | Difficulty is interesting, not consequential | | Recorded at the top severity | Input only | Severity ranks the harm, not the learning | ## Why "investigate everything" fails Attention is finite and roughly constant. A rule that investigates every closed defect spreads that fixed pool evenly across defects carrying no lesson and defects carrying a large one, so the instructive escape gets the same shallow pass as a misspelled label caught in review. The consequences are predictable: the write-up becomes a form; the cause field fills with restatements of the symptom; nobody reads the output because most of it is empty; and the practice acquires a bad name, so the one investigation that mattered is resented along with the rest. Volume also destroys the signal - a hundred recorded causes a month cannot be compared, grouped or acted on by anyone. ## Why "investigate nothing" fails the same way The mirror rule looks efficient. Each defect closes the moment the symptom stops, and nobody asks how it escaped. The same weakness then fires again, and the team pays the fix cost repeatedly plus the escape cost each time. Nothing accumulates: the checks never grow, and the team's knowledge of its own blind spots stays exactly where it was. The two failures are one failure. **Both replace a decision with a constant.** One spends a fixed amount on every defect regardless of what could be learned; the other spends zero on every defect regardless of what could be learned. Effort should track expected learning, which varies enormously between defects, and neither rule lets it vary at all. That is why the interviewer usually asks the two halves together: a candidate who only attacks the lazy end has not noticed that diligence fails identically. ## Making the rule cheap enough to actually apply 1. **Write the conditions down** - four or five of them, in one place, in plain language. A rule that lives in someone's head is re-argued from scratch on every defect. 2. **Check them when the defect closes**, not when it is filed. At closing time the escape route and the size of the change are both known; at filing time neither is. 3. **Allow an override with a recorded reason**, in both directions. Judgement beats a checklist, but an unrecorded override is indistinguishable from drift. 4. **Cap the investigations running at once.** Two open and finished beats ten open and abandoned. 5. **Revisit the conditions periodically.** If nothing has been selected in three months they are too tight; if almost everything is selected, they are too loose.

  • A defect took one line to fix, but nobody can explain how it reached customers. Does it earn an investigation?
    Yes. The size of the fix measures the code, not the lesson. The question worth an hour is not why the line was wrong but why every stage between the change and the customer passed it. That answer produces a missing check, which outlives the fix.
  • How do you stop the selection rule being re-argued on every defect?
    Write four or five conditions in one place, apply them at closing time, and let anyone override in either direction as long as the reason lands on the record. Reviewing the recorded overrides every few months tells you whether the conditions or the judgement needs adjusting.
  • Who should apply the rule, and when?
    Whoever closes the defect, at the moment they close it, because that is when the escape route and the size of the change are both known. Filing time is too early - the reporter usually knows the symptom and nothing about how it got out.

Auditing every household expense stops being an audit and becomes filing; auditing none is how a small leak runs for a year unnoticed. Both spend a constant instead of deciding.

saying these in an interview costs you the question

  • Every closed defect deserves a full written cause
  • Only crashes and outages are worth investigating
  • The fix was one line, so there is nothing to learn
  • Investigate whatever the loudest customer complained about
  • Repetition is bad luck, not a selection signal
  • Selection is a judgement call that cannot be written down