skip to content

Should a defect's root-cause investigation run before the fix, after it, or after release, and who owns it?

level: seniorimportance: should knowfreq 45%

answer

  1. Ask what the fix will destroy
  2. Preserve in minutes, explain in hours
  3. The change required is itself evidence
  4. Memory decays fast after a release
  5. One named owner, never the team

basics

~20 s

Perishable evidence decides. Investigate before the fix only when the failing state is the evidence; otherwise preserve it, fix, then explain, because the change the fix required is itself evidence. Ownership follows the evidence: the fixer plus the context holder.

solid answer

~40 s

Let **evidence perishability** decide. Investigate before the fix only when the failing state is the evidence - data about to be corrected, configuration about to be replaced, a transient condition nobody can recreate. Otherwise preserve that material and fix first, because delaying a customer-visible repair to write an explanation is rarely defensible. The default home is **after the fix, before the release**: the change the fix required is itself evidence of the weakness, the reproduction is confirmed, and the people involved still remember. Leave until after release only the questions the field answers - how far the defect actually spread, and which stage should have stopped it. Ownership follows the evidence: the engineer who made the fix, plus whoever holds the operational or design context, with one named person accountable rather than the team.

code

pseudocode · 12 lines
pseudocode
onDefectConfirmed(defect):
    if fixWouldDestroy(defect.failingData, defect.liveConfig, defect.transientState):
        defect.preserved = snapshot(defect.failingData, defect.liveConfig, defect.inputs)
    applyFix(defect)
    verifyFix(defect)
    close(defect)

onDefectClosed(defect):
    if earnsInvestigation(defect):
        evidence = defect.preserved plus defect.changeRequiredByFix
        owner    = defect.fixer plus holderOfContext(defect.area)
        schedule(investigation, before = nextRelease, timebox = 90 minutes)

go deeper

for a junior

Know that fixing a defect can destroy the information needed to explain it, and that when that is likely you save the failing inputs and configuration before you change anything.

for a middle

Be able to argue the default - explain after the fix but before the release - and say why: the change the fix required is evidence, the reproduction is confirmed, and everyone still remembers.

for a senior

Show the read you actually make: how fast is this evidence disappearing, what does the delay cost the people affected, and who holds the context the fixer does not. Name one accountable owner rather than a group.

for a principal

Own the policy and its decay. Decide what may be deferred past a release at all, what a deferred item must carry to be real, and how you notice that deferred investigations have quietly become investigations nobody does.

## The variable that decides placement The instinct is to argue about placement in terms of process discipline - explain before you patch, or patch before you explain. That framing does not survive contact with a real defect. The variable that actually decides is **how quickly the evidence disappears**, weighed against **what the delay costs the people affected**. Some evidence perishes the moment the fix lands: the rows in the wrong state, the configuration that is about to be replaced, the transient condition that only exists while the system is misbehaving. Other evidence gets **stronger** after the fix: the change the fix required is a precise statement of what was wrong, and a confirmed reproduction proves you understood the defect rather than merely quieted it. ## Before the fix Justified when the failing state is the only evidence there will ever be. Signs: nobody can reproduce it on demand; the misbehaviour depends on data that is about to be corrected; the condition is environmental and about to be recycled. Even then, the answer is usually not "delay the fix". It is **preserve, then fix**: copy the offending inputs, the surrounding records, and the configuration as it stands, so the explanation can be written later against material that no longer exists in the live system. Preserving takes minutes; a full investigation takes hours, and holding a customer-visible repair open for hours to write one is hard to defend. The owner here is whoever is already holding the failing state - the person operating or debugging the system - because they are the only one who can preserve it in time. ## After the fix, before the release This is the default home for most defects that earn an investigation at all, for three reasons: - The **change the fix required** is evidence. What had to be altered says more about the weakness than any recollection of the symptom. - The **reproduction is confirmed**, so the explanation is anchored to a demonstrated mechanism rather than a theory. - **Memory is intact.** The people who wrote the change, reviewed it and reported the symptom are all still available and still care. The owner is the engineer who made the fix, joined by whoever holds context the fixer lacks - whoever designed the area, whoever operates it, whoever wrote the check that did not fire. One of them is named as accountable for the output; "the team owns it" reliably produces no output at all. ## After the release Reserve this for the questions only the field can answer: how far the defect actually spread, which population met it, and whether the new check demonstrably fires on the next occurrence. It also suits class-level work, where several defects have to accumulate before a shape is visible. The cost is real - attention decays fast after a release, and an investigation scheduled for later frequently becomes an investigation never done. Treat a post-release slot as a commitment with a date and an owner, not an intention. | Placement | Evidence you get | Evidence you lose | Natural owner | | --- | --- | --- | --- | | Before the fix | Live failing state, exact inputs | Nothing yet, but the repair waits | Whoever holds the failing state | | After the fix, pre-release | The change required, confirmed reproduction | The transient state, unless preserved | The engineer who fixed it | | After release | Actual spread, which stage missed it | Attention, memory, and often the slot | Whoever owns that area | ## A default that survives contact 1. On confirmation, ask one question: **will fixing this destroy evidence?** If yes, spend ten minutes preserving it. If no, proceed. 2. Fix, verify, close. 3. At closing time, apply the selection conditions. Most defects stop here and that is correct. 4. For the selected ones, investigate before the release while the change and the memory are both fresh, with one named owner and a timebox. 5. Push to after the release only what the field genuinely answers, with a date attached. ## The mistake at each end Investigating everything before the fix looks rigorous and is not: it converts every defect into a stoppage, and the pressure that creates eventually collapses into skipping the explanation altogether. Deferring everything until after the release looks pragmatic and is not either: by then the change is one of many, the reproduction has been dismantled, and the person who knew why has moved on to the next thing. Placement is not a matter of principle - it is a read of how fast this particular evidence is disappearing.

  • Someone argues the fix must wait until the cause is fully explained. How do you answer?
    Separate preservation from explanation. Preserving the failing inputs, data and configuration takes minutes and protects everything the explanation will need; the explanation itself takes hours. Holding a customer-visible repair open for hours to write prose is rarely defensible, and the pressure it creates is what eventually kills the practice.
  • Why do post-release investigations so often never happen?
    Attention moves with the release. The change becomes one of many, the reproduction is dismantled, and the people who knew are on the next piece of work. Anything deferred past the release needs a date, a named owner and a specific question it will answer, or it is an intention rather than a commitment.

Evidence around a failure behaves like a footprint in soft ground: it is legible while the ground is still fresh, and gone once the repair crew has raked over it.

saying these in an interview costs you the question

  • No defect may be fixed before its cause is explained
  • Always explain it later, after the release
  • The fix itself tells you nothing about the weakness
  • Anyone can write it up whenever they have time
  • Preserving evidence and investigating are the same step