How do you spot a class of related defects, and what does investigating the class once give you?
answer
- Group by cause, never by symptom
- What single change prevents all of them?
- Read a few end to end first
- One explanation must survive every member
- The outlier is a second class
basics
~20 sGroup defects by cause shape (same area, same missing check, same escape route), never by shared symptom. One pass forces an explanation that must account for every member, a far harder test than explaining one.
solid answer
~50 sGroup by **cause shape, not symptom**: the same area of the system, the same absent check, the same escape route, the same data condition. Symptom grouping produces sets that share nothing but a message. Find candidates by scanning recently closed defect records for repeated areas and repeated conditions, then confirm by reading three or four of them end to end before declaring a class. What the single pass adds is a harder test: **the explanation must account for every member**. A story that convincingly explains one defect usually collapses when it has to explain five, and the surviving explanation is the durable one. The pass also yields a countermeasure sized to the weakness rather than one patch per defect, evidence of how often that weakness fires, and one owner instead of five people each spending twenty minutes. The risk is forcing unrelated defects together to justify the effort.
code
pseudocode · 13 linesbuckets = {}
for record in closedDefects(window = 180 days):
shape = (record.area, record.activatingCondition, record.stageThatMissedIt)
buckets[shape].append(record)
for shape, members in buckets:
if size(members) >= 3 and spans(members, releases = ">1"):
candidates = readInFull(members, limit = 6)
if onlyOneChangeWouldHavePreventedAll(candidates):
openOneInvestigation(shape, evidence = candidates, owner = onePerson)
else:
# a theme, not a class - split it or leave the defects alone
park(shape)go deeper
Know that repeated defects can be explained together rather than one at a time, and that what makes them one group is a shared cause rather than a shared error message.
Be ready to describe how candidates are found - repeats of area, condition and missed stage across recently closed defect records - and to apply the test that one change would have prevented all of them.
Show why the group pass is a stronger test than five separate ones: the explanation has to survive every member, and the members themselves become the frequency evidence that funds the change.
Own the boundary and the discipline. Decide how large a set is worth reading, what the remainder that is not explained gets recorded as, and how you stop teams retro-fitting members onto a conclusion they already like.
## What "a class of related defects" actually means A class of related defects is a set of separately reported, separately fixed defects that share a **cause shape**, not a symptom. Cause shape is the combination of where the weakness lives, what condition activates it, and which check failed to stop it. Three defects are one class when a single change would plausibly have prevented all three - even if they were reported by different people, in different parts of the product, with completely different user-visible symptoms. That last point is what makes the skill non-obvious. The natural grouping is by symptom, because symptoms are what the defect records describe. Symptom grouping is almost always wrong: the same error text can come from four unrelated weaknesses, and one weakness can surface as a wrong figure, a slow page and a rejected submission depending on who met it. ## Identifying the class - **Scan a window of closed defect records** - a quarter is usually enough - for repeats: the same area touched, the same condition named, the same stage that should have caught it. - **Read three or four candidates end to end** before declaring a class. Titles and cause fields are too compressed to judge from. - **Apply the one-change test.** Ask what single change would have prevented every candidate. If you cannot name one, you have a theme, not a class. - **Watch two false positives.** The busiest area of the product collects unrelated defects simply because most changes land there, and a shared symptom string collects unrelated causes. - **Keep the outlier honest.** A candidate the explanation does not cover is not a rounding error - it is either a second class or evidence that the first class is wrong. ## What one pass produces that per-defect work cannot | | One investigation per defect | One pass over the class | | --- | --- | --- | | The explanation must fit | one defect | every member of the set | | Typical output | a note per record, rarely reread | one countermeasure sized to the weakness | | Evidence of frequency | absent by construction | built in - the members are the evidence | | Cost shape | many small efforts, each unowned | one effort with one named owner | | Characteristic failure | a plausible story nobody can refute | a class forced together to justify effort | Four things follow from the table and they are the substance of the answer: 1. **A falsifiable explanation.** With one defect in front of you, almost any causal story survives, because there is nothing to contradict it. With five, most stories die on the second or third member. What survives has been tested rather than merely believed. 2. **Frequency evidence that funds the fix.** The hardest part of acting on a weakness is usually persuading someone to pay for the change. "This happened five times in ninety days and here are the five" is an argument; "this could happen again" is not. 3. **A countermeasure at the right size.** Per-defect work naturally produces per-defect remedies, each cheap and each covering one member. Seeing the set makes it obvious whether the right change is at the boundary all five crossed, rather than five separate guards. 4. **A boundary.** Deciding which candidates the explanation does not cover is real output. It splits a vague sense that "we keep having problems here" into one class you can act on and a remainder you know is still unexplained. ## Where it goes wrong - **The symptom class.** Everything with the same message is swept together, the explanation is generic, and the change fixes none of them. - **The class that becomes a project.** Twenty members are collected, the pass expands into a survey, and it delivers after everyone has stopped caring. Keep the set small enough to read - five or six well-chosen members carry nearly all the signal of twenty. - **Retro-fitting.** The team decides the cause first and then collects defects that agree with it. The one-change test is the defence: it must be applied to candidates chosen before the explanation existed. - **Losing the individual fixes.** Grouping is for explanation, not for repair. Each defect still gets fixed and closed on its own schedule; the class pass is what happens afterwards, over records that are already closed. - **No owner.** A class investigation with the whole team nominally responsible produces a shared document and no change. One named person, a small set, and a change with a due date.
- Two of the six members do not fit the explanation you have reached. What do you do?Do not stretch the explanation to cover them, which is how a class becomes unfalsifiable. Say plainly that the explanation covers four and leave the other two as an open remainder. If the two share their own shape, they are a second class; if not, they were never members.
- How is a class investigation kept from expanding into a survey that never lands?Bound it before it starts: a fixed window of records, a set small enough to read fully, one named owner, and a required output of one change with a due date. Five or six well-chosen members carry nearly the whole signal of twenty.
saying these in an interview costs you the question
- Group defects by the error message they show
- The more members in the class, the better the analysis
- Every defect in a busy area belongs to one class
- Stretch the explanation until it covers every member
- Wait for the class before fixing the individual defects