Leadership wants a per-team defect density target. How do you set defect measures without distorting behaviour?
answer
- A measure becomes a target
- Every input is a human decision
- Give each number a partner
- Direction beats a threshold line
- Say what would retire it
basics
~10 sRefuse the single per-team target and offer paired measures instead: every efficiency number gets a counter-measure, definitions are published and frozen, trends beat thresholds, and nothing is wired to ranking or pay.
solid answer
~50 sDefect measures are built entirely from human decisions — whether to file, what to classify it as, when to close — so a target moves those decisions before it moves the software. Predict the distortions out loud: under-filing, defects reclassified as questions or change requests, cosmetic closures, and duplication kept in the codebase because it inflates the denominator. Then design against them. Measure where the outcome is owned rather than per individual; pair every measure with its counter-measure, so density travels with escaped-defect rate and closure rate travels with reopen rate and backlog age; publish and freeze the definitions and window so a series means one thing; prefer direction over a threshold line; never attach the numbers to compensation or a league table; audit a sample of closed reports; and state up front what would retire each measure. Offer field-observed outcomes as the outer measure, since those sit outside the team's own bookkeeping.
go deeper
Understand the basic trap: a defect count is made of filing decisions, so rewarding a low count teaches people to file less. Recognising that is enough at this level.
Name concrete distortions you could expect — reclassification as a question or change request, cosmetic closures, keeping duplication so the per-thousand-lines ratio looks better — and be able to explain why each one follows from the target.
Bring the counter-measure pairs and be able to justify each pairing mechanically: which manoeuvre improves the first number and which series it damages. Be ready to spot a gamed series from its neighbours rather than from an accusation.
Own the whole measurement design: what is measured, at which level of ownership, under whose stable definition, with what review cadence and what retirement condition. Expect to negotiate with leadership rather than refuse, and to move the outer measures toward field-observed outcomes.
## Why defect measures distort faster than most Goodhart's law is the compact statement: a measure that becomes a target stops being a good measure. Campbell's law makes the social-science version, that the more a quantitative indicator is used for decision-making, the more it will be corrupted and the more it will distort the process it was meant to monitor. Defect measures are unusually easy to corrupt because of what they are made of. Every input is a human decision that costs nothing to adjust: whether to file at all, what to call it, whether to merge two reports or split one, which severity to propose, when to close, and what counts as closed. None of that requires anyone to lie. The intake path just moves. ## The distortions you should predict out loud - **Under-filing.** Defects reported verbally, over a desk, or in a chat thread that never reaches the tracker. - **Reclassification.** A defect becomes a question, a change request, a support ticket, expected behaviour, or a technical-debt item. - **Merging and splitting.** Separate faults collapsed into one report under a count ceiling; one fault split into six under a closure-rate target. - **Denominator inflation.** If density is measured per thousand lines, the cheapest way to improve it is to keep the duplication. - **Cosmetic closure.** A closure-rate target produces cannot-reproduce and defer-to-next-release resolutions, and the tell is a rising reopen rate. - **Mass deferral.** A backlog-size target produces a very long known-issue list. - **Collateral damage.** The worst outcome is social: once filing a defect costs a colleague something, testers and developers stop cooperating, and you lose the reporting culture that made any of the numbers meaningful. ## Design rules for a measure that survives being watched 1. **Measure where the outcome is owned.** A product or value stream, rarely a single team, never an individual. 2. **Pair every measure with its counter-measure.** Density with escaped-defect rate. Closure rate with reopen rate and backlog age. Defects found in testing with defects reported per active user after release. A pair is much harder to move in both directions at once than either half alone. 3. **Prefer trends to thresholds.** Direction across several releases under one stable definition beats a number with a line drawn through it. 4. **Freeze and publish the definitions.** What counts as a defect, what the denominator is, what the observation window is. Any change to a definition resets the series, and that reset must be visible. 5. **Never wire them to compensation, ranking or a public league table.** That single decision converts a diagnostic into a target faster than anything else on this list. 6. **Use the number to ask, not to judge.** The output of a metric is a question for the team that owns it, and the team's explanation is usually more informative than the number. 7. **Audit the data.** Sample 30 closed reports a quarter and check the classifications honestly. A measure nobody audits is a measure nobody can trust when it improves. 8. **Sunset it.** A measure introduced to answer a question should be retired when the question is answered, and you should say up front what would retire it. ## What to concede Leadership is asking for something legitimate: a way to know whether quality is moving without reading every report. Refusing to measure is not the principal answer. Offer a small paired set with an explicit review cadence, an explicit retirement condition, and a standing agreement that any number which starts looking unusually good gets audited before it gets celebrated. Push the outer measures toward field-observed outcomes — escaped defects per release, customer-reported severity mix, time to restore service — because those sit outside the team's own bookkeeping and are correspondingly harder to move by reclassification. ## The incident to have in your pocket A subscription renewal team was handed a ceiling of 2.0 defects per thousand lines per release. The next release reported 1.4 and was held up as a success. Six weeks later a currency-rounding drift arrived from a customer, and an audit of the tracker found 18 reports filed as "billing questions" in a support queue that never reached the defect database. Nobody falsified anything. The number was true, the target was met, and the measure had stopped measuring the thing it was named after.
- Which counter-measure would you attach to a target on defect closure rate?Reopen rate together with the age profile of the open backlog. A closure-rate target is most cheaply met by closing as cannot-reproduce, deferring to a known-issue list, or clearing the easy recent items — and each of those shows up in one of those two series. Reopen rate catches the cosmetic closures, backlog age catches the cherry-picking. Neither can be improved by the same manoeuvre that inflates closure rate, which is what makes the pair informative.
- A team's density rose sharply after they started writing better defect reports. How do you handle that in a review?Treat it as a data-quality change, not a quality regression, and say so before anyone else interprets it. A change in filing behaviour resets the series exactly like a change in definition, so I would annotate the break, compare only within each regime, and pair the density with the escape rate — which should be falling if the reporting really improved. Punishing that rise is the fastest way to teach the organisation to stop filing.
- Leadership insists on one number for a dashboard. What do you give them?An outcome measure from outside the team's own bookkeeping — escaped defects per release over a fixed observation window, or customer-reported severity mix — rather than an internal process ratio. It is harder to move by reclassification, it is closer to what they actually want to know, and it does not distort the reporting culture that every other measure depends on. I would attach the observation window and the retirement condition to it, and keep the diagnostic pairs for the team's own review.
- How would you know a defect measure has already started to be gamed?Watch for improvement that no process change explains, and check the neighbouring series: filing volume dropping while support contacts hold steady, a rising share of reports resolved as not-a-defect or as change requests, a closure rate improving while reopen rate climbs, or a backlog that stops growing because deferrals spiked. Then audit a sample of recent closures directly. An unexplained improvement deserves the same scrutiny as an unexplained regression.
A thermometer taped to a radiator still reads a real temperature. The reading is honest and the room is still cold.
saying these in an interview costs you the question
- Accepts a per-individual defect target without objection
- Proposes a single number with no counter-measure
- Changes the defect definition mid-series without annotating it
- Ties defect measures to performance reviews or a league table
- Celebrates a sudden improvement without auditing the data
- Refuses to measure anything rather than designing measures well