skip to content

What is defect density, and why does the denominator you choose change the answer?

level: juniorimportance: must knowfreq 62%

answer

  1. A raw count compares nothing
  2. Something has to go under the line
  3. The size unit is the argument
  4. Definition, unit, window travel together
  5. Removing duplication can worsen the ratio

basics

~20 s

Defect density is a count of confirmed defects divided by a size unit — a thousand lines of code, a functional size unit, a module, a feature. Change the unit and the same product scores differently.

solid answer

~50 s

Defect density normalises a raw defect count so that things of different size can be compared: defects found, divided by a chosen size unit, over a stated window. The denominator carries the whole argument. Lines of code reward verbose code and punish a rewrite that removes duplication; a feature or story unit ignores implementation size and is wildly uneven between units; a per-module count means nothing until you know how big the modules are. The numerator is just as contestable — confirmed defects only, or duplicates and rejected reports too, and found during which activity. A density figure is therefore only readable with three things attached: what counted as a defect, what the size unit was, and over what window and build range. Used honestly it points at which components deserve attention; used as a target it just moves the counting.

go deeper

for a junior

Be ready to state the formula and name at least two possible size units. The point that must land is that the number is meaningless without saying what was divided by what, and over which window.

for a middle

Explain the tradeoffs between size units: code volume is cheap but measures the solution, functional size compares across stacks but costs effort, feature units are legible but uneven. Show that removing duplication can make the ratio worse.

for a senior

Demonstrate that you read density next to an escape measure. An interviewer expects you to say a low density is equally consistent with clean code and weak testing, and to name the effort and definition confounders before interpreting a ranking.

for a principal

Own the definition. Decide what counts as a defect and which denominator the organisation uses, publish it, keep it stable across releases, and be ready to argue why you would not hand that ratio to leadership as a per-team target.

## The ratio Defect density is a normalised defect count: defects attributed to a piece of software, divided by a measure of that software's size, over a stated window. density = confirmed defects / size units Dividing is the whole point. A bare count of 84 defects says nothing — 84 in a three-week-old prototype and 84 in a five-year-old billing subsystem are not the same fact. Normalising lets you put two components, two releases, or two suppliers' deliverables side by side, and it is the only way to turn "is this module worse than that one?" into a question with an answer. ## Three things must travel with the number **The numerator — what counted as a defect.** Every filed report, or only confirmed ones? Are duplicates, cannot-reproduce, works-as-designed and enhancement requests stripped out? Are all severities in, or only those above trivial? Two teams with different filing cultures produce incomparable numerators from identical software. **The denominator — the size unit.** Thousand lines of code, functional size units, endpoints, screens, features, stories, or modules. **The window — where and when.** Defects found in which activity (review, unit testing, system testing, the field) and across which build range. A density from the first week of system testing and a density from twelve months in the field measure different things. Quote a density without all three and it is not readable, only repeatable. ## Why the denominator decides the answer **Lines of code** are cheap to obtain and the most common choice, but they measure the solution rather than the problem. Verbose or generated code inflates the denominator and flatters the ratio; a rewrite that removes duplication without changing behaviour shrinks the denominator and makes the density look worse. The counting rule is itself contested — comments, blank lines, test code, generated files and vendored code are all in or out depending on who counts. **Functional size** units count what the software does rather than how much text expresses it, so they compare across implementations and technology stacks — at the cost of a trained counter and real effort per release. **Feature, story or endpoint** units are cheap and legible to product people, but the units are wildly uneven: one story is an afternoon, another is a fortnight. **Per module or per file** is fine for ranking inside one codebase whose modules are broadly comparable, and close to meaningless across codebases. ## A worked example A subscription renewal job ships with 12,400 executable lines and 27 confirmed defects found in system testing — 2.18 per thousand lines. The team then extracts a duplicated proration routine, dropping the component to 7,900 lines and finding 3 further defects while doing it. The cumulative figure is now 30 defects in 7,900 lines: 3.80 per thousand. The code got better and the number got worse. Nothing about the product moved in the direction the metric moved; the denominator did. The same release's currency-rounding drift shows the numerator problem. It was filed four times by three testers before triage merged the reports. Counted raw, that single fault contributes four to the numerator; counted as confirmed-unique, one. A factor of four on a small numerator swamps any real difference you were hoping to see. ## Skew Defects are not spread evenly. It is a repeated observation in the testing literature — usually called defect clustering — that a small proportion of components carries a disproportionate share of the defects. The precise split varies by system and by study, so quoting a fixed ratio as a law overstates the evidence. The practical use is real, though: rank components by density and let the ranking direct extra review, extra testing, or a rewrite decision. Guard against the feedback loop. Attention finds defects, so whichever module you look at hardest keeps topping the table. Cross-check a density ranking against change frequency, structural complexity, and where field-reported defects actually land. ## What the number cannot tell you Density measures defects **found**, not defects **present**. A low density is equally consistent with clean code and with weak testing, so it is only readable next to a measure of what escaped — defects reported from the field over a fixed window after release. On its own it is not a quality verdict, and the moment it is set as a target the numerator starts to move for reasons that have nothing to do with the software.

  • Two components report the same density, but one is three times the size of the other. Which concerns you more?
    The larger one, in absolute terms — the same ratio over three times the size means roughly three times the defects to find, fix and confirm, and a larger blast radius per release. But I would not treat the equal ratios as equivalent quality either: the larger component almost certainly received more testing, so its denominator of found defects rests on a stronger search. I would check test effort per component before drawing any conclusion about the code.
  • Should defects reported after release be counted in the same density figure as defects found during testing?
    Not in the same figure without saying so. Found-in-testing density and escaped-defect density answer different questions: the first measures what the process caught, the second measures what reached users. Merging them hides the interesting comparison. I would report them side by side over a stated post-release window, because a low found-in-testing density is only good news when the escape rate is also low.
  • Defect counts are rarely spread evenly across modules. What do you do with that observation?
    Use it to direct effort: rank components by density and send extra review, extra testing or a refactoring decision to the top of the list. The literature calls the pattern defect clustering; the exact concentration varies by system, so I would not quote a fixed ratio as a law. I would also guard against the feedback loop — the module you scrutinise most keeps looking worst — by cross-checking the ranking against change frequency and where field reports actually land.

Like injury rates: deaths per thousand journeys and deaths per billion kilometres rank the same transport modes in completely different orders, and neither number is wrong.

saying these in an interview costs you the question

  • Quotes a density number with no size unit attached
  • Treats defect density as a direct quality score
  • Assumes a lower density always means better code
  • Counts every filed report, duplicates and rejects included
  • Compares densities across teams that define a defect differently
  • Never mentions that it measures defects found, not present

context