How is defect removal efficiency calculated, and what does a low value for one phase tell you?
answer
- How leaky is one filter
- Found here versus found later
- Escapes only show up over time
- State the post-release observation window
- Compare filters, not organisations
basics
~20 sDefect removal efficiency is the share of the defects present at a stage that the stage actually caught: defects found there, divided by that count plus the ones that escaped it and were found later, as a percentage.
solid answer
~50 sFor one phase it is defects found in the phase, divided by defects found in the phase plus defects that were present then but found later, times one hundred. At release level it is everything found before release over everything found before release plus everything reported from the field in a stated window. A low value for a phase means that filter let most of what it could have seen pass through to a later, more expensive stage — but it can also mean the defect class is simply invisible at that level, such as a concurrency fault under load being invisible to a component test. The figure is provisional by construction, because escapes are only discovered over time, so always state the observation window. It also needs each defect's injection point, or every early filter is understated.
code
pseudocode · 7 linesphase_efficiency(found_here, escaped) =
found_here / (found_here + escaped) * 100
system_testing = phase_efficiency(26, 9) // 74.3
integration_testing = phase_efficiency(34, 26 + 9) // 49.3
code_review = phase_efficiency(41, 58+34+26+9) // 24.4
release_overall = phase_efficiency(182, 9) // 95.3go deeper
Recall the shape of the formula: what a stage caught over what it caught plus what got past it. Knowing that the second term only becomes visible later is enough at this level.
Compute it for a phase and for a release from a table of counts, and explain why the injection point of each defect is needed before a per-phase figure is honest. Be ready to state the observation window as part of the answer.
Show that you read the escaped defects, not just the ratio. Explain when a low value means a weak filter and when it means the defect class is undetectable at that level, and refuse to benchmark the number against another organisation.
Use it to argue investment: which filter to strengthen, and at what cost. Be explicit that the cost-of-late-defect multipliers commonly quoted are contested, so the case rests on your own escape data rather than a borrowed number.
## What it measures Defect removal efficiency — also described as containment or filter effectiveness — asks how good one stage of work is at catching the defects that were already present when that stage ran. phase efficiency = found in the phase / (found in the phase + escaped the phase) x 100 "Escaped" means found later, in a subsequent phase or after release, but present at the time the phase ran. The release-level figure compares everything found before release against everything found before release plus everything reported after it, inside a fixed observation window. ## Why the injection point matters A defect introduced during integration was not available for an earlier code review to catch, so it is not an escape from that review. A correct per-phase figure therefore needs each defect's injection point, which is a classification judgement made at triage or during a post-fix analysis. Teams that skip that step and treat every later-found defect as an escape from every earlier stage systematically understate every early filter, and then draw the wrong conclusion about which filter to invest in. ## A worked example A subscription renewal job release records: requirement and design review 23, code review 41, component testing 58, integration testing 34, system testing 26, and 9 defects reported from the field in the first 90 days. Total known defects: 191, of which 182 were found before release. release efficiency = 182 / 191 = 95.3% system testing = 26 / (26 + 9) = 74.3% integration testing = 34 / (34 + 26 + 9) = 49.3% code review = 41 / (41 + 58 + 34 + 26 + 9) = 24.4% Read down the table: system testing is the strongest filter in this process; the code review lets roughly three quarters of what it could have seen through to later, more expensive stages. Whether that is worth fixing is a cost argument. The claim that a defect costs dramatically more to remove the later it is found is widely repeated and directionally uncontroversial; the specific multipliers often quoted (ten times per phase, a hundred times in the field) trace back to a small number of old studies and are genuinely contested. Argue the direction, not a number. The 9 escapes here include a currency-rounding drift that only appeared at a 1,200-request-per-minute peak — fractional cents accumulating across concurrently prorated renewals. No component-level filter was ever going to catch a defect whose trigger is load. That is the second way to read the table: a low figure sometimes means the filter is weak, and sometimes means that class of defect is invisible at that level. Deciding which one you are looking at requires reading the escaped defects, not just counting them. ## Provisional by construction The escape count is unknown at release. Any figure computed on release day is an upper bound that only ever falls as the field reports arrive. So: fix an observation window — 30, 90 or 180 days — state it alongside the figure, and expect a revision. Comparing one release's 30-day figure with another's 180-day figure is a bookkeeping error dressed as a trend. ## Where it is useful and where it is not Useful: comparing filters against each other inside one process; tracking a single filter's trend across releases under a stable definition; building the investment case for an earlier filter, because the measure names which stage is leaking rather than just asserting that quality is poor. Not useful: as an absolute benchmark between organisations, since defect definitions, phase boundaries and observation windows all differ; as an individual or team target, because the numerator is made of human filing decisions and moves the moment it is rewarded; and on tiny denominators, where a phase that found 3 and leaked 1 reports 75% with an enormous error bar around it. Finally, note what the measure cannot see at all: defects that nobody has found in either bucket. It scores the filters against each other, not against reality.
- Why is a release-level removal efficiency figure never final on release day?Because the denominator includes escapes that have not happened yet. On release day the escape count is zero or near it, so the figure starts at its maximum and only falls as field reports arrive. The fix is to fix an observation window — 30, 90 or 180 days — publish the figure with that window attached, and expect a revision. Comparing one release's 30-day figure with another's 180-day figure is a bookkeeping error, not a trend.
- A phase reports 75% efficiency from three defects found and one escaped. How much weight would you give it?Almost none. With a denominator of four, one reclassified report swings the figure by twenty-five points, so the confidence interval is wider than any difference you would act on. I would aggregate across several releases before reading a trend, or fall back to reading the escaped defects individually — with numbers that small, understanding the one that got out is worth more than the ratio.
- Is it fair to compare your removal efficiency against a published industry figure?No, not directly. The numerator depends on what the organisation counts as a defect, the phase boundaries differ, and the observation window is often unstated in published figures. Those three choices move the number more than any real difference in process quality. The measure earns its keep comparing filters against each other inside one process, and tracking one filter's trend across releases under a definition that has not changed.
saying these in an interview costs you the question
- Treats every later-found defect as an escape from every earlier phase
- Reports the figure with no observation window stated
- Reads a low phase value as laziness rather than filter capability
- Compares the number across organisations as a benchmark
- Draws a trend from a phase with three or four defects
- Claims a fixed cost multiplier per phase as established fact