skip to content

Why is the 100:1 late-defect cost curve contested, and how should you cite it responsibly?

level: seniorimportance: should knowfreq 33%

answer

  1. The direction survives, the constant does not
  2. Old, small, non-comparable project data
  3. The stages it plots often no longer exist
  4. Citation chain audited and found circular
  5. Replace the multiplier with your own measured range

basics

~20 s

The order-of-magnitude curve rests on small, old datasets from very different projects, and the round multiplier is usually quoted third-hand without its context. Cite the direction and the mechanism, back the size with your own measured repair cost, and give a range.

solid answer

~50 s

The familiar picture — a defect costing roughly ten times more at each successive stage, up to a hundred times or more in production — traces back to a handful of studies of large, long-cycle projects several decades old. The criticisms are specific: the datasets were small and are not directly comparable, "cost" was not measured the same way across them, the stage boundaries do not exist in short-cycle delivery, and the figure is now usually quoted from a slide that cites another slide rather than from any data. Audits of that citation chain have found much of it circular. What survives scrutiny is the **direction and the mechanism**: dependent work, dependent data and coordination accumulate on an undetected defect, so later detection generally costs more. The magnitude is not a constant. Cite it that way, and if you need a number, measure your own recent repairs and present a range.

go deeper

for a junior

Know that a well-known cost curve exists and that the round multiplier is disputed. Repeating the mechanism — later means more dependent work and data — is safer than repeating a figure you cannot source.

for a middle

Be able to name the concrete criticisms: small and dated datasets, inconsistent definitions of cost, stage boundaries that no longer exist in short-cycle delivery, and a citation chain that mostly leads back to other citations.

for a senior

Show you can appraise evidence, not just consume it. Separate the well-supported direction from the unsupported constant, split by defect class, and demonstrate how you would substitute measured repair costs from your own recent escapes.

for a principal

Own the credibility risk. Explain why using a weak figure that happens to favour you costs you every future argument, and how you would set your organisation's quality case on local, falsifiable measurement instead.

## Where the curve came from The cost-of-late-defects curve is one of the most cited pictures in software engineering and one of the least examined. Its usual ancestry runs back to cost-modelling work on large aerospace and enterprise projects published around 1981, most often attributed to Barry Boehm's software economics work and to contemporaneous inspection studies; a widely repeated restatement compresses it to "1:10:100" across requirements, development and production. Later publications repeat it with new decimal points, and it now appears in slides, certification syllabi and vendor material as though it were a measured constant of the discipline. ## The specific criticisms **The datasets are small, old and heterogeneous.** The underlying projects were large, long-cycle and defence- or enterprise-flavoured, with formal phase gates, multi-year schedules and specification documents measured in kilograms. Whether their ratios transfer to a team releasing several times a week is an empirical question that the citation does not answer. **"Cost" is not defined consistently.** Some sources count engineer-hours to correct; some include re-inspection and re-certification; some include field service. Comparing a ratio built one way with a ratio built another way is not meaningful, but the comparison is exactly what the compressed "1:10:100" invites. **The phases often do not exist.** The curve is drawn over requirements, design, code, test and operation. In continuous delivery a change may pass through all of those in an afternoon, and the interesting variable becomes *detection latency*, not *phase crossed*. A model whose x-axis is missing cannot give you a number. **The citation chain is largely circular.** Book-length audits of software-engineering folklore — Laurent Bossavit's examination of frequently repeated claims is the best known — followed these references back and reported that many of them lead to secondary sources, to papers that do not contain the claimed figure, or to data too thin to support it. Whatever one thinks of the underlying effect, the *provenance* of the round number is weak. **Counter-evidence exists for short-cycle work.** Several later analyses argue the curve is far flatter for small changes in short-cycle development with automated checks, because the components that grow with delay — dependent work, lost context, escaped data — have had far less time to grow. That, too, is not settled, and honest phrasing says so. ## What survives The **direction** survives, because it has a mechanism rather than merely a correlation: an undetected defect accumulates dependent work, dependent data, and people who must be involved in unwinding it. That reasoning is sound and does not require any dataset. What does not survive is a **universal multiplier**, and worse, a multiplier applied uniformly to every defect regardless of class. The spread across defects is enormous: a wrong screen label costs about the same whenever you find it, while a wrong write into a ledger costs whatever it takes to rebuild every record derived from it. ## How to cite it in a real conversation A warehouse-ledger team wants to fund earlier checking. Suppose a stale-cache read escapes, is caught nineteen days later, and the repair consumes 41 engineer-hours of diagnosis and fix, nineteen re-runs of a 6-hour nightly reconciliation, 2,847 hand-corrected ledger lines and 63 support-hours. Two ways of using that exist. The weak way: "industry data shows defects cost a hundred times more in production." The first sceptic asks which industry, which data, and over what definition of cost; the argument dies, and it takes the speaker's credibility on the whole subject with it. The strong way: "Our last four escaped defects cost between 9 and 41 engineer-hours plus support and data repair. Comparable defects caught in review cost under two hours. That is our own range, from our own records, over the last two quarters, and the mechanism is that a 6-hour batch job propagates a bad write before anyone sees it." It is smaller, local, defensible, and impossible to wave away with a citation dispute. ## The interview signal Interviewers asking this are testing whether you repeat industry folklore or appraise it. The strongest answer states the direction confidently, refuses the constant explicitly, names *why* the provenance is weak, distinguishes classes of defect, and then reaches for local measurement. Saying "this specific claim is contested and here is what I would use instead" is a stronger signal than any number.

  • What would convince you the effect is real even without trusting the multiplier?
    A mechanism plus local measurement. The mechanism is that dependent work, dependent data and coordination accumulate on an undetected defect, which does not need a dataset to be persuasive. The measurement is your own repair costs for recent escaped defects against comparable ones caught in-cycle, reported as a range and split by defect class rather than averaged into a single ratio.
  • How would you answer a stakeholder who cites the hundred-times figure in your favour?
    Correct it, even though it supports you. Say the direction is right but the number is not defensible, then substitute your own measured range. Letting an unsound figure stand because it favours you means the whole argument collapses the first time someone checks it, and you lose credibility on every later quality request.
  • Does the curve apply differently to a requirements misunderstanding than to a coding slip?
    Yes, and conflating them is the main misuse. A requirements misunderstanding invalidates dependent design, code, checks and often stored data, so its cost genuinely escalates with delay. A coding slip caught by an automated check costs about the same at any point before release. Any ratio averaged across both classes describes neither.

saying these in an interview costs you the question

  • Quotes the hundred-times multiplier as measured fact
  • Cannot say where the figure originally came from
  • Applies one ratio to every class of defect
  • Says the effect is pure myth with no mechanism
  • Defends a favourable number known to be unsound
  • Presents a single average instead of a range

context