skip to content

questions

4

Why does a defect found in production usually cost more to fix than the same defect found during development?

level: juniorimportance: must knowfreq 68%

answer

  1. The fix itself is the cheap part
  2. Count what the fix drags along
  3. Detection, reproduction, lost context, coordination
  4. Rework fans out to data and specs
  5. Delay measured in dependent work, not dates

basics

~20 s

A production defect costs more because the code change is only a fraction of the work. Someone must notice it, reproduce it and re-learn code written weeks ago, and the team then pays for data repair, support and an unplanned release.

solid answer

~50 s

The code change is usually the smallest part of the bill. A defect that survives into production has to be noticed by someone who was not looking for it, reproduced from an incomplete report, and diagnosed by an author who has moved on and lost the context. The rework then fans out: the specification, the tests, the documentation and often the data written by the faulty code all need correcting. Around that sit coordination costs — triage, a decision to release out of band, a possible rollback — and the consequences borne outside the team, from support handling to compensation. Caught while the author is still writing the code, the same defect costs minutes: context, artefacts and data are all still local and disposable. The driver is not the phase label; it is how much other work has been built on top of the defect and how many people are now involved.

go deeper

for a junior

Be ready to list what a production defect costs beyond the code change: noticing it, reproducing it, re-learning the code, fixing the data, and the support work around it. One concrete example from your own work beats any general statement.

for a middle

Explain the mechanism rather than the slogan. Walk through how rework fans out into specifications, cases, documentation and stored data, and why context loss makes diagnosis the dominant cost long before the fix is written.

for a senior

Show that you have paid this bill. Describe a real late defect, what the repair actually consisted of, how much of it was data or coordination rather than code, and what you changed afterwards so that class of defect is caught in-cycle.

for a principal

Own the caveat. Say plainly that this is a tendency driven by dependent work and dependent data, that it is weak for cosmetic defects, and use it to argue about feedback latency rather than to justify a blanket increase in checking.

## The cost of a defect is not the cost of the fix When people say a defect "costs more" later, they are not claiming that the corrected line of code becomes harder to type. They are claiming that the **total cost of ownership of one defect** — everything the organisation spends between the moment the mistake is made and the moment the world is consistent again — grows with the delay before detection. Understanding *which* components grow, and why, is what separates a candidate who can argue this from one who can only recite a multiplier. ### The components that grow with delay **Detection.** While you are writing the code, you are actively looking for problems and you get an answer in seconds. In production nobody is looking; the defect is found by a customer, a downstream team, or a reconciliation that fails. Detection is now someone else's interrupted day. **Reproduction.** A defect caught at the keyboard is already reproduced — you are standing in it. A reported defect arrives as a partial description of a symptom, often without the inputs, the timing or the state that produced it. Recreating those conditions can cost more than the fix. **Context re-acquisition.** The author has forgotten the change. Everything they knew in the moment — why the branch exists, which case it was guarding, what they nearly did instead — has to be rebuilt by reading code, history and notes. This is the single most under-counted item on the list. **Rework fan-out.** By the time a defect ships, other artefacts have been built on top of it: a specification that describes the wrong behaviour, cases that assert it, documentation and training material that teach it, downstream code written against it, and — critically — **data already written by the faulty code**. Correcting the behaviour without correcting the data leaves you with a half-fixed system. **Coordination.** A defect found before merge involves one person. A defect found in production involves triage, a severity decision, a decision about whether to release out of band, communication to whoever is affected, and often a review afterwards. Every one of those is several people's time. **External consequence.** Finally there are the costs that land outside engineering entirely: support handling, manual workarounds, credits or compensation, and lost trust. These are simultaneously the largest and the least measurable component. ### A worked example A warehouse stock ledger reads on-hand quantities through a cache. A change makes one read path serve a stale value under a specific interleaving, so a small number of ledger lines are written with quantities that were already superseded. Nothing fails loudly. The 6-hour nightly reconciliation run consumes those lines and propagates the wrong figures into the next day's opening balances. Caught in review or by a case at development time, this is a few minutes: fix the read path, add a case that pins the interleaving, done. Caught nineteen days later by a warehouse team counting shelves, the bill looks different: 41 engineer-hours to reproduce, diagnose and fix; a re-run of nineteen nights of reconciliation to rebuild balances; 2,847 ledger lines to correct by hand where the re-run could not decide; and 63 support-hours answering the warehouses whose counts had not matched for weeks. The corrected code is still a handful of lines. Everything else is the delay. Notice what actually drove the number: not the calendar date, but the fact that a **6-hour batch job ran nineteen times on top of the defect**. That is the general rule — cost tracks how much work has accumulated on the mistake, and the phase label is only a proxy for that. ### Where the pattern breaks This is a tendency, not a law, and a good candidate says so unprompted. A misspelled label on a screen found in production costs almost exactly what it would have cost during development: no data is wrong, no downstream artefact assumed anything, nobody's balance is off. Conversely, a misunderstanding of a requirement caught *before* any code exists can still be expensive if a quarter of planning was built on it. What genuinely scales with delay is the amount of dependent work and dependent data, and the number of people who must now be involved. ### The vocabulary to reach for The cost-of-quality vocabulary splits the same spending into work done to prevent defects, work done to detect them, rework on defects found before release (**internal failure**) and everything spent after release (**external failure**). The whole late-defect argument is, in that language, simply the observation that external-failure cost per defect is the largest of the four and the one you have the least control over once it is incurred.

  • Which single component of that cost do teams most often forget to count?
    Correcting the data the faulty code already wrote. Teams estimate the code change and the re-release, then discover that months of records are wrong and that no automated path can decide the ambiguous ones. In a ledger-style system this repair is frequently larger than the engineering fix, involves people outside the team, and cannot be rolled back.
  • Give an example where a defect found in production costs no more than one found at development time.
    A cosmetic defect with no downstream dependants — a wrong label, a misaligned column, a typo in a message. Nothing was computed from it, no stored data is wrong, and nobody built work on top of it, so the repair is the same edit either way. These cases are the reason the cost argument should be made about classes of defect rather than about defects in general.
  • How does a long-running batch job change the shape of this cost?
    It converts one wrong write into many. Each run propagates the bad values further and creates more derived records that must later be rebuilt, so cost grows with the number of runs rather than smoothly with time. It also delays detection, because the visible symptom appears one cycle after the cause. Systems with long batch cycles therefore need faster in-cycle checks, not just more of them.

Correcting a recipe while you are still writing it costs an eraser. Correcting it after ten thousand copies are printed and the meals are cooked costs a reprint, an apology and the wasted ingredients.

saying these in an interview costs you the question

  • Claims the cost is exactly ten times higher per phase
  • Counts only the engineering time to change the code
  • Forgets that shipped code has already written wrong data
  • Treats the pattern as a law that holds for every defect
  • Cannot name a single cost component beyond the fix
  • Says late defects are expensive only because of reputation

context

open as a page

What are the four cost-of-quality categories, and how does spending move between them?

level: middleimportance: should knowfreq 46%

basics

~20 s

Cost of quality splits spending into prevention (stopping defects being made), appraisal (looking for them), internal failure (rework before release) and external failure (everything after release). The argument is that prevention and appraisal spending buys down the two failure categories.

open as a page

Why is the 100:1 late-defect cost curve contested, and how should you cite it responsibly?

level: seniorimportance: should knowfreq 33%

basics

~20 s

The order-of-magnitude curve rests on small, old datasets from very different projects, and the round multiplier is usually quoted third-hand without its context. Cite the direction and the mechanism, back the size with your own measured repair cost, and give a range.

open as a page

How would you build the economic case for moving quality work earlier without leaning on a defect-cost multiplier?

level: principalimportance: should knowfreq 38%

basics

~20 s

Measure your own recent escaped defects, price the repair in hours and consequences, and compare that against the cost of the specific earlier check that would have caught them. Argue at the margin, state the range, and name what would prove you wrong.

open as a page