Not every wrong answer from a generative product feature costs the same. How do you separate cheap misses from expensive ones?
answer
- How often is not how much
- Some misses cost only a shrug
- Detectable, reversible, or already acted on
- One tolerance per consequence class
- Over-sample the rare expensive branch
basics
~20 sClassify each wrong answer by what it costs rather than how often it happens: whether the user notices at the time, whether it can be undone, and whether the feature acted or only proposed. Then set a separate tolerance for each class.
solid answer
~50 sAn aggregate error rate treats a suggestion somebody deletes and a wrong figure that reaches a customer as the same event. Before recommending release, sort the observed wrong answers along three axes: **detectability**, whether the user notices at the moment it happens; **reversibility**, what undoing it costs once noticed; and **agency**, whether the feature acted or only proposed something a person accepted. Those give a handful of consequence classes, each with its own tolerance — a few percent may be fine for discarded proposals and unacceptable for anything irreversible. The recommendation then reads as a rate per class rather than one number, and it changes the work: the cheapest move is usually not lowering the rate but demoting the expensive class, turning an action into a proposal, showing the material an answer rests on, or requiring confirmation on the irreversible branch.
code
pseudocode · 15 linesclassify(miss):
if miss.caused_an_irreversible_operation: return "acted"
if miss.left_the_organisation_unchecked: return "escaped"
if miss.was_visible_to_the_user_at_the_time: return "discarded"
return "corrected_late"
tolerance = {
"discarded": 0.05,
"corrected_late": 0.02,
"escaped": 0.002,
"acted": 0.0
}
for cls in tolerance:
report(cls, observed_rate(cls), tolerance[cls], attempts_judged(cls))go deeper
Be ready to give two wrong answers from the same feature that cost very different amounts, and explain why counting them as one unit each hides the difference that matters.
Explain the axes that set the cost — whether the user notices, whether it can be undone, and whether the feature acted or only suggested — and how each produces a different tolerance.
Show that you would report a rate per consequence class on its own sample, and that your first move against the expensive class is to change its shape rather than chase the overall rate down.
Own the tolerance table itself: which classes exist for this product, who is allowed to accept each one, and why a single figure covering all of them would be dishonest to whoever signs it.
## Why one number is the wrong unit A residual error rate answers "how often", and the release decision needs "how much". The two separate as soon as the wrong answers stop being interchangeable. A drafting suggestion a user glances at and deletes costs a second of attention. The same feature returning a plausible but wrong figure that somebody pastes into a document sent to a customer costs an apology, a correction, and part of the trust that made the feature worth building. Multiplying one rate by one cost gives an expected loss that is wrong in both directions: it overstates the cheap misses, which dominate the count, and understates the expensive ones, which dominate the consequences. So the unit that belongs in a release recommendation is a small set of **consequence classes**, each with its own observed rate and its own tolerance. Classify by what a wrong answer costs, not by what produced it. ## Three axes that decide the cost - **Detectability.** Does the person find out at the moment it happens? An answer that is obviously wrong is nearly free, because they simply ask again. An answer wrong in a way only a specialist would catch is the expensive one, precisely because it gets used. - **Reversibility.** Once noticed, what does undoing it take? A field that can still be edited before submission is cheap. A message already delivered, a payment already made, or a record already exported into another system is not. - **Agency.** Did the feature *do* something, or *propose* something a person accepted? A proposal keeps a human between the error and its consequences. Acting removes them, and moves the whole error rate into the expensive class. Two further questions sharpen the picture: **who bears the cost** — the person using the feature, a colleague, or a third party who never chose to use it — and **whether the error looks uncertain**, since an answer that is confidently formatted and wrong is worse than one that is visibly hedged and wrong. ## A worked set of classes | Class | Shape | What it costs | Sensible tolerance | |---|---|---|---| | Discarded | the user sees it is wrong and asks again | seconds of attention | a few percent is normal | | Corrected late | wrong but plausible; caught before it leaves the team | rework and annoyance | low, with a check inside the flow | | Escaped | wrong, unnoticed, carried into a document or a decision | apology, correction, trust | small, and only with a way to find out | | Acted | the feature performed an irreversible operation | the operation itself | effectively zero; gate it | The numbers are not universal; they belong to the product. What is universal is that a single tolerance across all four rows is either far too strict for the top row or far too loose for the bottom one. ## The cheapest lever is the shape, not the rate Once misses are classified, the recommendation stops being only "is the rate low enough" and becomes "can the expensive class be moved somewhere cheaper". That is ordinary product and engineering work, and it is usually far cheaper than driving the error rate down: 1. **Demote actions to proposals.** Anything irreversible waits for a person to confirm, and the confirmation states in concrete terms what is about to happen. 2. **Make errors detectable.** Show the material the answer rests on so the claim can be checked without leaving the flow, and present uncertain answers differently from confident ones. 3. **Contain the escape routes.** Put a check between the feature's output and anything that leaves the organisation. 4. **Narrow the surface.** Restrict the feature to the inputs where it is reliably right and route the rest to the previous path, rather than accepting a worse rate everywhere. A recommendation that says "the irreversible class is now zero because that branch requires confirmation, and every remaining error is visible to the user at the moment it occurs" is far stronger than one reporting that the rate improved by two points. ## Sampling for the class that matters The expensive class is usually the rare one, which means a sample drawn to mirror ordinary traffic contains almost none of it and the aggregate is dominated by the cheap class. Judge the classes on separate samples: deliberately over-sample the inputs that lead into the expensive branch, report that class on its own denominator, and never let a comfortable aggregate stand in for it. If you cannot gather enough attempts in the expensive class to say anything, that is itself a finding — report it as unmeasured rather than implying it was measured and small.
- The expensive class is rare, and your sample contains three of them. What do you report?Report that class on its own denominator and state that the denominator is three. Three attempts support no rate at all, so present it as unmeasured rather than as a reassuring percentage, and either judge more attempts on inputs deliberately chosen to reach that branch or recommend gating the branch until you can.
- Which is usually cheaper: lowering the error rate, or changing what an error costs?Changing the cost, almost always. Requiring confirmation before an irreversible operation, showing the material an answer rests on, or routing uncertain inputs to the previous path are ordinary engineering changes with predictable outcomes. Driving the rate down is open-ended work with uncertain returns, and it never reaches zero.
Counting misses without their cost is like counting dropped items without noticing that most were paper and one was a glass.
saying these in an interview costs you the question
- Applies one tolerance to every kind of wrong answer
- Counts wrong answers without asking what each costs
- Assumes users always notice an incorrect answer
- Lets the feature take irreversible actions from its own output
- Measures the expensive class on a sample containing almost none