skip to content

A product feature's generated answers are wrong in 4% of attempts and will not reach zero. How do you make the ship call?

level: principalimportance: should knowfreq 42%

answer

  1. Compared against what, exactly
  2. The current process has a rate too
  3. Containment and detectability, not the number
  4. Bounded and visible, or unbounded and silent
  5. Who accepted the residual, by name

basics

~20 s

Compare against the process the feature replaces, not against zero. Establish whether the figure is trustworthy, what each wrong answer costs, whether errors are detectable by the user and by the team, and what reverting takes. Then record the number, its conditions and who accepted it.

solid answer

~50 s

Zero is not on offer, so the comparison is between 4% and whatever the current path does — which is rarely error-free and often slower. Four things settle the call. **Is the figure trustworthy**: judged over attempts resembling real input, with a stated rule, and enough of them that the range is narrow. **What the 4% costs**: split by consequence, with the irreversible share driven to nothing by gating rather than by hope. **Is being wrong detectable**, both by the user in the moment and by the team afterwards through a signal recorded from day one. **What reverting takes**, concretely. Then decide, and write down the number, the conditions it holds under, the classes and their tolerances, the reversal condition and its owner, and the name of the person who accepted the residual — the product decision-maker, never the engineer who measured it.

code

yaml · 13 lines
yaml
decision:
  feature: generated reply drafts in the support workspace
  measured_residual: 4.0% of attempts wrong (240 judged, repeats 3.1% to 5.2%)
  judging_rule: recorded; two reviewers, disagreements resolved by a third
  classes:
    visible_to_user: 3.6% accepted
    escaped_to_customer: 0.4% accepted and watched
    irreversible: 0%, enforced by requiring send confirmation
  comparator: manual drafting, sampled at 6.5% wrong and four times slower
  reversal: edit rate 1.4x baseline holds expansion; 2.0x reverts
  watcher: named owner, read daily, for six weeks
  residual_accepted_by: the product decision-maker, by name
  revisit: at full exposure plus one month

go deeper

for a junior

Be ready to say that a feature built on generated answers is never error-free, and that the decision compares it with the process it replaces rather than with a perfect result nobody was offered.

for a middle

Explain what has to be true of the measurement before any percentage is worth arguing about: representative inputs, a stated judging rule, and enough judged attempts for the range to be narrow.

for a senior

Show how containment changes the call — gating the irreversible class, making errors visible to the user in the moment, and recording a signal that reveals degradation after release.

for a principal

Own the asymmetry. Argue when the cost of holding dominates the cost of shipping and when it reverses, and insist the residual is accepted by the person who owns the product outcome and revisited on a stated date.

## The comparison is never against zero The question usually arrives as "can we accept 4%", which quietly compares the feature with a perfect version that was never on offer. The real options are the feature at 4%, the process it replaces, and no feature at all. The process it replaces has an error rate of its own and it is almost never measured: people mistype, skip steps, work from stale copies and get tired. If the existing path is wrong in 7% of attempts and takes four minutes, then a feature wrong in 4% of attempts and taking ten seconds is an improvement, and the fact that its errors are new and attributable to a system is a matter of perception rather than of expected loss. Perception is nonetheless a real cost. Errors a system makes are noticed, aggregated and remembered in a way individual human errors are not, and people forgive their own mistakes more readily than a product's. The comparison is honest only when it also accounts for who gets blamed and what the errors spend from trust in the rest of the product. ## The two defensible positions **Ship it, contained.** The residual is measured, the expensive classes are gated so the remaining errors are cheap and detectable, users are told plainly what the feature is and is not, and a rate-based reversal condition exists with an owner and a fallback. The argument: expected value is positive against the real alternative, the downside is bounded by containment, and holding indefinitely for a number that will never reach zero is a decision never to ship at all. **Do not ship this shape.** Narrow it to the input range where it is reliably right and route the rest to the previous path; or demote it from acting to suggesting; or hold until the containment work exists. The argument: the residual is invisible to the person using it, the team has no way to discover afterwards that it was wrong, or the expensive class cannot be gated without removing the point of the feature. Both positions accept the same 4%. What separates them is containment and detectability, not the number. ## What picks between them 1. **Is the figure real?** Judged on inputs resembling what people will actually send, with the judging rule written down, over enough attempts that the range is narrow enough to decide on. A 4% measured over a hundred convenient inputs is not a 4%. 2. **What does the wrong 4% cost?** Split by consequence. If the irreversible slice is not zero after gating, that alone settles it. 3. **Can the user tell?** A feature whose errors are visible in the moment is bounded by the user's attention. One whose errors are plausible, confident and invisible transfers the whole risk to whoever reads the output later. 4. **Can we tell, afterwards?** If no recorded signal moves when the feature degrades, you are not shipping a 4% feature; you are shipping an unmeasured feature that was 4% once. 5. **What is the exposure plan?** Starting narrow and expanding on a reading of that signal turns one large bet into several small ones — but only if a breach genuinely holds the next expansion. 6. **Who carries the residual?** A cost borne by the person who chose to use the feature differs from one borne by a customer downstream who did not. ## The cost of being wrong in each direction | Direction | What it costs | |---|---| | Shipped and should not have | errors reach people; trust is spent on the whole product rather than this feature; withdrawal after adoption is far more expensive than never launching; sometimes an obligation to correct something outside the team | | Held and should have shipped | the worse existing process continues at its own unmeasured rate; the work stales and is quietly defunded; the team learns nothing, because real input only exists in production | The asymmetry between those two rows is what the decision actually turns on, and it is specific to the product. Where the downside is bounded, reversible and visible, the cost of holding usually dominates — the only route to a better rate runs through exposure. Where the downside is unbounded or silent, the cost of shipping dominates, because you will not find out in time to act. ## What gets written down - The measured rate, the number of judged attempts, and the spread across repeats. - The rule that judged right from wrong, and the input mix it was measured on. - The consequence classes and the tolerance accepted for each — particularly the irreversible one at zero, and how that zero is enforced. - The reversal condition: its signal, its levels, its window, and the person who reads it. - The date or event at which the decision is revisited, so an acceptance does not become permanent through silence. - **Who accepted the residual, by name.** That is the person who owns the product outcome, not the engineer who measured the rate. Stating a number is reporting; accepting it is deciding, and blurring the two is what turns an acceptance into a private judgement nobody can find again.

  • You are asked to sign the acceptance of the residual yourself. What do you say?
    That reporting and accepting are different acts. I can state the rate, the conditions it holds under, what each class of error costs and what would reverse the decision. Accepting the residual is a product outcome, so it belongs to whoever owns that outcome. Recording their name is not bureaucracy — it is what makes the acceptance findable and revisitable when the rate moves.
  • The process the feature replaces has never been measured. Does the comparison still hold?
    Only as an argument to measure it, cheaply and roughly. Sample recent outcomes of the existing path and judge them by the same rule you applied to the feature. An unmeasured comparator lets both sides assert whatever suits them, and the usual result is an unspoken assumption that the current process is error-free, which it never is.
  • What makes a staged expansion more than a formality?
    A breach at the current share of traffic must hold the next increase, and somebody must be empowered to hold it. If expansion runs on a calendar while the signal is merely noted, the staging bought nothing: the same total exposure arrives a little later, and the early group absorbed risk without purchasing any information.

Refusing to ship until the rate is zero is like refusing to hire because no candidate is perfect: the empty seat has a cost too, and it is being paid the whole time.

saying these in an interview costs you the question

  • Treats zero errors as the alternative actually on offer
  • Holds indefinitely for a rate that cannot reach zero
  • Ships with no recorded signal that would show degradation
  • Lets the engineer who measured the rate also accept it
  • Assumes withdrawal after adoption is as cheap as never launching