skip to content

A defect model reads 94% accuracy on a test set rebalanced to 50/50, but live defects run at 0.3%. What does that number tell you?

level: seniorimportance: should knowfreq 44%

answer

  1. the test prior is not the live prior
  2. accuracy is a prevalence-weighted average
  3. on a 50/50 set it is balanced accuracy
  4. class-conditional rates travel, accuracy does not
  5. recombine at 0.003 and compare to 99.7

basics

~20 s

Almost nothing about production. Accuracy depends on the class mix it was measured on, so a 50/50 test set reports balanced accuracy, not the live figure at a 0.3% defect rate. Recombine the per-class rates at the true base rate.

solid answer

~50 s

Accuracy is prevalence-dependent, so a number measured on an evaluation set whose class mix differs from production does not transfer. On a set rebalanced to 50/50, accuracy is the simple average of the two class-conditional rates, which is exactly balanced accuracy - here 0.94 might come from 92% recall on defects and a 4% false-alarm rate on good units. What survives the change of prior are those class-conditional rates, because each is computed within one true class. Recombine them at the live mix: `accuracy = 0.003 * 0.92 + 0.997 * 0.96 = 0.96`, which is **below** the 99.7% always-good baseline. And the operational picture is worse than the metric: 4% of 99700 good units is nearly 4000 false alarms against roughly 276 real catches. Fix it by evaluating on a holdout that carries the production prior and reporting the class-conditional rates.

go deeper

for a junior

Recall that accuracy depends on the class mix of the data it was measured on, so a number from a rebalanced test set cannot be quoted as the production figure.

for a middle

Explain the decomposition accuracy = p*recall_positive + (1-p)*recall_negative, and show that at p = 0.5 it collapses to balanced accuracy. Recompute the live figure from the two rates and the true prevalence.

for a senior

Diagnose it end to end: identify the prior mismatch, recompute against the live base rate, check whether the discarded rows were dropped at random, and put the false-alarm volume and the trivial baseline in front of the decision-maker.

for a principal

Own the evaluation contract: prevalence-preserving holdouts, class-conditional rates as the reported artifact, the evaluation-set class mix published with every number, and prior drift monitored so a shifting base rate is never mistaken for model improvement.

## What went wrong Someone built an evaluation split with equal numbers of defective and good units - typically by keeping every defect and discarding most of the good units - and reported accuracy on it. The reported 94% is a real measurement of *something*, but it is not the accuracy anyone will observe on the line, because **accuracy is a function of the class mix of the set it is computed on**. Write it out. Let *p* be the share of positives (defects), *r+* the recall on defects and *r-* the recall on good units (specificity): ``` accuracy = p * r+ + (1 - p) * r- ``` The class-conditional rates *r+* and *r-* are properties of the model. The weights are properties of the evaluation set. Change the set's composition and the same model reports a different accuracy without a single prediction changing. ## The 50/50 identity worth knowing On a set rebalanced to exactly 50/50, p = 0.5, so ``` accuracy = 0.5 * r+ + 0.5 * r- = (r+ + r-) / 2 ``` which is precisely **balanced accuracy**. So the reported 94% is not a corrupted number - it is the model's balanced accuracy wearing the wrong label. That reframing is the crisp answer to give: 'that is balanced accuracy, and it is fine as a class-balanced summary, but calling it accuracy invites everyone to compare it against a production baseline it was never measured against.' ## Recomputing for the live prior Suppose the underlying rates are r+ = 0.92 and r- = 0.96 (a 4% false-alarm rate on good units); their mean is the reported 0.94. At the live prevalence of 0.003: ``` accuracy_live = 0.003 * 0.92 + 0.997 * 0.96 = 0.00276 + 0.95712 = 0.9599 ``` So the honest production accuracy is about **96%** - a *lower* number than the 94% figure would lead a reader to expect, and, far more importantly, **below the 99.7% majority-class baseline**. The model that looked like a 94% success is, on the metric as reported, worse than never flagging anything. The volume picture is what actually decides deployment. Per 100000 units: 300 defects, of which 276 are caught and 24 missed; 99700 good units, of which 4% - 3988 units - are falsely flagged. The inspection queue holds 4264 units, of which 276 are real. Nothing in a 94% figure hints at that ratio, and no amount of arguing about accuracy will surface it; the flagged-set view has to be reported separately. ## Why rates transfer and accuracy does not Each class-conditional rate is computed **within** one true class - defects caught over defects present - so it is unaffected by how many rows of the *other* class are in the set. Discarding good units cannot change the defect recall. Accuracy, by contrast, is an average across classes weighted by their frequencies, so it inherits the set's composition directly. The assumption that makes the recomputation valid is worth stating in the interview, because it is the part candidates skip: the *conditional distributions must be unchanged*. Discarding good units at random preserves what a good unit looks like, so r- transfers. Discarding them non-randomly - keeping only the hard ones, or only a recent week, or only one production shift - changes what the class *is*, and then neither rate transfers and the whole recomputation is invalid. ## What to do about it 1. **Evaluate on a set that carries the production prior.** Rebalancing the *training* data is a legitimate modelling choice with its own tradeoffs; rebalancing the *evaluation* data destroys the only property that made the number interpretable. Hold out a random, prevalence-preserving slice for scoring. 2. **If the rare class is too small for stable estimates**, keep the natural prior and accept wide intervals on the rare-class rate, or oversample the majority for *speed* while carrying explicit row weights that restore the true prior in the metric. Never let the unweighted number out of the room. 3. **Report class-conditional rates as the primary artifact**, plus the raw confusion matrix and the evaluation-set class mix. Then any reader can recompute for whatever prior their site, season or segment has. 4. **Always print the majority-class baseline next to the headline.** In this case it is the single line that would have caught the error immediately: a 94% figure quoted for a domain whose trivial baseline is 99.7% should never have survived review. 5. **Watch for prior drift after launch.** Even a correctly built evaluation set goes stale if the defect rate moves - a process change that halves defects silently raises accuracy while the model gets no better, which is why prevalence-independent rates belong on the monitoring dashboard. ## The shape of the mistake The general lesson generalises past this scenario: metrics that mix across classes - accuracy, error rate, and anything else summed over rows - are only comparable between two sets that share a class mix. Metrics conditioned on the true class are portable. When a number crosses a boundary between populations, ask which kind it is before you compare it to anything.

  • Suppose you cannot rebuild the evaluation set. How do you report an honest production number?
    Report the class-conditional rates - recall on defects, false-alarm rate on good units - and recombine them at the live prevalence: accuracy = p*r+ + (1-p)*r-. State the assumption that the good units were dropped at random, so their conditional distribution is unchanged, and print the majority-class baseline next to the result.
  • Why do per-class rates survive a change of prior when accuracy does not?
    Each rate is computed inside one true class, so the number of rows in the other class cannot affect it. Accuracy is those rates averaged with weights equal to the class frequencies, so it inherits the evaluation set's composition directly. The caveat is that the within-class distributions must be unchanged.
  • The live defect rate drops from 0.3% to 0.15% after a process change. What happens to accuracy?
    It rises, with no change to the model. Halving the positive share shifts weight onto the good-unit rate, so accuracy drifts up toward that rate while the majority-class baseline rises faster still. That is why prevalence-independent, class-conditional rates belong on the monitoring dashboard rather than accuracy.

saying these in an interview costs you the question

  • Treats accuracy as a property of the model alone
  • Compares numbers measured on sets with different class mixes
  • Rebalances the evaluation set the same way as the training set
  • Recomputes for the live prior without checking the sample was random
  • Never contrasts the recomputed figure with the majority-class baseline
  • Assumes a lower prevalence means the model degraded

context