skip to content

Why is 99.7% accuracy uninformative for a defect detector when 0.3% of units are defective?

level: juniorimportance: must knowfreq 88%

answer

  1. compare the model against doing nothing
  2. what would a constant prediction score
  3. the negative class supplies almost every row
  4. 99.7 is the floor, not the ceiling
  5. only three-tenths of a point of headroom

basics

~20 s

Because a model that never predicts 'defective' already scores 99.7% on that data. The majority class supplies almost every row, so accuracy measures performance on good units and says nothing about whether a single defect was ever caught.

solid answer

~50 s

At a 0.3% defect rate, the constant rule 'always say not-defective' is correct on 99.7% of units, so 99.7% accuracy is the floor, not an achievement. This is the accuracy paradox: accuracy sums over rows, and 997 of every 1000 rows belong to the good-unit class, so the score is dominated by the class nobody is trying to detect. A model could catch zero defects and still hit the headline number. The first thing I do with any accuracy figure is compute the **majority-class baseline** - the share of the largest class - and ask how far above it the model sits. On this data the useful questions are per-class: of the defects that existed, how many were caught, and how many good units got flagged. If the answer to the first is near zero, the 99.7% is describing the base rate, not the model.

go deeper

for a junior

Be ready to state the constant-prediction score out loud from the base rate alone, and to say that 99.7% accuracy at a 0.3% defect rate is the floor. This is a near-universal screening question.

for a middle

Explain the mechanics: accuracy sums over rows, so the majority class owns the score, and perfect detection of the rare class moves the total by at most the base rate. Show the arithmetic rather than asserting the conclusion.

for a senior

Demonstrate the review habit - baseline first, then per-class rates, then the raw confusion matrix - and be able to explain why a model can land below the trivial baseline once false alarms are counted.

for a principal

Own the reporting standard: require every headline classification number to be presented next to its majority-class baseline and the evaluation-set class mix, so a stakeholder deck cannot pass off a base rate as model performance.

## The setup A production line inspects wafers. Historically **0.3% of units are defective** - 3 in every 1000. A classifier is trained to flag defective units and reported at **99.7% accuracy**. The number looks superb and it is completely uninformative. ## Why the number is the floor Consider the dumbest possible model: a constant rule that outputs *not defective* for every unit, ignoring its inputs entirely. On 1000 units it is right 997 times and wrong 3 times. Its accuracy is **99.7%**. So the reported figure is exactly what you get from a rule that has learned nothing and catches no defects at all. The confusion matrix of that constant rule, per 1000 units, is TP = 0, FP = 0, FN = 3, TN = 997. Accuracy = (0 + 997)/1000 = 0.997. Its recall on defects - the share of true defects it caught - is 0/3 = 0. This is the **accuracy paradox**: on skewed data, the score that sounds nearly perfect can coincide with total failure at the actual task. It is not a quirk of one dataset. It follows from the definition. Accuracy sums over *rows*, and when one class owns 99.7% of the rows, the score is essentially a report card on that class: ``` accuracy = prevalence * recall_positive + (1 - prevalence) * recall_negative ``` with prevalence = 0.003, the first term can contribute at most 0.003 to the total. Everything the model does on defects - the entire reason it exists - is capped at three-tenths of one percentage point of the headline number. Even perfect defect detection only moves accuracy from 99.7% to 100%, which is inside the noise of most test sets. ## The majority-class baseline The discipline that prevents this embarrassment is simple and should be automatic: **before reading any accuracy figure, compute the share of the largest class in the same evaluation set.** That is the majority-class baseline, and it is the score of the trivial constant model. Then state the model's accuracy *relative to it*. - Baseline 99.7%, model 99.7% - the model adds nothing measurable. - Baseline 99.7%, model 99.4% - the model is *worse than doing nothing*, which is a real and common outcome once false alarms appear. - Baseline 50%, model 84% - now the accuracy figure is carrying information. The same trap appears wherever a rare event is being predicted. An HR attrition flag on a workforce with a 4% annual quit rate, presented in a stakeholder deck as '96% accurate', is describing a workforce that is 96% stayers. The correct first question in that meeting is not about the model - it is 'what does always predicting *stays* score?' If the answer is also 96%, the deck contains no evidence of a working model. ## Why intuition fails here People read a percentage against an implicit reference of 50%, the coin flip. That reference is only correct when the classes are balanced. Under imbalance the correct reference is the base rate of the majority class, and the usable range of accuracy shrinks accordingly: at a 0.3% defect rate, the entire interesting span of the metric lives between 99.7% and 100%. A metric with three-tenths of a point of headroom cannot rank candidate models meaningfully - ordinary sampling noise in the test set will move it further than a real modelling improvement does. ## What to look at instead Stay per-class. The four cells of the confusion matrix separated by true class answer the questions accuracy has blurred together: - **Of the defects that actually existed, how many did the model catch?** That is the defect-class recall, and for the constant model it is zero. - **Of the good units, how many were needlessly flagged?** That is the false-alarm side, and it determines whether the inspection queue is workable. A single summary that respects both is **balanced accuracy**, the mean of the per-class recalls, whose trivial-model score is 0.5 rather than 0.997. Reporting the raw confusion matrix alongside any scalar is the most robust habit: a reader can then compute whatever they care about, and no rare class can hide inside a rounding error. ## What accuracy is actually good for Accuracy is not a bad metric; it is a metric with a precondition. It is a fair headline when the classes are roughly balanced *and* the two error types cost roughly the same *and* the model emits hard labels at a fixed operating point. Multi-class problems with comparable class sizes - and balanced benchmark datasets, which is where most people first meet the metric - satisfy this, which is exactly why the habit transfers so badly to rare-event problems in production.

  • An HR attrition flag is presented as 96% accurate where 4% of staff leave each year. What is your first question in that meeting?
    'What accuracy does always predicting *stays* get?' The answer is 96%, identical to the model, so the deck shows no evidence the model has learned anything. The follow-up is per-class: of the people who actually left, how many did the flag catch, and how many stayers were wrongly flagged.
  • Could a model score below the majority-class baseline, and what does that mean?
    Yes, easily. Any correct defect it catches adds a tiny amount, while every false alarm on the huge good-unit class subtracts. A model at 99.4% against a 99.7% baseline is trading many false alarms for a few catches - which may still be the right business call, but accuracy is the wrong number to defend it with.
  • When is accuracy still a fair headline metric?
    When the classes are close to balanced, the two error types carry comparable cost, and the model is being reported at one fixed operating point. Balanced multi-class problems are the honest home of accuracy. Outside that, quote per-class rates or a class-balanced summary instead.

Predicting 'no rain' every day in a desert makes you a 99%-accurate forecaster. The score measures the desert, not the forecaster.

saying these in an interview costs you the question

  • Reads 99.7% against an implicit 50% coin-flip reference
  • Never computes the majority-class baseline before judging a score
  • Claims high accuracy proves the rare class is being detected
  • Thinks the paradox is a data bug rather than the metric's definition
  • Assumes a model can never score below the trivial baseline

context