Why can accuracy mislead when scoring prompt candidates on an imbalanced defect set?
answer
- Majority class dominates the number
- Predict-nothing baseline scores high
- Score the class you care about
- Precision and recall on defects
- Harmonic mean punishes both degenerate extremes
basics
~20 sAccuracy is dominated by the majority class. On a manufacturing set that is 2% defective, a candidate that never flags a defect still scores 98%, so the optimizer ranks it top while it fails completely on the class you built the system for.
solid answer
~50 sAccuracy counts every example equally, so when 98% of parts are non-defective it is mostly a measure of how well a candidate says "fine". A prompt that never predicts "defect" scores 98% and can outrank a genuinely useful prompt that catches most defects at the cost of a few false alarms. Because an optimization loop ranks candidates purely by the number you give it, that is not a reporting inconvenience — it is the search actively selecting for the degenerate prompt. The fix is to score the minority class explicitly: precision and recall on "defect", combined into F1, or an F-beta weighted toward whichever error is more expensive. I would also keep the confusion matrix visible for the top candidates so I can see whether a score improved by catching more defects or by cutting false alarms.
go deeper
Be able to state the trap plainly: on a 2% positive set, always answering "no defect" scores 98%. Name precision, recall and F1 on the positive class as what you would score instead.
Explain the mechanics — how F1's harmonic mean punishes both degenerate strategies, what macro versus micro averaging does, and why an automated search amplifies a bad objective that a human reviewer would have caught.
Show you would tie the metric to error costs with an F-beta or a precision floor, inspect confusion matrices for top candidates, and account for the variance of a score resting on a handful of positive examples.
Frame it as choosing what the organization is optimizing for: the weighting between missed defects and false alarms is a cost decision that belongs to the business, and freezing the wrong weighting into an automated loop quietly sets policy for every downstream team.
## The setup Suppose you are optimizing a prompt that classifies inspection notes from a production line as "defect" or "no defect", and the real defect rate is 2%. Your scoring dataset mirrors that rate: 4 defects in every 200 examples. You wire accuracy — the fraction of examples labelled correctly — into the search loop as the objective. ## Why the number lies Accuracy weights every example the same. With 98% of the set in one class, the score is essentially a measurement of performance on the majority class. The pathological case makes it obvious: a prompt that answers "no defect" unconditionally is correct on 196 of 200 examples and scores 98%. A far more useful prompt that catches 3 of the 4 defects but raises 6 false alarms gets 193 of 200, or 96.5%. Ranked by accuracy, the useless candidate wins. In a hand-tuned workflow a human would notice. In an automatic loop nobody does: the search consumes the score, promotes the higher number, and mutations that push candidates toward "never flag anything" get reinforced generation after generation. The optimizer is not broken — it is doing exactly what you asked. ## What to score instead Measure the minority class directly. - **Precision on defects** = of the items the prompt flagged, how many really were defects. Low precision means inspectors waste time on false alarms. - **Recall on defects** = of the real defects, how many the prompt flagged. Low recall means defective parts ship. - **F1** is the harmonic mean of the two, which is the usual single number to hand a search loop because a candidate has to be decent at both to score well. The harmonic mean is deliberately unforgiving: a candidate with recall 1.0 and precision 0.04 (flag everything) scores about 0.077, so the opposite degenerate strategy is punished too. - **F-beta** shifts the balance when the two errors have different costs. On a safety-critical line a missed defect costs far more than a re-inspection, so you weight recall higher; in a high-volume, low-stakes setting you may weight precision higher to keep the alarm queue usable. For multi-class tasks, macro-averaged F1 averages the per-class F1 scores with equal weight per class, which stops rare classes from disappearing into the average; micro-averaging weights by frequency and reintroduces the same majority-class dominance accuracy has. ## Practical guardrails inside the loop Even with F1 as the objective, two things are worth keeping. First, **look at the confusion matrix for the top candidates**, not only the scalar. Two candidates with the same F1 can have completely different error profiles, and which profile you want is a business call the metric cannot make. Second, **watch the noise floor**. With 4 positives in 200 examples, a single extra caught defect swings recall by 25 percentage points. That means F1 on a small imbalanced set is extremely high-variance, and the search will chase noise. The standard responses are to enlarge the number of *positive* examples in the scoring set (oversampling the rare class in the eval set, while remembering the score is then no longer a population estimate), to stratify so every batch contains positives, and to require a candidate's lead to survive re-evaluation before promoting it. ## Common wrong answers Saying "just use F1" without explaining why accuracy fails does not demonstrate understanding. Neither does claiming accuracy is always wrong: on a balanced task with symmetric error costs it is a perfectly reasonable objective and is easier to interpret. The skill is recognizing when class balance and error costs make it a broken objective for an automated search.
- On a defect line, would you optimize toward precision or recall, and how do you encode that?It depends on error costs. A missed defect that reaches a customer usually costs far more than a re-inspection, so I would weight recall higher using an F-beta with beta above 1, and set a precision floor as a hard gate so the queue stays workable. The point is to make the cost asymmetry explicit in the objective rather than hoping F1's equal weighting matches the business.
- What does macro-averaged F1 give you that micro-averaging does not?Macro averaging computes F1 per class and averages with equal weight, so a rare class counts as much as a common one and a candidate cannot hide total failure on it. Micro averaging pools all decisions before computing the score, which weights by frequency and reproduces the same majority-class dominance that makes plain accuracy misleading on skewed data.
- With only four positive examples in the scoring set, what problem remains even after switching to F1?Variance. One extra caught defect moves recall by 25 points, so F1 differences between candidates are mostly noise and the search will chase it. Fix it by enlarging the positive sample — stratifying or oversampling the rare class in the eval set — and by requiring a candidate's lead to hold up on re-evaluation before it is promoted.
saying these in an interview costs you the question
- Reporting 98% accuracy as evidence the prompt works
- Believing high accuracy implies the rare class is handled
- Using F1 without saying which class it is computed on
- Treating accuracy as wrong in all situations, including balanced tasks
- Ignoring how few positive examples the score rests on