skip to content

How does label noise put a ceiling on the accuracy any model can reach?

level: seniorimportance: must knowfreq 55%

answer

  1. wrong labels, wrong scoreboard
  2. a perfect model still disagrees
  3. the ceiling is one minus the error rate
  4. clean the evaluation set first
  5. beating the ceiling means copying mistakes

basics

~20 s

If a share of labels are wrong, even a model that predicts the true answer every time scores below 100% against those labels. Two radiologists disagreeing on one scan set a floor under any model's error.

solid answer

~50 s

Suppose 7% of the labels are wrong. A model that perfectly recovers the truth disagrees with the recorded label on exactly those 7%, so measured accuracy tops out near 93% -- and anything above that means the model has started reproducing the annotators' mistakes rather than the underlying reality. The effect lands on both sides. In training, wrong targets pull the boundary toward regions it should not cover; symmetric flips mostly shrink predicted probabilities toward 0.5 and waste capacity, while class-conditional noise, where positives are systematically marked negative, biases both the boundary and the base rate you calibrate against. In evaluation it is worse, because a noisy test set makes the measurement itself untrustworthy and can invert the ranking of two candidate models. The practical answer is to spend the cleaning budget on a small, adjudicated gold evaluation set first, then treat that ceiling as the target, not 100%.

go deeper

for a junior

Be able to say that if some labels are wrong, no model can score perfectly against them, and that the first question about an accuracy target is how reliable the labels are.

for a middle

Work the arithmetic out loud: an error rate of e caps measured accuracy near 1-e. Explain how wrong targets pull the fitted boundary and why probabilities shrink toward the middle under noise.

for a senior

Demonstrate that you clean the evaluation labels before the training labels, can build and defend an adjudicated gold set, and know how to surface suspect rows with out-of-fold predictions without auto-deleting them.

for a principal

Own the conversation with stakeholders about a target that exceeds the label ceiling, and decide how the annotation budget splits between a gold evaluation set, redundancy, and raw volume.

## The arithmetic of the ceiling Let a fraction `e` of the recorded labels differ from the truth. Take the best possible model -- one that outputs the true label for every case. On every clean row it matches the record; on every corrupted row it disagrees. Its measured accuracy is therefore `1 - e`. That is the ceiling, and it is set by the data, not the algorithm. This is why the first question to ask when a stakeholder demands 99% accuracy is not about the model, it is: how often do two competent annotators labelling the same case produce the same answer? If two radiologists reading the same chest scan disagree on 8% of cases, no model reading those scans can be scored above roughly 92% against single-reader labels -- and a model that scores 96% is not better than the radiologists, it is agreeing with one reader's idiosyncrasies more often than chance. ## Two shapes of noise, two different harms **Symmetric noise** flips labels at the same rate in both directions. Its effect on the *optimal decision rule* is surprisingly mild: as long as the flip rate is below 50%, the class with the higher true probability still has the higher observed probability, so the ideal boundary is unchanged. What does change is the posterior -- the observed probability becomes `(1-e)*p + e*(1-p)`, which pulls every prediction toward 0.5. So calibration degrades even when the ranking survives, and with finite data the learner burns capacity trying to fit points it cannot possibly explain. **Class-conditional noise** flips at different rates per class -- for example, positives are frequently recorded as negatives because the annotator only marks what they are sure of. Now the harm is bias, not just variance. The observed base rate is wrong, the boundary shifts toward the under-reported class, and precision and recall measured against those labels are systematically off in a direction you can predict but not remove without cleaning. ## The evaluation set is the expensive half Most people's instinct is to clean the training set. That is the less important half. A model can tolerate a fair amount of training noise, especially with enough data and a loss that does not chase outliers. But a noisy *test* set corrupts the measurement, and every decision downstream -- which model ships, whether the retrain helped, whether an A/B result is real -- is made from that measurement. Two failure modes matter: - **Downward bias.** A genuinely better model is punished for disagreeing with wrong labels, so your reported number understates it. - **Rank inversion.** If the noise is class-conditional, a model that shares the annotators' bias scores higher than a model that does not. You ship the worse model and cannot see why production disagrees with the offline number. The standard remedy is a **gold set**: take a modest evaluation sample -- often a few thousand rows -- have it labelled independently by two or three of your best annotators, adjudicate every disagreement, and freeze it. Report headline metrics on the gold set and use the large noisy set only for training and for trend monitoring. ## Finding the suspect rows You can locate probable mislabels without any ground truth. Train a model with cross-validation and collect **out-of-fold** predictions, so every row is scored by a model that never saw it. Then rank rows by how confidently the model contradicts the recorded label -- high loss, or a confident probability on the other class. Methods in the confident-learning family formalise this by estimating the joint distribution between the recorded label and the likely true label from the confidently-predicted counts, which gives per-class flip-rate estimates rather than a single ranked list. A simpler neighbourhood check works too: flag a row whose nearest neighbours in feature space are overwhelmingly the other class. The crucial discipline is what you do with the list. **Do not auto-delete it.** Confidently-contradicted rows are a mixture of genuine mislabels, genuinely hard cases, and legitimate minority patterns the model has not learned yet. Deleting all of them removes exactly the tail you most need, and it flatters your metric because you have quietly deleted the hard test. Send the top of the list for re-annotation, measure what share were actually wrong, and let that rate tell you how much noise remains in the rest. ## Training under noise When re-labelling is not affordable for the whole set, the levers are: choose a loss that does not explode on a single contradicted point (absolute or Huber-style errors for regression rather than squared error, which weights a badly wrong target quadratically); stop before the model drives training error to zero, since the last fraction of fit is where memorising wrong labels happens; and downweight rather than delete the suspect rows so an incorrect suspicion is recoverable. ## What to say in the interview Name the ceiling explicitly, tie it to a measured annotator disagreement rate rather than a guess, insist that the evaluation labels are cleaned first, and be honest that above the ceiling the extra points are the model learning the annotators' errors.

  • How would you find the likely mislabelled rows in a million-row training set?
    Cross-validate and keep out-of-fold predictions so every row is scored by a model that never trained on it, then rank rows by how confidently the model contradicts the recorded label. Confident-learning methods extend this into per-class flip-rate estimates. Send the top of the ranking for re-annotation rather than deleting it -- the list mixes true mislabels with genuinely hard cases and rare-but-real patterns you cannot afford to lose.
  • Your model measures 88% and adjudication says roughly 7% of test labels are wrong. Is a 95% target realistic?
    No. With 7% of the labels wrong, a model that recovers the truth every time measures about 93%, so 95% is only reachable by reproducing annotator mistakes. The honest move is to rebuild the evaluation set with adjudicated labels, restate the ceiling to stakeholders, and reset the target beneath it. Chasing the original number rewards a model that has learned the noise.
  • Does symmetric label noise hurt differently from class-conditional noise?
    Yes. Symmetric flips below 50% leave the optimal decision rule intact but pull predicted probabilities toward 0.5, so ranking survives while calibration and sample efficiency suffer. Class-conditional noise -- say positives routinely recorded as negatives -- biases the boundary and the base rate, so precision, recall and any calibrated probability are systematically wrong in a direction that more data will not fix.
  • You can only afford to clean part of the data. Training set or evaluation set?
    Evaluation set, first and without hesitation. A model tolerates a good deal of training noise, but a noisy test set corrupts every decision made from the number: which model ships, whether a retrain helped, whether an experiment moved anything. Build a small adjudicated gold set, report headline metrics on it, and keep the large noisy set for training and trend monitoring.

Grading an exam against an answer key with mistakes in it: the best student in the room still loses marks, and the only way to score full marks is to have made the same mistakes as whoever wrote the key.

saying these in an interview costs you the question

  • Promises 99% accuracy without asking how good the labels are
  • Treats every model-label disagreement as a model error
  • Cleans the training set but leaves evaluation labels noisy
  • Deletes all high-loss rows, discarding the hard cases
  • Assumes more training data cures systematically wrong labels
  • Celebrates beating the annotator disagreement rate

context