In self-training, how do a model's own pseudo-labels reinforce its errors?
answer
- the model nominates its own training data
- confidence is not correctness
- miscalibration lets errors clear the bar
- fitting an error raises confidence in it
- only real labels can contradict the loop
basics
~20 sSelf-training adds the model's own confident predictions to its training set as ground truth. Some are wrong, and training on them raises confidence in the same mistakes, so more of them clear the threshold each round.
solid answer
~50 sSelf-training loops: fit on the labelled rows, predict the unlabelled pool, promote the predictions above a confidence threshold to pseudo-labels, refit on labelled plus pseudo-labelled, repeat. The trap is that the confidence score is produced by the very model whose errors you are trying to catch, and it is usually not calibrated — a wrong prediction at 0.98 looks identical to a right one. Those errors enter the next round as ground truth, the model fits them, its confidence on similar rows rises, and more of the same errors clear the threshold. Raising the threshold slows the drift but cannot stop it, because it selects for confidence, not correctness. Two symptoms show up fast: accuracy on a truly-labelled validation set falls while training confidence rises, and the pseudo-label class mix drifts toward the majority class.
go deeper
Be able to describe the loop in order — train, predict the unlabelled pool, keep the confident predictions as labels, retrain — and to say that those labels are guesses, not ground truth, so mistakes can be learned as facts.
Explain the mechanism precisely: selection is on confidence, models are typically overconfident, and fitting a wrong pseudo-label raises confidence on similar rows so more of the same error is promoted next round. Mention the majority-class drift too.
Demonstrate operating discipline — a human-labelled validation set that gates every round, per-class quotas instead of one global threshold, re-deriving pseudo-labels rather than accumulating them, and a stopping rule agreed before the run starts.
Own the tradeoff between a cheap self-training programme and paying for labels. Set the evidence standard for shipping a self-trained model, and be clear about the systemic risk: a model that quietly stops predicting a minority class can pass an accuracy gate while failing the business.
## The loop Self-training (also called pseudo-labelling) is the simplest way to use unlabelled data, and it needs no special algorithm: 1. Train a model on the labelled rows. 2. Predict every row in the unlabelled pool. 3. Keep the predictions whose confidence exceeds a threshold, and treat those predicted classes as if they were real labels. 4. Retrain on labelled plus pseudo-labelled rows. 5. Go back to step 2 and repeat, usually a fixed number of rounds. On a call-transcript classifier with 8,000 labelled and 2,000,000 unlabelled transcripts, round one might promote 300,000 transcripts at a 0.95 threshold and the validation score might genuinely improve. Round four is where teams get hurt. ## Why the errors compound The selection rule and the error source are the same object. You are asking the model to nominate the rows it is most sure about, and then you are training it to be even surer about exactly those rows. Three properties turn this into a spiral. **Confidence is not correctness.** A predicted probability of 0.97 is a number the model produced, and most models are miscalibrated — many are systematically overconfident, so the set of rows above the threshold contains more errors than the threshold implies. Crucially, the errors that clear a high threshold are not random noise: they are the model's *systematic* mistakes, the ones it makes consistently on a particular region of the input space. **Training on an error makes it more confident.** Once a mistaken pseudo-label is in the training set, the next fit reduces loss on it, which raises confidence on that row and on its neighbours. Rows that previously sat just below the threshold now clear it — with the same mistake attached. Each round therefore recruits more of the same error, and this is why the loop is described as a death spiral rather than a plateau. **The mistake becomes invisible.** The training loss goes down and the average confidence goes up every round, so every internal signal says things are improving. Only a held-out set with real human labels contradicts it. ## The class-drift symptom There is a second, related failure that shows up as skew rather than as pure error. High-confidence predictions are disproportionately the majority class, because that is where the model has the most evidence. So round after round the pseudo-labelled pool becomes more majority-heavy than the true data, the effective class prior shifts, the decision threshold moves, and the minority class quietly starves. A model that started at a usable minority recall can end up almost never predicting the minority class, while overall accuracy — dominated by the majority — barely moves. Watching per-class pseudo-label counts across rounds catches this immediately. ## Why a higher threshold is not the fix The intuitive response is to raise the threshold from 0.95 to 0.99. That reduces the *number* of pseudo-labels and slows the drift, but the mechanism is untouched: you are still selecting on confidence, which correlates with correctness only as well as calibration allows, and you are still selecting the model's systematic errors preferentially over its random ones. Raise it far enough and you promote only rows the model already handles, which adds no information at all. The threshold trades speed of collapse against usefulness; it does not remove the failure mode. ## What actually helps - **Gate every round on truly-labelled data.** Keep a human-labelled validation set that never receives a pseudo-label, score after each round, and stop the moment it stops improving. This single rule turns the death spiral into a bounded experiment. - **Balance selection by class.** Instead of a single global threshold, take a fixed quota per class, or a fixed top percentile per class, so the pseudo-label mix cannot drift toward the majority. - **Re-derive, do not accumulate.** Recompute the pseudo-label set from scratch each round from the current model, rather than permanently freezing earlier rounds' guesses into the training set. Wrong labels then get a chance to be revised instead of being locked in. - **Down-weight pseudo-labels.** Give them a smaller weight in the loss than human labels, so a wrong one cannot dominate a right one. - **Calibrate first.** Fit a calibration step on held-out labelled data so that "0.95" means something closer to a 95% chance of being right, which at least makes the threshold interpretable. - **Perturb the student.** Training the next round on augmented or noised inputs while the pseudo-labels came from clean inputs makes the loop learn something beyond its own opinion, rather than memorising it. ## The interview framing A strong answer names the loop, names the mechanism (selection on confidence, not correctness, plus fitting reinforces the selected error), names one measurable symptom (validation score falling while training confidence rises, or per-class pseudo-label counts drifting), and gives a mitigation that does not consist solely of moving the threshold. A weak answer treats the threshold as the safety mechanism and never mentions evaluation on real labels.
- What would you monitor across self-training rounds to catch this early?Three things per round: the score on a human-labelled validation set that never receives pseudo-labels, the class distribution of the newly promoted pseudo-labels against the expected prior, and the number of rows promoted. Validation falling while promotions and confidence rise is the signature. If per-class counts skew toward the majority, the loop is drifting even when overall accuracy looks flat.
- Would raising the confidence threshold from 0.95 to 0.99 solve it?No. It slows the drift by promoting fewer rows, but it still selects on confidence rather than correctness, and the errors that clear a very high bar are the model's systematic ones, which are the most damaging to reinforce. Push it high enough and you only promote rows the model already gets right, adding no information. Class-balanced quotas, per-round validation gating and re-deriving pseudo-labels each round address the mechanism; a threshold does not.
- When is self-training a reasonable first choice despite this risk?When the unlabelled pool is enormous relative to the labelled set, the classes are reasonably separated, and you can afford a human-labelled validation set to gate each round. It needs no second view, no rule-writing and no annotation budget, so it is the cheapest thing to try. Run it for few rounds, verify against the supervised baseline, and stop as soon as validation stops improving.
It is a student who grades their own mock exams against an answer key they wrote themselves. Every confident mistake gets copied into the key and revised from again, so the more they study, the surer they are of the wrong answer.
saying these in an interview costs you the question
- Treats the confidence threshold as a safety guarantee
- Assumes predicted probabilities are calibrated by default
- Evaluates the model on its own pseudo-labelled rows
- Accumulates every round's pseudo-labels permanently
- Ignores the class mix of the promoted rows
- Runs self-training with no supervised-only baseline