skip to content

Under domain shift, why is a network's confidence a poor filter for target pseudo-labels?

level: seniorimportance: should knowfreq 40%

answer

  1. the filter is the weak part, not the idea
  2. a peaked softmax is not a calibrated posterior
  3. high scorers are the most source-like examples
  4. class histogram of accepted labels collapses
  5. confidence rises while accuracy does not

basics

~20 s

Softmax confidence is not a calibrated probability, and shift makes it worse: a network stays confident where it is wrong. A fixed high threshold then keeps the most source-like, easiest examples, skewing the pseudo-label set toward classes the model already handles.

solid answer

~50 s

Self-training on an unlabelled target domain predicts labels, keeps the confident ones and retrains on them, so the filter is the weak point. Deep classifiers are overconfident even in distribution, and under shift the softmax maximum drifts further from a true posterior, so a 0.95 cut admits confidently wrong examples. It is also biased in a specific direction: the target examples clearing a high bar are the ones that most resemble the source, concentrated in the easy, frequent classes. The pseudo-label set therefore adds little information while skewing the class distribution, and each round bakes the model's own early errors in harder. Fix it structurally: select a top proportion within each predicted class, require agreement across augmentations or an ensemble, keep source data in the mix, and hold a small labelled target set for evaluation.

go deeper

for a junior

Remember that a network's confidence and its correctness are different things, and that a wrong pseudo-label enters training as a genuine instruction rather than as noise.

for a middle

Explain both failure mechanisms: overconfidence that worsens under shift, and the selection bias that admits the most source-like, easiest and most frequent classes first.

for a senior

Show the working design — per-class selection, agreement across augmentations or an ensemble, source data retained, a threshold curriculum — plus the label-free monitors that let you stop a run that has begun feeding on itself.

for a principal

Decide when self-training is worth running at all against buying target labels, and set the team's rule that a small labelled target evaluation set is non-negotiable before any unsupervised adaptation programme starts.

## The procedure and where the risk sits Self-training on an unlabelled target domain is simple: run the source-trained network over target data, treat high-scoring predictions as if they were labels, retrain or fine-tune including them, repeat. Everything hinges on the selection step, because a pseudo-label is indistinguishable from a real one once it enters the training set. A wrong pseudo-label is not noise that averages out; it is a supervised instruction to be wrong. ## Why the softmax number is not a probability The maximum softmax output is a monotone score, not a calibrated posterior. Trained to convergence with cross-entropy, deep classifiers push logits apart until predictions are near-saturated, so a value of 0.99 routinely corresponds to an empirical accuracy well below 99 percent. Any calibration measured on held-out **source** data does not transfer: calibration is itself a property of the input distribution, and it degrades under shift in the same direction as accuracy. So the one number the threshold reads is least trustworthy exactly where you are relying on it. ## The selection bias Even if confidence were perfectly calibrated, a global high threshold selects a biased subsample. The target examples that score highest are the ones lying closest to the source distribution — the frames that already look like the training data, the clean audio, the canonical viewpoints. Two consequences follow: 1. **Low information gain.** You retrain on the examples the model already gets right. The hard region — the part of the target that actually caused the accuracy drop — is exactly what the filter excludes. 2. **Class skew.** Confidence is not uniform across classes. Easy, frequent, visually distinctive classes clear a 0.95 bar constantly; rare or ambiguous classes almost never do. The pseudo-label set therefore misrepresents the target's class distribution, and training on it shifts the decision boundaries further toward the majority classes, which raises their confidence again next round. ## The spiral Put those together and the loop is self-reinforcing. Round one admits a confident, easy, class-skewed set including some confident errors. Training on it increases confidence in the same regions — including in the errors, which are now explicitly supervised. Round two's threshold therefore admits more of the same, plus new errors that inherited confidence from the first round's mistakes. Average target confidence climbs monotonically while target accuracy stagnates or falls. The rising-confidence curve looks like progress, which is what makes this failure so effective at surviving review. ## Selection rules that hold up better - **Per-class quotas rather than a global cut.** Take the top fixed proportion of examples *within each predicted class*. The pseudo-label set keeps a sane class balance, and rare classes contribute their best examples instead of none. - **Agreement instead of magnitude.** Require the same prediction across several augmented views of the input, or across an ensemble or across weight-averaged copies of the model. Consistency under perturbation is a far better correctness signal under shift than a single peaked softmax. - **A curriculum on the threshold.** Start strict and loosen over rounds as the model genuinely improves, rather than fixing one number for the whole run. - **Keep source data in the mix.** Continuing to train on labelled source examples anchors the model against drifting into its own fiction. - **Soft targets or loss down-weighting.** Weighting a pseudo-labelled example below a real one limits the damage a wrong one can do. - **Refresh cheap corrections first.** Adapt normalization statistics before generating the first round of pseudo-labels, so the predictions being filtered are the best free predictions available. ## Monitoring it without target labels Even with no labels you have three usable signals. Track the **class histogram of accepted pseudo-labels** across rounds — collapse toward a few classes is the spiral becoming visible. Track the **number of examples clearing the threshold**: an explosive rise means confidence inflation, not learning. Track **prediction churn**, the fraction of target examples whose predicted class changes between rounds; healthy adaptation churns then settles, a spiral freezes early. And if it is at all affordable, hand-label a few hundred target examples for evaluation only. That set never trains anything, and it is the only thing that will tell you the truth about whether the last three rounds helped.

  • What selection rule would you use instead of a single global threshold?
    Select within each predicted class — take a fixed top proportion per class — so the pseudo-label set keeps a plausible class balance and rare classes still contribute. Combine that with an agreement criterion: keep an example only if the prediction is stable across augmented views or across an ensemble. Consistency under perturbation is a much better correctness signal under shift than the size of one softmax value.
  • With no target labels, how would you notice the spiral starting?
    Watch three label-free curves across rounds. The class histogram of accepted pseudo-labels collapsing toward a few classes, the accepted count rising sharply while nothing else improves, and prediction churn freezing early all indicate the model is reinforcing itself rather than learning the target. Average target confidence rising monotonically is the classic tell, since genuine adaptation does not make everything easier at once.
  • Would calibrating the network on held-out source data fix the threshold?
    No. Calibration is a property of a particular input distribution, and it degrades under shift along with accuracy, so a temperature fitted on source data gives no guarantee on target data. Calibration would have to be fitted on labelled target examples — and if you have those, they are usually better spent as an honest evaluation set than as a way to keep a global threshold alive.

saying these in an interview costs you the question

  • Treats the softmax maximum as a calibrated probability
  • Believes a higher threshold alone removes the problem
  • Assumes source-fitted calibration transfers to the target
  • Reads rising average confidence as evidence of adaptation
  • Never checks the class balance of accepted pseudo-labels

context