In distillation, how do you weight the teacher-KL term against hard-label cross-entropy when labels are noisy?
answer
- one term never sees the label at all
- the hard-label weight is your noise exposure
- ask what the teacher was trained on
- memorised noise gets laundered as confidence
- tune only on a hand-cleaned holdout
basics
~20 sThe hard-label cross-entropy term is the only path by which a wrong label reaches the student, so lower its weight and lean on the teacher — but only if the teacher was not itself trained to convergence on the same corrupted labels.
solid answer
~50 sI treat the hard-label weight as the noise-exposure knob: the ground-truth term is the only channel through which a wrong label reaches the student, so with a few percent of labels known bad I shift weight toward the teacher-KL term and sweep the blend on a small hand-cleaned holdout. The prior question is what the teacher saw. A high-capacity teacher trained to convergence on the same corrupted set may have memorised those examples, in which case its soft targets repeat the errors and raising its weight buys nothing. A teacher trained on cleaner or much larger data, or stopped early, usually still puts real mass on the true class where the label is wrong, and leaning on it is a genuine denoiser. I never tune the blend against a validation split carrying the same label noise, because that rewards a student for reproducing it.
go deeper
Be able to name the two terms in the distillation loss and say which one uses the recorded label and which one uses the teacher's output.
Expect to explain that the hard-label term is the only entry point for label noise, and that the blend weight and temperature must be swept together with the T-squared factor in place.
Show the diagnostic instinct: check whether the teacher memorised the same bad labels, insist on a hand-cleaned holdout for tuning, and mine teacher-student disagreement to find mislabelled examples.
Own the call between compensating for noise in the loss and repairing the labelling pipeline, including what teacher-only training forfeits when labels carry signal the teacher predates.
## The two terms and what each one is for The distillation objective is a blend: `L = alpha * T^2 * KL(teacher_softened || student_softened) + (1 - alpha) * CE(hard_label, student)`. The first term pulls the student toward the teacher's softened distribution; the second anchors it to the recorded ground truth. Alpha is the dial. The key structural observation for noisy data: **the hard-label term is the only place a wrong label enters the objective.** The KL term never sees the label at all — it sees the teacher's output on that input. So `1 - alpha` is, quite literally, your exposure to label noise. That reframing is what an interviewer is looking for; everything else follows from it. A useful default before noise even enters the picture is that distillation setups typically put the *majority* of the weight on the soft term, with the hard-label term contributing a minority share. Starting from there, noisy labels push you further in the same direction rather than into unfamiliar territory. ## The prior question: what did the teacher see? Raising alpha only helps if the teacher's targets are cleaner than the labels. Three cases, and you must know which you are in: 1. **Teacher trained on a cleaner or much larger corpus.** Best case. Where the recorded label is wrong, the teacher usually puts substantial mass on the true class, having generalised from many correctly-labelled similar examples. Raising alpha is a real denoiser, and alpha close to 1 can beat any blend that respects the labels. 2. **Teacher trained on the same corrupted set, but large and regularised, or stopped before convergence.** Networks fit the clean, consistent structure of a dataset before they memorise its inconsistent examples, so an early-stopped teacher often *generalises past* the noise: its prediction on a mislabelled example reflects what similar correctly-labelled examples say, not the wrong label. Raising alpha usually still helps, though less dramatically. 3. **Teacher trained on the same corrupted set to full convergence with enough capacity to memorise.** Its output on a mislabelled training example may be confidently the wrong class. Now the soft targets carry the same errors dressed up as confident predictions, and shifting weight toward the teacher launders noise rather than removing it. This is the trap. A cheap diagnostic separates case 3 from the others: take a small set of examples you *know* are mislabelled, and look at the teacher's distribution on them. Confident agreement with the wrong label is memorisation; mass on the plausible true class is generalisation. ## How to actually pick alpha - **Hand-clean a small holdout.** A few hundred to a couple of thousand carefully re-labelled examples is enough. This is the single highest-value step, because every downstream decision is measured against it. - **Never tune on a validation split drawn from the same noisy pipeline.** A noisy validation set scores a student higher for reproducing the noise, so it will push you toward exactly the wrong alpha — the failure is self-concealing, since your metrics look fine. - **Sweep coarsely**: alpha in roughly 0.3, 0.5, 0.7, 0.9, 1.0, jointly with the temperature, since the two interact — a high temperature with a high alpha spreads mass broadly and can dilute the true class as well as the wrong one. - **Keep the T-squared factor in place** while sweeping, or alpha and temperature will be confounded and the sweep will be uninterpretable. - **Consider alpha = 1 honestly.** Teacher-only training caps the student at the teacher's fidelity, so you forfeit any signal the labels carry that the teacher lacks — new classes, recently corrected annotations, distribution shift the teacher predates. If the labels are only a few percent wrong, that forfeit is often larger than the noise you avoided. ## Per-example refinement A global alpha treats every example as equally suspect, and they are not. Where the teacher confidently disagrees with the recorded label, that example is a strong candidate for being mislabelled. Two productive responses: downweight or drop the hard-label term for exactly those examples, keeping the full hard-label weight everywhere else; or route them to human review and fix the labels. The second is usually the better investment — teacher-student disagreement is one of the cheapest mislabelled-example detectors available, and unlike a loss-weighting trick it improves every model you train afterwards. The honest framing to offer an interviewer: tuning alpha is damage control, not a fix. If the noise is systematic — a whole annotator, a class pair the guidelines never disambiguated, a broken ingestion path — the right move is to repair the labelling process. Blend weights compensate for random noise reasonably well and for systematic noise badly, because a systematic error is likely present in the teacher too. ## What good looks like You should be able to say: which term carries the noise, what the teacher's training history implies about whether it can outvote the labels, that the blend is tuned against clean data only, and that teacher-student disagreement is a signal worth mining rather than a nuisance to smooth over.
- Your teacher was trained on the same mislabelled set. Does raising its weight still help?It depends on whether it memorised the bad examples. A large teacher trained to convergence can fit them confidently, in which case its soft targets repeat the errors and raising alpha launders noise. A regularised or early-stopped teacher usually generalises past them, since networks fit consistent structure before memorising inconsistent examples. Check directly: inspect the teacher's distribution on examples you know are mislabelled.
- How do you validate the blend weight when your validation labels carry the same noise?You hand-clean a holdout. A few hundred carefully re-labelled examples are enough, and there is no substitute — a validation set sharing the training noise rewards a student for reproducing it, so the sweep would select the worst alpha while every metric looked healthy. It is the one place in a noisy-label project where manual annotation effort is unambiguously worth it.
- Would you ever prefer fixing the labels over tuning the blend?Usually, especially for systematic noise. A blend weight compensates reasonably for random errors but poorly for a bad annotator or an ambiguous class pair, because that error is probably in the teacher too. Teacher-student disagreement is a cheap mislabelled-example detector; routing those examples to review fixes the dataset once and improves every model trained on it afterwards.
saying these in an interview costs you the question
- Thinks the KL term can propagate a wrong ground-truth label
- Raises the teacher weight without asking what the teacher trained on
- Tunes the blend on a validation set with the same label noise
- Assumes teacher-only training is strictly safest
- Treats blend weighting as a substitute for fixing systematic labelling errors