skip to content

Soft Targets and Temperature

The student minimizes divergence from the teacher's temperature-softened probabilities, blended with ordinary cross-entropy on the true label. Interviewers ask what the temperature actually changes.

on this pageshow

questions

3

In knowledge distillation, why train the student on the teacher's full probability vector instead of the one-hot label?

level: juniorimportance: must knowfreq 58%

answer

  1. one-hot discards everything except the answer
  2. look at the runner-up scores, not the argmax
  3. which classes get confused with which
  4. input-dependent similarity structure
  5. dark knowledge

basics

~20 s

The teacher's wrong-class probabilities encode which classes resemble each other, structure a one-hot label throws away. Each example then supplies a whole similarity-ranked distribution rather than a single index, giving the student a richer and lower-variance training signal.

solid answer

~50 s

A one-hot label says only which class is right; it says nothing about how the remaining classes relate to the input. The teacher's full softmax vector does. On a fine-grained bird-species classifier, a photo of one warbler puts most mass on the true species but leaves a ranked tail across the visually similar warblers and essentially nothing on a pelican. That ranking is the teacher's learned similarity structure, usually called dark knowledge, and it is the part the student is really copying. Practically, each training example now carries a whole distribution instead of one index, so the per-example gradient is far more informative and less noisy, and the student can approach the teacher's accuracy on a fraction of the data. The training objective blends a KL term pulling the student toward the teacher's softened distribution with ordinary cross-entropy against the true label.

go deeper

for a junior

Be ready to state plainly that the teacher's probability vector ranks the wrong classes by similarity and that a one-hot label cannot, then give one concrete example of two classes that look alike.

for a middle

Expect to describe the two-term objective — KL toward the teacher's softened distribution plus cross-entropy on the hard label — and to explain why a full distribution gives a lower-variance gradient than a single index.

for a senior

Show that you validate distillation by teacher-student agreement and confusion-matrix overlap rather than accuracy alone, and that you think about whether the transfer set matches deployment traffic.

for a principal

Own the argument for when distillation is the right lever at all: the serving-cost case for collapsing an ensemble, what a good teacher is beyond top-1 accuracy, and what you give up by tying a product model to a teacher's behaviour.

## The setup Knowledge distillation trains a small **student** network to imitate a large, already-trained **teacher**. The teacher is frozen: it does a forward pass over the training inputs and its output distribution becomes the target. The student never needs the teacher's weights or architecture, only its outputs on the inputs you feed it — which is why distillation works across completely different model families. The interesting choice is *what* the student is trained against. Two options exist for every example: the recorded label, and the teacher's probability vector. ## What a one-hot label carries A one-hot target is a vector with 1 in the true class and 0 everywhere else. It asserts exactly one fact: this input is class k. It asserts, with equal force, that every other class is *equally* wrong. For a 200-class bird dataset it says a photo of a Blackpoll Warbler is exactly as un-like a Bay-breasted Warbler as it is un-like a pelican. That is false, and the network is forced to fit the falsehood. In information terms, one example delivers at most log2(200) bits, and the gradient it produces mostly pushes one logit up and the rest down uniformly. ## What the teacher's vector carries A converged teacher on the same photo might output 0.86 for the true warbler, 0.09 for the near-identical species, 0.03 for a third warbler, and values around 1e-7 spread over everything else. Read the vector as a *ranking with magnitudes*: it tells the student which classes this input is nearly confusable with and by how much. That is the same relational structure the teacher spent its whole training run discovering, delivered per example, for free. This is what the distillation literature calls **dark knowledge** — the information sitting in the wrong-class scores, invisible if you only ever look at the argmax. Two teachers with identical top-1 accuracy can carry very different dark knowledge, and the one whose runner-up ordering is sensible is the better teacher to distill from. ## Why it actually helps the student Three effects, worth separating: 1. **More signal per example.** A full distribution over C classes constrains the student's whole output layer at once, not just one coordinate. Gradients are correspondingly lower-variance, and distillation is famously data-efficient — a modest transfer set can carry most of the teacher's behaviour. 2. **The targets are input-dependent.** The teacher tells the student something different about *this* image than about the next one. That is what distinguishes soft targets from simply spreading a fixed amount of mass over all wrong classes for every example: the latter injects no similarity structure, because the same flat tail is attached to every input regardless of what it shows. 3. **The teacher has already smoothed the hard cases.** Where the recorded label is ambiguous or plain wrong, a well-trained teacher often hedges or points elsewhere, and the student inherits that hedge instead of the raw label. ## The objective in words The student minimises a weighted sum of two terms: a KL divergence from the teacher's softened distribution to the student's, and ordinary cross-entropy against the recorded hard label. Because the teacher's distribution is fixed, minimising that KL is equivalent to minimising cross-entropy against the teacher's probabilities — the teacher's own entropy is a constant with respect to the student's parameters. The blend weight decides how much the student is allowed to disagree with the recorded labels in favour of the teacher. ## The transfer set The data the teacher is queried on does not have to be the labelled training set. Unlabelled in-domain data works, because the teacher supplies the target — which is often the practical unlock, since unlabelled data is cheap. What matters is that the transfer inputs resemble what the student will see at deployment; querying the teacher on out-of-distribution inputs produces targets that teach nothing useful. ## How you check it worked Accuracy alone is a weak check. The sharper diagnostic is **agreement with the teacher**: on a held-out set, how often does the student's argmax match the teacher's, and how close are the full distributions? A student that matches the teacher's accuracy but disagrees with it on a quarter of examples has learned a different function that happens to score similarly — often a sign the soft term is under-weighted or the targets are too sharp to carry any tail information at all. ## Common misreadings Soft targets are not a regulariser you sprinkle on; they are a different supervision signal whose content depends on the input. And the wrong-class scores are not noise to be argmaxed away — argmaxing the teacher's output before training the student discards the entire reason to distill and reduces the whole procedure to relabelling.

  • Why would you distill an ensemble of five differently-seeded teachers into a single student the size of one member?
    The ensemble's averaged distribution is both more accurate and better calibrated than any member, and its wrong-class tail reflects where the members disagreed — a genuinely richer target than one member's output. Distilling it collapses five forward passes at serving time into one while keeping most of the ensemble's gain, which is the usual reason ensembles are affordable to train but not to deploy.
  • Does the student have to be trained on the same labelled data the teacher saw?
    No. The teacher supplies the target, so any in-domain inputs work, including unlabelled ones — this transfer set is often much larger than the original labelled set and is a large part of why distillation is data-efficient. The requirement is distributional: query the teacher on inputs resembling deployment traffic, since its outputs on out-of-distribution inputs are unreliable targets.
  • How would you check that the student actually absorbed the teacher's dark knowledge rather than just matching its accuracy?
    Measure agreement, not just accuracy: the rate at which student and teacher pick the same class on held-out data, and the divergence between their full distributions. Comparing their confusion matrices helps too — a student that inherited the similarity structure should confuse the same class pairs the teacher does, even where both are wrong.

A one-hot label is a grader writing only WRONG on your paper. The teacher's distribution is a tutor saying you were not merely wrong, you confused this species with the one it most resembles — the mistake itself tells you where the boundary is.

saying these in an interview costs you the question

  • Says soft targets are just label smoothing with a uniform prior
  • Treats the teacher's wrong-class scores as noise to argmax away
  • Claims distillation only compresses and can never help accuracy
  • Thinks you need the teacher's weights rather than just its outputs
  • Assumes the transfer set must be the original labelled training set

context

open as a page

In knowledge distillation, what does dividing teacher and student logits by a temperature above 1 accomplish?

level: middleimportance: should knowfreq 64%

basics

~20 s

Dividing logits by a temperature above 1 flattens the teacher's softmax output, lifting near-zero wrong-class probabilities into a range that actually influences the student's gradient. Teacher and student share that temperature during training; the deployed student uses T = 1.

open as a page

In distillation, how do you weight the teacher-KL term against hard-label cross-entropy when labels are noisy?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

The hard-label cross-entropy term is the only path by which a wrong label reaches the student, so lower its weight and lean on the teacher — but only if the teacher was not itself trained to convergence on the same corrupted labels.

open as a page