In knowledge distillation, what does dividing teacher and student logits by a temperature above 1 accomplish?
answer
- a converged teacher is nearly one-hot
- you need entropy in the target
- divide the logits before exponentiating
- same knob on both networks, training only
- gradient magnitude shrinks, so rescale
basics
~20 sDividing logits by a temperature above 1 flattens the teacher's softmax output, lifting near-zero wrong-class probabilities into a range that actually influences the student's gradient. Teacher and student share that temperature during training; the deployed student uses T = 1.
solid answer
~50 sBoth networks' logits are divided by T before the softmax, so the target becomes `q_i = exp(z_i / T) / sum_j exp(z_j / T)`. A converged teacher is usually close to one-hot, so at T = 1 its wrong-class probabilities sit around 1e-6 and contribute almost nothing to the loss — the similarity structure distillation exists to transfer is effectively invisible. Raising T flattens the distribution and lifts that tail into a range the gradient can see. On a 100-class head, T = 2 and T = 5 typically expose the ranking, while T = 20 pushes everything toward the uniform 0.01 and washes the signal out, so T is a hyperparameter you sweep. The same T is applied to teacher and student for the distillation term only; the hard-label term uses the student's ordinary outputs, and the deployed student runs at T = 1.
code
python · 11 linesimport math
logits = [4.0, 2.5, 2.0, -1.0, -2.0]
def soft_targets(z, T):
e = [math.exp(v / T) for v in z]
s = sum(e)
return [round(x / s, 3) for x in e]
for T in (1, 2, 5, 20):
print("T =", T, soft_targets(logits, T))go deeper
Be ready to write the temperature-scaled softmax and say that T above 1 flattens the distribution while leaving the class ordering untouched.
Expect to explain why a converged teacher's near-zero wrong-class scores are useless at T = 1, where the useful temperature range sits, and that the same T applies to teacher and student during training only.
Demonstrate that you sweep temperature jointly with the blend weight, that you know the soft loss carries a T-squared factor, and that you would suspect a missing rescale when raising temperature appears to do nothing.
Own the framing that temperature trades transferred similarity structure against the teacher's genuine confidence, and be able to argue how that tradeoff shifts with class count, teacher calibration and student capacity.
## The knob The temperature-scaled softmax is `q_i = exp(z_i / T) / sum_j exp(z_j / T)`, where `z` are logits and T is a positive scalar. T = 1 is the ordinary softmax. T > 1 divides every logit down before exponentiating, which shrinks the *gaps* between them and produces a higher-entropy, flatter distribution. T < 1 does the reverse, sharpening toward one-hot. In the limit of very large T every class approaches 1/C. Note what T does *not* do: it never changes the ordering of the classes, because dividing all logits by a positive constant is monotone. Argmax is unaffected. Only the relative magnitudes move. ## Why distillation needs it A well-trained teacher on a training example it has effectively memorised is close to one-hot: 0.9995 on the true class, and the remaining mass spread as values around 1e-6 or smaller. The ranking inside that tail is precisely the dark knowledge you want to transfer, but at those magnitudes it barely moves the loss — the student can match the teacher's top class, drive the rest to zero, and the objective is essentially satisfied. The similarity structure is present in the numbers and absent from the gradient. Raising T rescales the tail into a range the loss responds to. Sweeping a 100-class head illustrates the tradeoff directly: - **T = 1** — near one-hot; the tail is numerically present but contributes almost nothing. - **T = 2** — the top few classes separate from the rest; the confusable neighbours become visible. - **T = 5** — the similarity ranking is clearly expressed across a broad tail; typically the useful range. - **T = 20** — everything is drifting toward the uniform 0.01, the true class is barely distinguished, and the teacher's genuine confidence has been discarded along with the ranking. So T is not monotonically better. Too low and you transfer nothing beyond the label you already had; too high and you transfer near-uniform mush, amplifying whatever arbitrary noise sits in the teacher's smallest logits. It is a hyperparameter, swept on validation, and it interacts with class count — heads with many classes generally tolerate and need less flattening than a small head, where even a modest T pushes you close to uniform. ## Applied where, exactly Three details candidates routinely get wrong: 1. **Both networks, same T.** The teacher's targets and the student's predictions must be softened by the same temperature for the distillation term, or you are asking the student to match a distribution of a different shape. 2. **Only the distillation term.** The hard-label cross-entropy uses the student's ordinary T = 1 outputs against the true label. Softening that term as well would degrade the true-label signal for no reason. 3. **Inference uses T = 1.** Temperature is a training-time device. The deployed student computes its ordinary softmax, and since T does not change the ordering, argmax predictions would be identical anyway — but calibrated probabilities would not be, so leaving T > 1 in the serving path silently under-confidences every output. ## The T-squared rescaling With temperature-softened distributions, the gradient of the soft-target cross-entropy with respect to a student logit takes the form `(1 / T) * (q_i - p_i)`, where `q` is the softened student probability and `p` the softened teacher probability. There is a 1/T out front, and the probability difference itself shrinks as the distributions flatten — so the magnitude of the soft-target gradient falls roughly as 1/T-squared. The consequence is practical: change T and the effective weight of the soft term relative to the hard-label term changes with it, so a blend weight tuned at T = 2 is wrong at T = 5. The standard fix is to multiply the soft-target loss by T-squared, restoring comparable gradient magnitude and making the blend weight roughly temperature-independent. Skipping the rescale is a classic silent bug: raising T appears to stop helping, when in fact you have quietly turned the distillation term down. ## The high-temperature limit If T is made large relative to the logit magnitudes and the logits are zero-mean per example, the softened cross-entropy gradient reduces to a term proportional to the difference between student and teacher logits. In that regime distillation is essentially matching logits directly rather than probabilities. This is a useful thing to know because it clarifies what the middle range buys you: at moderate T the loss still weights the classes the teacher considers plausible more heavily than the ones it has rejected outright, whereas logit matching treats a very negative logit as just as worth reproducing as a competitive one. Very negative logits are typically noisy and constrained by little training signal, so forcing the student to reproduce them wastes capacity. ## Choosing T in practice Start with a small sweep — 2, 4, 8 — on a validation set, jointly with the blend weight, since the two interact. Larger students and larger class counts tend to tolerate higher T. If your teacher was trained with a technique that already produces a high-entropy output, less flattening is needed. And always confirm the T-squared factor is present before concluding a temperature does not help.
- Why is the distillation loss multiplied by T-squared?The gradient of the softened cross-entropy with respect to a student logit carries a 1/T factor, and the probability differences themselves shrink as the distributions flatten, so the soft term's gradient magnitude falls roughly as 1/T-squared. Multiplying the soft loss by T-squared restores comparable magnitude, so the blend weight against the hard-label term stays meaningful when you change temperature instead of silently drifting.
- What temperature should the deployed student use at inference after training at T = 5?T = 1, the ordinary softmax. Temperature is a training-time device for shaping the target. Because dividing logits by a positive constant is monotone, leaving T = 5 in the serving path would not change any argmax prediction, but it would flatten every reported probability and make the model look systematically under-confident to anything downstream that thresholds on those scores.
- What happens to distillation as the temperature grows very large?With logits zero-meaned per example, the softened objective's gradient reduces to a term proportional to the difference between student and teacher logits, so distillation becomes plain logit matching. That is usually worse than a moderate temperature, because it forces the student to reproduce very negative logits that carry little training signal and are largely noise, spending capacity on classes the teacher already rejected.
Temperature is the exposure control on an underexposed photograph. The detail in the shadows was captured, but at default exposure it is indistinguishable from black; turn it up and the structure appears. Turn it up too far and everything is washed to grey.
saying these in an interview costs you the question
- Thinks temperature changes which class the teacher predicts
- Applies the temperature to the teacher only, not the student
- Leaves temperature above 1 in the serving path
- Believes higher temperature is always better for transfer
- Omits the T-squared factor then concludes temperature does not help