skip to content

In born-again self-distillation the student copies the teacher's architecture exactly, so why does it improve?

level: seniorimportance: nice to knowfreq 26%

answer

  1. same size, so not compression
  2. fresh initialisation, trained on teacher outputs
  3. targets vary per example
  4. gains saturate after a few generations
  5. agreement with teacher stays low

basics

~20 s

Nothing is compressed, so the gain cannot come from a smaller model. Training against a trained network's output distribution replaces one-hot labels with smoother, per-example targets that regularise and reweight examples - a training-signal effect.

solid answer

~50 s

In born-again training the second network has the identical architecture and parameter count, is initialised fresh, and is trained against the first network's outputs; it often ends up more accurate than the network that taught it, and a third generation can improve again before gains flatten. Since capacity is unchanged, the explanation cannot be compression. What changed is the target: instead of a fixed one-hot label, each example gets a target that varies with that example, which regularises the fit and effectively reweights confident and ambiguous examples differently. The practical lesson for an interview is that distillation is a *training* technique that happens to also be useful for compression. Two caveats: each generation costs a full training run for a modest gain, and studies measuring student-teacher agreement find fidelity stays low even when test accuracy rises - so the student is not simply reproducing the teacher's function.

go deeper

for a junior

Recall that a network can be trained against another network's predictions instead of the raw labels, and that the second network need not be smaller than the first.

for a middle

Be ready to explain why identical capacity rules out compression as the cause, and to name what changed: a smoother, per-example target rather than a fixed one-hot label.

for a senior

Demonstrate that you know the gain is modest and saturating, that a full training run is the price per generation, and that student-teacher agreement stays low even as accuracy rises.

for a principal

Own the budget call: argue when another generation is a worse use of compute than more data or a longer single run, and set the policy that any claimed gain is credited only against a matched-budget control.

## The setup **Self-distillation** is distillation where the student is not smaller than the teacher. The cleanest version is the *born-again* procedure: train a network normally; freeze it; then train a **second network with the identical architecture**, from a fresh random initialisation, using the first network's outputs as its training target. The result that makes this interesting is that the second network frequently ends up *more accurate on held-out data than the network that taught it*. Repeat, using generation two as the teacher for generation three, and accuracy often improves again for a generation or two before flattening. Averaging the predictions of several generations helps further, though that is an ensemble at inference cost. ## Why the result matters The usual story told about distillation is a compression story: a small student cannot learn the task from labels alone, so it learns from a big teacher that can. Born-again training breaks that story cleanly, because **the student has exactly the same capacity as the teacher**. Whatever is happening, it is not the student being rescued from a capacity shortfall. This is the single most useful thing to say about the result in an interview: distillation is a way of *changing the training signal*, and compression is one application of it rather than its definition. ## What actually changes The only thing that changed is the target each example is trained against. 1. **The target is no longer a hard one-hot vector.** A trained network almost never outputs a degenerate distribution, so the student is fit to a smoother objective. Smoother targets reduce the pressure to drive the correct output to an extreme and generally reduce overfitting to individual examples. 2. **The target is input-dependent.** This distinguishes it from a fixed smoothing scheme applied identically to every example. The teacher's confidence varies from example to example, so easy, unambiguous examples and hard, ambiguous ones are effectively weighted differently in the loss. An example the teacher finds obvious contributes a sharper target and a stronger pull; an example it finds genuinely ambiguous contributes a softer one. This per-example reweighting is a real, measurable difference from any input-independent regulariser. 3. **Optimisation is easier.** The gradient signal from a smooth, self-consistent target has lower variance than the signal from one-hot labels, particularly late in training when a fit to hard labels keeps pushing outputs further apart with little accuracy left to gain. Note what is *not* in this list: the student is not inheriting a better architecture, more parameters, more data, or more supervision than the teacher had. It sees the same inputs; only the labels it is asked to reproduce changed. ## The fidelity result A sharp piece of evidence sits alongside this. If you measure how often a distilled student's *prediction* agrees with its teacher's prediction - fidelity, as distinct from accuracy - the agreement is often surprisingly low, and it stays low even when the student's own test accuracy has improved. Equal-capacity students frequently fail to match their teacher's function even on the training set. Two conclusions follow. First, distillation in practice is not solved function-matching; it is an optimisation problem that is only partly solved, and the residual is not simply a capacity shortfall. Second, accuracy gains and fidelity gains are different objectives - if you actually need the student to *behave like* the teacher (say, because downstream systems were tuned against the teacher's mistakes), you must measure agreement directly, because accuracy will not tell you. ## Related variants The same idea appears in other shapes. One is a within-network form, where deeper parts of a single network supply targets to shallower exits during a single training run, so no second run is needed. Another is generational training on unlabelled or weakly labelled data, where each generation labels more data for the next. All share the property that the teacher is not larger than the student, so none of them can be explained by compression. ## Costs and honest caveats - **Each generation is a full training run.** The accuracy gain is real but modest; on a fixed compute budget you should ask whether another generation beats spending the same compute on data, augmentation, or a longer single run. - **Gains saturate.** Beyond two or three generations, improvements typically stop, and the sequence can drift if the teacher's errors are systematic - each generation inherits and can amplify the previous one's biases, especially on rare classes. - **Attribution requires a control.** Because the mechanism is partly generic regularisation, a claimed born-again gain should be compared against the same architecture trained with a comparable simple regulariser and a matched compute budget before it is credited to the teacher specifically. ## The interview answer State the setup precisely (same architecture, fresh initialisation, trained on the teacher's outputs), state that identical capacity rules out compression as the explanation, and name the mechanism as input-dependent target smoothing that regularises and reweights. Then add the fidelity nuance - the student does not actually reproduce the teacher's function - and the cost caveat. That combination shows you have read past the headline.

  • If the mechanism is just smoother targets, why not apply a fixed smoothing scheme and skip the second training run?
    Because a fixed scheme spreads the same off-target mass over every example, while a trained network's target varies with the input - it is sharper on unambiguous examples and softer on genuinely ambiguous ones, which reweights the loss per example. Fixed smoothing captures part of the gain, which is exactly why it belongs in your control run, but it is not equivalent.
  • How many generations would you run in practice, and what stops you?
    Two or three. The first generation gives most of the gain, the next is smaller, and after that improvements typically vanish while each generation still costs a full training run. Stop when the held-out gain no longer exceeds run-to-run seed variance - and watch rare-class metrics, since successive generations can inherit and amplify the previous teacher's systematic errors.
  • Your distilled student is more accurate than the teacher but disagrees with it on many inputs. Is that a problem?
    It depends on what depends on the teacher. Higher accuracy with low agreement is the normal outcome and is fine if the student is judged on its own metrics. It is a problem when downstream thresholds, rules or human workflows were calibrated against the teacher's specific error pattern - then you must measure and optimise agreement explicitly, because accuracy will hide the shift.

saying these in an interview costs you the question

  • Says the student improves because it has fewer parameters
  • Claims distillation only makes sense for compression
  • Assumes the student reproduces the teacher's predictions closely
  • Expects gains to keep growing with more generations
  • Credits the teacher without a matched-budget control run

context