skip to content

Why can a distilled 3B student beat a 3B model trained from scratch?

level: middleimportance: should knowfreq 38%

answer

  1. one-hot targets are a sparse signal
  2. the wrong answers carry information too
  3. teacher compute is amortised, not free
  4. the student inherits the teacher's ceiling
  5. compression has a capacity floor

basics

~20 s

The student learns from a strong teacher's full output distribution and worked outputs rather than from raw one-hot next tokens, so each training example carries far more signal. Part of the teacher's compute is effectively transferred into the smaller model.

solid answer

~50 s

Training from scratch gives the model a single correct token per position and no information about the alternatives. A teacher instead supplies a probability distribution over the whole vocabulary — which wrong answers were nearly right, which were absurd — and, for reasoning work, complete worked traces the student would rarely stumble into on its own. That is a much denser learning signal per token, so the student converges faster and to a better point at the same student-side budget. Seen as compute allocation, distillation is not free: the teacher's pretraining compute is part of the true total. It pays when that cost is amortised — one teacher distilled into many students, or one student serving enormous traffic. The honest caveat is that the student inherits the teacher's ceiling and its blind spots, and comparisons at genuinely matched *total* compute are much less clear-cut than the headline results suggest.

go deeper

for a junior

Know that distillation means training a small model on a bigger model's outputs, and that this is faster and more effective than learning from raw text alone.

for a middle

Explain the signal-density argument: one-hot targets say only which token was right, while a teacher distribution also says which alternatives were close, so each token teaches more.

for a senior

Frame it as compute allocation. Insist that the teacher's pretraining and generation cost belongs in the total, explain when amortisation makes that worthwhile, and name the limits — teacher ceiling, inherited errors, capacity floor.

for a principal

Own the portfolio question: whether to invest in one strong teacher that seeds a family of specialised students, and what it means strategically that your model line inherits another model's ceiling, biases and licensing terms.

## The mechanism: signal density Standard pretraining is next-token prediction against a one-hot target: for each position, exactly one token is correct and every other token in the vocabulary is equally wrong. That is a very sparse teaching signal. The model must discover, across trillions of tokens, that several alternatives were nearly as reasonable. Distillation replaces or augments that target with the teacher's own probability distribution over the vocabulary. Now each position tells the student not just the answer but the *shape* of the answer space: this token at 0.6, that plausible synonym at 0.25, this category error at 0.001. The relative weights among the wrong answers — often called dark knowledge — encode a great deal of what the teacher learned about similarity and structure. Per training token, the student receives far more bits of supervision. There are variants worth distinguishing: - **Logit or token-level distillation** matches the teacher's per-position distribution. - **Sequence-level distillation** trains on complete teacher-generated outputs, treating them as targets. For reasoning models this means training on full worked traces — the intermediate steps a small model would essentially never produce spontaneously. - **On-policy distillation** samples from the *student*, then has the teacher score or correct those samples, which concentrates supervision on the states the student actually visits rather than the ones the teacher prefers. ## The compute-allocation view The headline claim — a distilled small model beats a same-size model trained from scratch on the same budget — is true and important, but the accounting deserves care. The student's budget is not the total cost. Producing the teacher consumed an enormous pretraining run, and generating the distillation corpus consumes teacher inference at scale. What makes the sum favourable is amortisation. The teacher is trained once and then distilled into many students: different sizes, different specialisations, different deployment targets. Its cost divides across all of them and across all the traffic they serve. From that vantage, distillation is a way of moving compute across the training-inference boundary — you spend it once, at frontier scale, in a form that is then compressed into models cheap enough to serve. This also explains why distillation and over-training are complements rather than alternatives. Both aim at the same target: the strongest possible model at a parameter count you can afford to run. One pushes the data axis; the other improves the quality of the signal on that axis. Production small models typically get both. ## What the student inherits A distilled student is bounded by its teacher in a specific way. It inherits the teacher's factual errors, stylistic tics, refusal behaviour and blind spots, because those are what it was shown. It generally does not exceed the teacher on the distilled distribution. And it can inherit a confident tone without the underlying competence — matching the teacher's surface distribution on cases the teacher happened to get right does not guarantee the same behaviour off-distribution. There is also a capacity floor. Distillation compresses; it does not eliminate the need for parameters. Push the student small enough and there is simply nowhere for the teacher's knowledge to go, and the gap re-opens. Where that floor sits depends on the task — narrow, well-specified tasks compress far better than broad general capability. ## Status in practice As of mid-2026, distilling large reasoning models into small ones is a routine production step rather than a research technique: a frontier model generates or verifies training material, and the small deployed model is trained on it. The interesting open question is not whether it works but how much of the reported advantage survives a genuinely matched total-compute comparison, and whether repeated generations of distillation compound or degrade. ## Answering it well Lead with the signal-density mechanism, since that is the actual reason. Then reframe as compute allocation and be explicit that the teacher's cost is real and only makes sense amortised. Finish with the limits: the teacher's ceiling, inherited errors, and the capacity floor below which compression stops working. A candidate who claims distillation gives you frontier capability for free has missed the accounting.

  • Can a distilled student ever exceed its teacher?
    On the distilled distribution, generally no — its supervision came from the teacher, so the teacher's behaviour is the ceiling. It can beat the teacher on narrower axes: latency, cost, and consistency on a specialised slice where the extra capacity of the teacher was never being used. Genuine capability gains beyond the teacher require an outside signal, such as reinforcement learning against verifiable rewards, not more distillation.
  • How small can the student go before distillation stops helping?
    Until capacity binds. Distillation compresses knowledge into parameters but cannot remove the need for them, so below some size there is nowhere for the teacher's behaviour to live and the gap re-opens. That floor is task-dependent: narrow, well-specified tasks compress far more aggressively than broad general capability. It is an empirical question, answered by sweeping student sizes rather than by a formula.
  • Is it fair to say distillation gives you frontier quality on a small training budget?
    Only if you exclude the teacher from the accounting, which is not honest. The real total includes the teacher's pretraining run plus the inference cost of generating the distillation corpus. What is fair to say is that this cost amortises: one teacher supplies many students and enormous downstream traffic, so per deployed model the allocation is excellent even though the absolute compute is large.

Learning from a marked exam that only says right or wrong is much slower than learning from an expert who shows their working and explains which wrong answers were close.

saying these in an interview costs you the question

  • Says distillation is free because only the student is trained
  • Claims the student routinely surpasses its teacher
  • Thinks distillation only copies outputs, not the distribution
  • Assumes any model can be distilled to any size
  • Ignores that the student inherits the teacher's errors

context