Why can a much larger teacher distill worse into a small student than a mid-sized teacher does?
answer
- teacher accuracy is not student accuracy
- the gap, not the level
- student cannot fit training set either
- insert an intermediate-size assistant
- pick teachers by student results
basics
~20 sA very large teacher represents a function the small student cannot fit - often not even on training data. Distilled accuracy tracks the teacher-student gap rather than teacher accuracy, so a mid-sized teacher often transfers better.
solid answer
~50 sTeacher accuracy and student accuracy are different objectives. As the teacher grows, its function becomes more complex and more confidently fit, and a small student has neither the capacity nor the optimisation luck to reproduce it - measurably so, since such students often fail to match a large teacher even on the training set. Past some point the extra teacher quality is unusable and student accuracy falls even while teacher accuracy rises; this is the **capacity gap**. Practical remedies: distil through a **teacher assistant** of intermediate size, so each hop is a gap the student can actually close; use an earlier, less-converged teacher checkpoint; lower the weight on the teacher term; or use an ensemble of mid-sized teachers instead of one giant. The governing rule is to select the teacher by measured *student* accuracy in a short trial, never by the teacher's own leaderboard position.
go deeper
Recall that a bigger teacher does not automatically produce a better student, and that the size difference between the two matters as much as the teacher's own accuracy.
Be ready to explain both causes: the student may be unable to represent the large teacher's function, and it may fail to fit it even on training data, which is an optimisation failure rather than a capacity one.
Show that you diagnose before you fix - measure training-set agreement, try an earlier teacher checkpoint or a smaller teacher first, and select the teacher by short-trial student accuracy rather than by the teacher's own numbers.
Own the cost side: an assistant chain buys accuracy with an extra training run, so set when the team pays for it, and make short teacher-selection trials a standing part of the compression process rather than an ad-hoc rescue.
## The observation The intuitive expectation is monotone: a better teacher should produce a better student. Empirically it is not. Hold the student architecture fixed, distil it from a sequence of progressively larger and more accurate teachers, and student accuracy typically rises, peaks, and then **falls**, even as teacher accuracy keeps climbing. The distance between teacher and student capacity, not the teacher's own quality, is what governs transfer. This is usually called the **capacity gap**. ## Why it happens Three mechanisms, and a good answer distinguishes them. **1. Representational capacity.** The student's hypothesis space simply may not contain a function close to the large teacher's. Asking it to reproduce that function means asking for something it cannot express, and the best it can do is a compromise that may be worse than what it would have found by fitting the task more directly. **2. Optimisation, not just expressiveness.** The sharper and less obvious point: even when the student *could* in principle represent something close, training frequently fails to get there. Students distilled from very large teachers often cannot match those teachers on the **training set** - the objective is not being minimised, quite apart from generalisation. That reframes the gap as partly an optimisation failure, which is why remedies that make the target easier to chase (a smaller intermediate teacher, a less converged teacher) work at all. **3. Target sharpness.** A very large, thoroughly converged network is also a very confident one. Its per-example targets sit close to the extremes, which supplies less usable structure than a moderately confident teacher's targets and reduces the objective towards something close to plain label fitting - at which point you have paid for a teacher and received little. ## Remedies **Teacher assistant chain.** Insert an intermediate-size network: distil the assistant from the large teacher, then distil the small student from the assistant. Each hop spans a gap that the receiving model can actually close. The chain costs an extra training run and adds a size choice to make, and it stops paying once the individual hops are already small. **Use an earlier teacher checkpoint.** A teacher stopped before full convergence is less confident and less contorted, and often distils better than the same teacher at its own accuracy peak. This is cheap - the checkpoints already exist - and it is the first thing to try. **Down-weight the teacher term.** If the teacher's target is partly unusable, give it less influence relative to the task objective and let the student spend its capacity on the task. **Ensemble mid-sized teachers.** Averaging several moderate teachers gives a smoother, better-calibrated target without the extreme sharpness of one giant model. Training cost goes up, but only at training time. **Change what you ask for.** Instead of asking the student to reproduce the teacher's outputs, ask for a weaker correspondence - matching coarse internal summaries, or ranking rather than exact values. A weaker constraint is one a small student can satisfy. ## The selection rule All of this collapses into one operating principle: **choose the teacher by measured student accuracy, not by teacher accuracy.** Run short distillation trials with two or three candidate teachers of different sizes and compare the students on your own validation set. This costs far less than a full training cycle and routinely overturns the choice you would have made from the teachers' own numbers. Record the choice, because it is student-specific - re-shrinking the student later can flip which teacher is best. ## Diagnosing it in your own run The distinguishing measurement is **training-set agreement between student and teacher**. If the student is close to the teacher on training data but poor on validation data, you have a generalisation problem and the capacity gap is not your issue. If the student cannot match the teacher even on training data, you are in the capacity-gap regime, and the remedies above apply. A second useful check is whether the student distilled from a deliberately smaller teacher does better - a two-run experiment that either confirms the diagnosis or eliminates it. ## Common mistakes The most common is treating teacher accuracy as a proxy for transfer quality and reaching for the largest available model on principle. The second is concluding from a single failed distillation that distillation does not work for this student, when the actual finding is that *this teacher* does not work for this student. The third is inserting an assistant chain reflexively: it is a fix for a diagnosed large gap, and when the gap is already modest it just adds a training run and a hyperparameter for no gain.
- What measurement tells you a failed distillation is a capacity gap rather than an overfitting problem?Student-teacher agreement on the **training set**. If the student tracks the teacher closely on training data but not on held-out data, the objective is being met and the problem is generalisation. If the student cannot match the teacher even on training data, the target itself is out of reach - the capacity-gap regime - and the fixes are a smaller or less converged teacher, an assistant hop, or a weaker matching constraint.
- Why does an earlier, less converged teacher checkpoint often distil better than the final one?A fully converged large network is extremely confident and its function is more contorted; both make it a harder target and a less informative one. An earlier checkpoint is smoother and easier for a small student to chase, and the checkpoints already exist, so trying two or three of them is nearly free compared with training an intermediate assistant.
- When is the teacher-assistant chain not worth it?When the diagnosed gap is already modest. Each hop costs a full training run and adds an assistant-size hyperparameter, and if the student can already fit the teacher on training data there is no gap to bridge. Try the cheap moves first - an earlier teacher checkpoint, a smaller existing teacher, a lower weight on the teacher term - and reach for the chain only when those leave a measurable shortfall.
A research professor and a first-year student are both excellent, but the professor's explanations may be too compressed to follow. A graduate teaching assistant who has just crossed the same ground often transfers more, not because they know more, but because the step size is one the listener can take.
saying these in an interview costs you the question
- Assumes the most accurate teacher always yields the best student
- Never measures student-teacher agreement on training data
- Concludes distillation fails when only that teacher failed
- Adds an assistant chain without diagnosing the gap first
- Ignores cheap fixes like an earlier teacher checkpoint