Your team reports a distilled student beating its from-scratch baseline; what controls do you demand?
answer
- one variable at a time
- distillation runs are usually longer
- tune both arms equally
- add a plain-regulariser control
- some gaps are information, not training
basics
~20 sDemand a from-scratch run of the same architecture with matched epochs, augmentation, tuning budget and seeds, plus a simple-regulariser control. Distillation adds training compute and a smoothing effect, so an unmatched baseline credits ordinary training gains to the teacher.
solid answer
~50 sThe claim only means something if the two runs differ in exactly one thing. Distillation runs are usually longer and better tuned than the baseline they are compared against, so I ask for: the identical student architecture; the same data and augmentation; a matched epoch or compute budget; an equal hyperparameter search budget on both sides; and several seeds with the spread reported, since small students are seed-noisy. I also ask for a plain regularised baseline, so we can see how much of the gain is generic target smoothing rather than anything teacher-specific. Finally, the comparison must be on the deployment metric at the deployment latency, not top-line accuracy alone. If the teacher had access to information the student cannot see - many video frames against the student's one - I want that ceiling quantified separately, because no distillation loss can transfer information a single frame does not contain.
go deeper
Recall that any claim about a training trick needs a comparison run that differs in only that trick, trained on the same data for the same number of epochs.
Be ready to list the specific confounds in a distillation comparison: longer schedules, borrowed augmentation policies, and tuning effort spent on one arm only.
Show that you would rerun the baseline rather than reuse an old number, report seed spread, and add a generic-regulariser arm that could plausibly kill your own result.
Own the standard: define what a reportable compression result must contain before the team runs the experiment, and separate gains that are recoverable by training from ceilings set by the student's input, so effort is not spent against a physical limit.
## Why the baseline is the whole claim "Distillation improved the small model" is a causal claim: swapping the teacher term in changed the outcome. It survives only if everything else was held fixed. In practice it usually was not, because the two runs are produced by different people at different times with different levels of care - the distillation run is the interesting experiment and gets attention, while the from-scratch run is the thing someone trained months ago to see if the small architecture was viable at all. ## The controls, in order of how often they are missing **1. Matched training budget.** Distillation runs tend to be longer, because the smoother objective keeps improving after hard-label training has plateaued. If the baseline ran for 90 epochs and the distilled run for 300, part of the gap is simply schedule length. Match epochs, or better, match wall-clock or step budget and report both points. **2. Matched augmentation and data.** The distillation recipe often arrives bundled with a stronger augmentation policy from whatever reference recipe it was copied from. Any augmentation the distilled run gets, the baseline gets. **3. Matched tuning budget.** The distilled run typically has had its loss weight and learning rate searched; the baseline often has neither. Give both the same number of trials. An under-tuned baseline is the single easiest way to manufacture a win. **4. A generic-regulariser control.** Part of distillation's benefit is that the target is smoother than a one-hot label, and that portion is available without a teacher. Run the student with a simple smoothing regulariser at matched budget. If most of the gap closes, what you have discovered is a regularisation effect, not a transfer of anything from the teacher - a much cheaper finding, and one that changes what you build next. **5. Seeds and spread.** Small students are noisy. Three to five seeds per arm with the spread reported, and a gap that is smaller than the seed spread is not a result. **6. The deployment metric, at the deployment operating point.** Top-line accuracy is not usually the shipping criterion. Compare on the metric the product uses, at the latency or memory budget that motivated the small model in the first place, and check calibration and tail-class behaviour - distilled students can trade rare-class recall for overall accuracy in ways an aggregate number hides. ## The information ceiling A control that is easy to forget applies whenever the teacher and student do not see the same input. A concrete case: a video action-recognition teacher that consumes many frames per clip, distilled into a student that sees a single frame, aiming for a cheap real-time model. The teacher's advantage comes substantially from motion, and motion is not present in one frame. Distillation can push the student to use every cue a single frame *does* contain - object identity, pose, scene context, characteristic postures - and that is a genuine and useful effect. What it cannot do is transfer information the student's input does not carry. So the honest report separates two numbers: how much of the teacher's advantage was single-frame-recoverable, and what the residual gap is that no amount of distillation will close. The residual is the argument for either accepting the ceiling or changing the student's input, and confusing the two leads teams to spend quarters tuning a loss against a physical limit. ## What the comparison should look like when it is done properly A reportable result is a table with one row per arm - from-scratch at matched budget, from-scratch with a generic regulariser, distilled - each with mean and spread over seeds, evaluated on the deployment metric, at the same inference cost, with the tuning trials per arm stated. That table survives being read by someone who was not in the room. The single number in a message does not. ## What to do when the controls kill the result This happens, and it is not a failure of the effort. If the matched-budget baseline closes most of the gap, you have learned that the small architecture was under-trained rather than under-informed, which is worth knowing and is cheaper to act on. If the regulariser control closes it, you have a one-line change instead of a teacher-serving pipeline in your training loop. Both outcomes save more engineering time than the original claim would have. ## The judgment being tested An interviewer asking this is checking whether you will defend a compression result under scrutiny, not whether you can recite an experiment protocol. The tell of a strong answer is naming the *specific* ways a distillation comparison is usually unfair - longer schedule, better-tuned run, borrowed augmentation - and insisting on a control that could plausibly kill your own result. The tell of a weak one is treating the from-scratch model as a fixed, already-known number that needs no rerun.
- Why insist on a simple smoothing-regulariser arm rather than only the from-scratch and distilled runs?Because part of distillation's benefit is that the target is smoother than a one-hot label, and that part needs no teacher. The regulariser arm splits the measured gain into a generic component and a teacher-specific one. If the generic component explains most of it, you ship a one-line change instead of running a teacher in the training loop for every future model.
- A many-frame video teacher is distilled into a single-frame student. What extra number do you ask for?The single-frame ceiling: how well any model restricted to one frame can do on this task. Distillation can teach the student to exploit every static cue - pose, objects, scene - but it cannot supply motion the input does not contain. Reporting the recoverable share and the residual gap separately stops the team from tuning a loss against a physical limit.
- The distilled student wins by less than the seed-to-seed spread. What do you do?Treat it as no result and say so. Either run more seeds to see whether a real effect emerges, or accept that on this student and dataset distillation is not buying anything and spend the compute elsewhere. Shipping a training-pipeline dependency on a difference inside the noise band means every future model carries that cost for an unproven gain.
saying these in an interview costs you the question
- Compares against an old, under-trained from-scratch run
- Gives the distilled arm more epochs than the baseline
- Tunes only the distillation run's hyperparameters
- Reports a single seed with no spread
- Never runs a plain regulariser control
- Assumes distillation can transfer information the student's input lacks