skip to content

Model Efficiency and Compression

Shrinking a trained network without wrecking it: teacher-student distillation, magnitude and structured pruning, blocks designed cheap from the start. Interviewers probe the accuracy-for-speed trade.

on this pageshow

explore

questions

page 1 of 2

In knowledge distillation, why train the student on the teacher's full probability vector instead of the one-hot label?

level: juniorimportance: must knowfreq 58%

answer

  1. one-hot discards everything except the answer
  2. look at the runner-up scores, not the argmax
  3. which classes get confused with which
  4. input-dependent similarity structure
  5. dark knowledge

basics

~20 s

The teacher's wrong-class probabilities encode which classes resemble each other, structure a one-hot label throws away. Each example then supplies a whole similarity-ranked distribution rather than a single index, giving the student a richer and lower-variance training signal.

solid answer

~50 s

A one-hot label says only which class is right; it says nothing about how the remaining classes relate to the input. The teacher's full softmax vector does. On a fine-grained bird-species classifier, a photo of one warbler puts most mass on the true species but leaves a ranked tail across the visually similar warblers and essentially nothing on a pelican. That ranking is the teacher's learned similarity structure, usually called dark knowledge, and it is the part the student is really copying. Practically, each training example now carries a whole distribution instead of one index, so the per-example gradient is far more informative and less noisy, and the student can approach the teacher's accuracy on a fraction of the data. The training objective blends a KL term pulling the student toward the teacher's softened distribution with ordinary cross-entropy against the true label.

go deeper

for a junior

Be ready to state plainly that the teacher's probability vector ranks the wrong classes by similarity and that a one-hot label cannot, then give one concrete example of two classes that look alike.

for a middle

Expect to describe the two-term objective — KL toward the teacher's softened distribution plus cross-entropy on the hard label — and to explain why a full distribution gives a lower-variance gradient than a single index.

for a senior

Show that you validate distillation by teacher-student agreement and confusion-matrix overlap rather than accuracy alone, and that you think about whether the transfer set matches deployment traffic.

for a principal

Own the argument for when distillation is the right lever at all: the serving-cost case for collapsing an ensemble, what a good teacher is beyond top-1 accuracy, and what you give up by tying a product model to a teacher's behaviour.

## The setup Knowledge distillation trains a small **student** network to imitate a large, already-trained **teacher**. The teacher is frozen: it does a forward pass over the training inputs and its output distribution becomes the target. The student never needs the teacher's weights or architecture, only its outputs on the inputs you feed it — which is why distillation works across completely different model families. The interesting choice is *what* the student is trained against. Two options exist for every example: the recorded label, and the teacher's probability vector. ## What a one-hot label carries A one-hot target is a vector with 1 in the true class and 0 everywhere else. It asserts exactly one fact: this input is class k. It asserts, with equal force, that every other class is *equally* wrong. For a 200-class bird dataset it says a photo of a Blackpoll Warbler is exactly as un-like a Bay-breasted Warbler as it is un-like a pelican. That is false, and the network is forced to fit the falsehood. In information terms, one example delivers at most log2(200) bits, and the gradient it produces mostly pushes one logit up and the rest down uniformly. ## What the teacher's vector carries A converged teacher on the same photo might output 0.86 for the true warbler, 0.09 for the near-identical species, 0.03 for a third warbler, and values around 1e-7 spread over everything else. Read the vector as a *ranking with magnitudes*: it tells the student which classes this input is nearly confusable with and by how much. That is the same relational structure the teacher spent its whole training run discovering, delivered per example, for free. This is what the distillation literature calls **dark knowledge** — the information sitting in the wrong-class scores, invisible if you only ever look at the argmax. Two teachers with identical top-1 accuracy can carry very different dark knowledge, and the one whose runner-up ordering is sensible is the better teacher to distill from. ## Why it actually helps the student Three effects, worth separating: 1. **More signal per example.** A full distribution over C classes constrains the student's whole output layer at once, not just one coordinate. Gradients are correspondingly lower-variance, and distillation is famously data-efficient — a modest transfer set can carry most of the teacher's behaviour. 2. **The targets are input-dependent.** The teacher tells the student something different about *this* image than about the next one. That is what distinguishes soft targets from simply spreading a fixed amount of mass over all wrong classes for every example: the latter injects no similarity structure, because the same flat tail is attached to every input regardless of what it shows. 3. **The teacher has already smoothed the hard cases.** Where the recorded label is ambiguous or plain wrong, a well-trained teacher often hedges or points elsewhere, and the student inherits that hedge instead of the raw label. ## The objective in words The student minimises a weighted sum of two terms: a KL divergence from the teacher's softened distribution to the student's, and ordinary cross-entropy against the recorded hard label. Because the teacher's distribution is fixed, minimising that KL is equivalent to minimising cross-entropy against the teacher's probabilities — the teacher's own entropy is a constant with respect to the student's parameters. The blend weight decides how much the student is allowed to disagree with the recorded labels in favour of the teacher. ## The transfer set The data the teacher is queried on does not have to be the labelled training set. Unlabelled in-domain data works, because the teacher supplies the target — which is often the practical unlock, since unlabelled data is cheap. What matters is that the transfer inputs resemble what the student will see at deployment; querying the teacher on out-of-distribution inputs produces targets that teach nothing useful. ## How you check it worked Accuracy alone is a weak check. The sharper diagnostic is **agreement with the teacher**: on a held-out set, how often does the student's argmax match the teacher's, and how close are the full distributions? A student that matches the teacher's accuracy but disagrees with it on a quarter of examples has learned a different function that happens to score similarly — often a sign the soft term is under-weighted or the targets are too sharp to carry any tail information at all. ## Common misreadings Soft targets are not a regulariser you sprinkle on; they are a different supervision signal whose content depends on the input. And the wrong-class scores are not noise to be argmaxed away — argmaxing the teacher's output before training the student discards the entire reason to distill and reduces the whole procedure to relabelling.

  • Why would you distill an ensemble of five differently-seeded teachers into a single student the size of one member?
    The ensemble's averaged distribution is both more accurate and better calibrated than any member, and its wrong-class tail reflects where the members disagreed — a genuinely richer target than one member's output. Distilling it collapses five forward passes at serving time into one while keeping most of the ensemble's gain, which is the usual reason ensembles are affordable to train but not to deploy.
  • Does the student have to be trained on the same labelled data the teacher saw?
    No. The teacher supplies the target, so any in-domain inputs work, including unlabelled ones — this transfer set is often much larger than the original labelled set and is a large part of why distillation is data-efficient. The requirement is distributional: query the teacher on inputs resembling deployment traffic, since its outputs on out-of-distribution inputs are unreliable targets.
  • How would you check that the student actually absorbed the teacher's dark knowledge rather than just matching its accuracy?
    Measure agreement, not just accuracy: the rate at which student and teacher pick the same class on held-out data, and the divergence between their full distributions. Comparing their confusion matrices helps too — a student that inherited the similarity structure should confuse the same class pairs the teacher does, even where both are wrong.

A one-hot label is a grader writing only WRONG on your paper. The teacher's distribution is a tutor saying you were not merely wrong, you confused this species with the one it most resembles — the mistake itself tells you where the boundary is.

saying these in an interview costs you the question

  • Says soft targets are just label smoothing with a uniform prior
  • Treats the teacher's wrong-class scores as noise to argmax away
  • Claims distillation only compresses and can never help accuracy
  • Thinks you need the teacher's weights rather than just its outputs
  • Assumes the transfer set must be the original labelled training set

context

open as a page

Why can a network with 5x fewer FLOPs still be slower than its rival on a mobile CPU?

level: middleimportance: must knowfreq 70%

basics

~20 s

FLOPs counts arithmetic only. Latency also pays for moving data, for per-layer fixed costs, and for how well each operation uses the device's parallel units, so a low-FLOP network built from poorly-supported operations can lose on the stopwatch.

open as a page

How does factorizing a 4096x4096 dense layer into two rank-256 matrices cut parameters?

level: middleimportance: must knowfreq 50%

basics

~20 s

One 4096x4096 matrix holds about 16.8M weights; replacing it with a 4096x256 matrix times a 256x4096 matrix holds about 2.1M, an 8x cut. The saving exists only while the rank stays below half the layer width.

open as a page

How do you measure per-layer sensitivity before pruning a trained network?

level: middleimportance: must knowfreq 52%

basics

~20 s

Prune one layer at a time to a fixed ratio, leave every other layer dense, and record the held-out drop. Repeating over several ratios gives a per-layer sensitivity curve that ranks which layers absorb cuts.

open as a page

In unstructured magnitude pruning, which weights are zeroed and why must the model be retrained?

level: middleimportance: must knowfreq 68%

basics

~20 s

Unstructured magnitude pruning zeroes the individual weights with the smallest absolute values, wherever they sit in the tensor, using |w| as a cheap proxy for importance. Retraining with those weights held at zero lets the survivors re-fit what was lost.

open as a page

In quantization-aware training, what does a fake-quantization node do in the forward pass?

level: middleimportance: must knowfreq 55%

basics

~20 s

A fake-quantization node clips a tensor, rounds it onto the low-precision grid, then scales it back to floating point. Values stay float but carry real rounding and clipping error, so the network trains against the error deployment imposes.

open as a page

What makes structured channel pruning of a convolutional network actually reduce latency?

level: middleimportance: must knowfreq 62%

basics

~20 s

Deleting whole filters leaves a smaller dense layer, so every multiply shrinks into a shape the hardware already runs fast, with no special kernel. Fewer operations and fewer bytes moved — but accuracy falls faster per removed parameter.

open as a page

A 90-percent unstructured-sparse model runs no faster than the dense one on CPU — why?

level: seniorimportance: must knowfreq 57%

basics

~20 s

Because the zeros are scattered. Dense matrix kernels multiply every element regardless of value, and a sparse kernel that skips zeros must load an index for each surviving value, whose irregular memory access usually costs more than the multiplies it saves.

open as a page

Why do a recommender's 500-million-parameter embedding tables barely affect its per-request FLOPs?

level: juniorimportance: should knowfreq 50%

basics

~20 s

An embedding table is read, not multiplied. A request looks up a handful of rows, so the table dominates model size and memory while contributing almost no arithmetic; the small dense layers on top do nearly all the compute.

open as a page

Why does an inverted residual block expand the channel count with a 1x1 before projecting back?

level: middleimportance: should knowfreq 42%

basics

~20 s

Because the block's middle operator is cheap per channel, so extra width there costs little while giving the nonlinearity room to work. The expensive parts and the residual stay on the narrow ends, which keeps parameters and stored activations small.

open as a page

How does EfficientNet's compound scaling split a compute budget across depth, width and resolution?

level: middleimportance: should knowfreq 45%

basics

~10 s

Compound scaling raises depth, width and input resolution together as fixed powers of one coefficient, so extra compute is spread over all three dimensions instead of poured into one, which saturates on its own.

open as a page

Why add an intermediate feature-matching loss when distilling into a smaller student?

level: middleimportance: should knowfreq 48%

basics

~20 s

Output-only distillation supervises just the final layer, giving a deep student little guidance on its internal representation. A loss on intermediate feature maps adds per-stage supervision, which helps thin, deep students that are otherwise hard to optimise.

open as a page

In knowledge distillation, what does dividing teacher and student logits by a temperature above 1 accomplish?

level: middleimportance: should knowfreq 64%

basics

~20 s

Dividing logits by a temperature above 1 flattens the teacher's softmax output, lifting near-zero wrong-class probabilities into a range that actually influences the student's gradient. Teacher and student share that temperature during training; the deployed student uses T = 1.

open as a page

Why do compression recipes exempt the input stem and final classifier from aggressive cuts?

level: middleimportance: should knowfreq 46%

basics

~20 s

Both sit at the network's boundary, where an error has no later layer to correct it, and the input stem holds a negligible share of the parameters. Exempting them costs almost no budget while protecting real accuracy.

open as a page

When should sparsity be ramped up gradually during training rather than cut once after training?

level: middleimportance: should knowfreq 45%

basics

~20 s

Ramp gradually when the target sparsity is high enough that a single cut would remove more than the network can recover from. One-shot pruning of the finished model plus a fine-tune is cheaper and adequate at modest sparsity.

open as a page

Why does quantization-aware training need a surrogate gradient for the rounding step?

level: middleimportance: should knowfreq 48%

basics

~20 s

Rounding is a staircase: its derivative is zero almost everywhere, so exact backpropagation would send zero gradient to every quantized weight. The surrogate pretends the rounding step is the identity and passes the incoming gradient straight back.

open as a page

How do you rank a conv layer's channels for pruning: filter norm or activation-based importance?

level: middleimportance: should knowfreq 52%

basics

~20 s

Filter norm ranks channels by the L1 or L2 magnitude of their weights: cheap and data-free, but it assumes magnitude means influence. Activation and first-order Taylor scores use a calibration batch to estimate how much the loss moves when a channel goes.

open as a page

A segmentation network with 60 MB of weights fails on a 4 GB edge device — why?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Peak activation memory, not weight memory, is the binding cost. At high input resolution the intermediate feature maps that must be live at once run to hundreds of megabytes each, and an encoder-decoder holds high-resolution skip tensors alive across the whole decoder.

open as a page

How does a trained weight matrix's singular-value spectrum tell you whether to factorize it?

level: seniorimportance: should knowfreq 35%

basics

~20 s

By how fast the singular values decay. A steep decay means a few directions carry the map, so a low-rank replacement loses little; a flat, well-conditioned spectrum means every direction matters and factorizing at any useful rank will cost accuracy.

open as a page

How would you specify a neural architecture search for an on-device image classifier?

level: seniorimportance: should knowfreq 33%

basics

~20 s

A neural architecture search needs three pieces: the space of operations and connections a candidate may use, an objective that scores it, and a strategy that explores. For an on-device target, score latency measured on that device.

open as a page

Why can a much larger teacher distill worse into a small student than a mid-sized teacher does?

level: seniorimportance: should knowfreq 41%

basics

~20 s

A very large teacher represents a function the small student cannot fit - often not even on training data. Distilled accuracy tracks the teacher-student gap rather than teacher accuracy, so a mid-sized teacher often transfers better.

open as a page

Why judge layer pruning sensitivity after recovery retraining rather than right after the cut?

level: seniorimportance: should knowfreq 38%

basics

~20 s

The instant drop mixes real information loss with recoverable disturbance such as stale normalization statistics. Layers that crater immediately can return to baseline after a short fine-tune, so only the post-recovery ranking predicts the shipped model.

open as a page

What constraint does 2:4 semi-structured sparsity place on a weight matrix, and what does it cost?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Every contiguous group of four weights along the reduction dimension may keep at most two nonzeros, fixing sparsity at exactly 50 percent. The cost is local selection: within an important group you must drop two weights even if all four matter.

open as a page

After you delete an output filter from a conv layer, what else must be removed to keep the network valid?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Deleting a filter also removes its bias, the following normalization layer's scale, shift and running statistics for that channel, and the matching input slice of every consumer. Layers joined by a residual sum must drop identical indices.

open as a page

Your team reports a distilled student beating its from-scratch baseline; what controls do you demand?

level: principalimportance: should knowfreq 34%

basics

~20 s

Demand a from-scratch run of the same architecture with matched epochs, augmentation, tuning budget and seeds, plus a simple-regulariser control. Distillation adds training compute and a smoothing effect, so an unmatched baseline credits ordinary training gains to the teacher.

open as a page

When is quantization-aware training worth a day of retraining over simply converting the model?

level: principalimportance: should knowfreq 38%

basics

~20 s

When the converted model misses a product threshold on the real task metric, cheaper fixes are exhausted, and the deployment target is stable enough that a second training pipeline is re-run rarely. A weekly refresh makes it a bad trade.

open as a page

How do you measure a model's on-device latency so another team can trust the number?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

State the conditions, then measure under load. Fix device, input shape, precision and batch size; discard warm-up iterations; run long enough to expose thermal throttling; and report a median and a 95th percentile rather than a mean or a best run.

open as a page

In born-again self-distillation the student copies the teacher's architecture exactly, so why does it improve?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Nothing is compressed, so the gain cannot come from a smaller model. Training against a trained network's output distribution replaces one-hot labels with smoother, per-example targets that regularise and reweight examples - a training-signal effect.

open as a page

In distillation, how do you weight the teacher-KL term against hard-label cross-entropy when labels are noisy?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

The hard-label cross-entropy term is the only path by which a wrong label reaches the student, so lower its weight and lean on the teacher — but only if the teacher was not itself trained to convergence on the same corrupted labels.

open as a page

When does a learned clipping threshold beat a min-max calibration range in QAT?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Whenever a layer's activations have a long outlier tail. A min-max range stretches to the largest value ever seen, so ordinary activations collapse into a few levels; a trainable threshold clips the tail and buys back resolution.

open as a page

showing 1–30 of 35