Why judge layer pruning sensitivity after recovery retraining rather than right after the cut?
answer
- the shipped model is the retrained one
- instant drop mixes two effects
- stale normalization statistics
- equal fine-tune budget per candidate
- rankings reorder after recovery
basics
~20 sThe instant drop mixes real information loss with recoverable disturbance such as stale normalization statistics. Layers that crater immediately can return to baseline after a short fine-tune, so only the post-recovery ranking predicts the shipped model.
solid answer
~50 sWhat you ship is the retrained network, so sensitivity has to be measured on the retrained network. The accuracy right after zeroing weights contains two different things: information the layer genuinely held, and disturbance the network can adapt away — normalization running statistics that no longer match the new activation distribution, shifted scales, biases that are now off-centre. Recalibrating those statistics on a few hundred unlabeled batches, with no gradient steps at all, often recovers a large part of the drop. A short fine-tune recovers more, and layers reorder: a layer that looked catastrophic can come back to baseline while one that looked harmless does not. The practical protocol is to give every candidate an identical, short recovery budget so the comparison is fair, use that as the ranking signal, and then confirm the final chosen configuration with the full retraining schedule.
go deeper
Remember that a pruned network is fine-tuned before it ships, so the accuracy measured one moment after the cut is not the accuracy the model will have.
Explain what the instant drop is made of: genuine lost information plus recoverable disturbance such as stale normalization statistics and shifted activation scales, only the first of which survives fine-tuning.
Show the protocol — recalibrate statistics, give every candidate an identical short recovery budget, use it only as a ranking proxy, then confirm the final configuration with the full retraining schedule.
Own the measurement economics: decide how much of the layers-times-ratios retraining cost the team buys, and require that any published sensitivity profile state the recovery budget it was measured under.
## The trap The cheap version of a sensitivity sweep zeroes a layer's weights, evaluates immediately, and ranks layers by the drop. It is cheap because nothing is retrained. It is also measuring the wrong quantity: nobody ships a network in the state it is in one second after a cut. The deployed model has been fine-tuned afterwards, so the number that matters is how much accuracy is missing *after recovery*, and the two orderings are not the same. ## What the instant drop actually contains The immediate degradation is a sum of two things. **Genuine information loss.** The removed weights were computing something the rest of the network cannot reconstruct. No amount of fine-tuning brings it back, because the capacity is gone. **Recoverable disturbance.** The cut changed the distribution of activations flowing through the network, and several things downstream are now mismatched in ways that a small amount of adaptation fixes: - Normalization layers carry running statistics — the mean and variance estimates collected during training and used at inference. After a cut, the activations they normalize have a different distribution, so the stored statistics are stale and every downstream layer receives badly scaled inputs. Recomputing those statistics by running a few hundred batches of unlabeled data forward, with no gradient steps and no labels, frequently recovers a large fraction of the drop on its own. Reporting a sensitivity number without doing this attributes a normalization bookkeeping problem to the pruning. - Remaining weights and biases are no longer centred for the new activation scale, which a short fine-tune corrects quickly. - The surviving units can often re-specialize to cover part of what was removed, provided they are allowed to train. Only the first component is sensitivity in the sense you care about. The naive protocol measures their sum and calls it sensitivity. ## Why the ranking flips A wide, redundant layer can produce an enormous instant drop — a large perturbation entering a deep stack — and yet recover completely, because the information was distributed and the surviving weights can carry it. A narrow bottleneck can look almost harmless at first and never fully recover, because what it lost was not represented anywhere else. Ranking on instant drop therefore protects the wrong layers: you spend budget defending a layer that would have healed and cut one that will not. ## A protocol that survives review 1. **Recalibrate first.** After each candidate cut, recompute normalization running statistics on a fixed set of unlabeled batches before measuring anything. This is nearly free and removes the largest source of misleading drop. 2. **Give every candidate the same short recovery budget.** A fixed small number of steps at a low learning rate, identical across layers and ratios. Equality is what makes the comparison a comparison; a layer given twice the fine-tuning will look twice as robust. 3. **Keep the proxy honest about what it is.** The short fine-tune is a ranking device, not the shipped recipe. If you let it grow long enough to be a training run in its own right, you are no longer measuring sensitivity — you are measuring whether the architecture can be retrained from a degraded start, which is a different and much less useful question. 4. **Confirm the chosen configuration with the full schedule.** Once the uneven allocation is picked, retrain properly: a fraction of the original schedule, low learning rate with a short warmup, the whole network trainable so the untouched layers can adapt around the cut ones. Then re-measure the layers you cut hardest, because the profile has moved. 5. **Watch the recovery curve, not just its endpoint.** If a configuration's accuracy is still climbing when the recovery budget ends, the comparison is budget-limited rather than capacity-limited and you have not learned what you think you have. ## The cost tension Full recovery retraining for every layer at every ratio is layers times ratios fine-tunes, which is prohibitive for anything real. That is precisely why the two-stage design exists: a cheap no-retrain sweep to eliminate the obviously safe and obviously fatal ratios, then short equal-budget recovery runs on the shortlist, then one full retrain of the final candidate. Interviewers like this question because it separates candidates who have run a compression project from candidates who have read about one — the instant-drop trap costs a real team a week before anyone notices the retrained rankings disagree. ## Recovery after an uneven cut One last practical point specific to uneven allocations: the layers that were cut hardest need the most adaptation, but you do not train them alone. Gradients through the untouched layers are what let the network route around the damage, so the standard recipe trains everything, at a reduced learning rate, for longer than the instinct suggests — heavily compressed regions often keep improving well past the point where accuracy looks converged on the lightly cut ones.
- Which step recovers accuracy after a cut before any gradient step is taken?Recalibrating the normalization layers' running statistics. Their stored mean and variance were estimated on the pre-cut activation distribution and are now stale, so a few hundred forward passes over unlabeled data to recompute them often restores a large share of the drop for essentially no cost. Skipping it blames the pruning for a bookkeeping mismatch.
- Why must the recovery budget be identical across the layers you are comparing?Because fine-tuning length is itself a lever on accuracy. If one candidate gets more steps than another, you are measuring the budget difference rather than the layers. A fixed small number of steps at a fixed low learning rate makes the comparison valid; it just has to stay short enough that it stays a ranking proxy rather than becoming a training run.
- A layer drops accuracy sharply right after the cut but returns to baseline after fine-tuning. What is your conclusion?That it is not sensitive. The removed weights were redundant and the surviving units re-specialized to cover them; the instant drop was a large but recoverable perturbation. It is a candidate to fund the budget, not to protect — though I would confirm the recovery holds under the full retraining schedule and in the joint configuration.
- How do you plan the recovery retraining after an uneven cut across the network?Train the whole network, not only the cut layers, so untouched layers can adapt around the damage. Use a fraction of the original schedule at a reduced learning rate with a short warmup, and run longer than instinct suggests: the heavily compressed regions keep improving after the lightly cut ones have plateaued. Stop when the validation curve is genuinely flat, not when it merely looks close.
Judging sensitivity from the accuracy immediately after a cut is like judging a patient's prognosis in the recovery room. Some of what you see is the anaesthetic wearing off, not the injury.
saying these in an interview costs you the question
- Ranks layers on the accuracy immediately after zeroing weights
- Skips normalization statistic recalibration and blames the pruning
- Gives candidates different fine-tune budgets and compares them
- Assumes retraining always restores the baseline accuracy
- Lets the ranking fine-tune grow into a full training run