How do you measure per-layer sensitivity before pruning a trained network?
answer
- one layer at a time
- hold every other layer dense
- drop versus ratio, per layer
- find each layer's knee
- marginal effect, not joint
basics
~20 sPrune one layer at a time to a fixed ratio, leave every other layer dense, and record the held-out drop. Repeating over several ratios gives a per-layer sensitivity curve that ranks which layers absorb cuts.
solid answer
~50 sI run a one-layer-at-a-time sweep. For each layer I zero its least important weights to 30, 50, 70 and 90 percent while every other layer stays dense, evaluate on a fixed held-out subset, and restore the layer before moving on. Plotting validation drop against sparsity gives one curve per layer: some stay flat until 80 or 90 percent, others fall off a cliff at 30 percent. The knee of each curve is the ratio I am willing to spend there, and the flat layers are where the compression budget should come from. Two caveats I state up front: the sweep measures a *marginal* effect, so the drops do not simply add up once several layers are cut together, and any drop smaller than the evaluation noise on that subset is not a signal. It costs layers times ratios forward evaluations and no retraining, so it is cheap.
go deeper
Be ready to state the procedure: cut one layer, keep the rest dense, measure on held-out data, restore, repeat. Know that the result is a curve of accuracy drop against compression ratio, one curve per layer.
Explain why isolating one layer at a time is what makes the number interpretable, how to read a knee off the curve, and why the sweep is cheap: it is forward evaluations only, no retraining in the first pass.
Show the operational judgment: fix the evaluation subset, establish metric noise before trusting small drops, and say plainly that a marginal sweep does not predict the joint drop, so the chosen configuration is still validated as a whole.
Own the question of how much measurement is worth buying. Argue when a full sweep pays for itself against a two-tier default, and treat the sensitivity profile as an artifact that must be re-derived whenever the architecture or data distribution moves.
## The problem the sweep solves A uniform compression ratio treats every layer as equally expendable, and networks are not built that way. Some layers are heavily over-parameterized for the function they compute and lose nothing at 90 percent sparsity; others are narrow bottlenecks where removing a third of the weights destroys the model. A *sensitivity sweep* is the cheap measurement that turns that guesswork into a per-layer number before you commit to any cut. ## The procedure Start from the trained network and a fixed held-out subset of data that was never used for training. 1. Pick a layer. Rank its weights by whatever criterion the final recipe will use (typically magnitude) and zero the lowest-ranked fraction, say 50 percent. 2. Leave every other layer untouched at full density. This is the whole point: you are isolating one layer's contribution. 3. Evaluate the task metric on the held-out subset. Record the drop from the uncompressed baseline. 4. Restore the layer's original weights. 5. Repeat for a ladder of ratios, commonly 30 / 50 / 70 / 90 percent, and for every layer in the network. The output is one curve per layer: metric drop on the vertical axis, compression ratio on the horizontal. Overlay them and the network's structure becomes visible at a glance. ## Reading the curves A **flat** curve, where accuracy is unchanged at 90 percent, says the layer is redundant relative to its job. These layers are the donors: the size budget you want to save should come from them first. Flat does not mean the layer is useless — it means its function survives a low-rank, sparse approximation of its weights, not that you can delete it. A **steep** curve, where the metric falls at 30 percent, marks a fragile layer. Fragility is not proportional to size. Layers with few parameters per unit of function are typically the fragile ones — a depthwise convolution holds only a small kernel per channel, so there is nothing redundant to remove, while the wide pointwise layer next to it holds most of the parameters and absorbs heavy pruning. Boundary layers behave the same way for a different reason: their errors have no downstream layer to correct them. The **knee** — the ratio where the curve turns from flat to steep — is the practical per-layer allowance. In practice you set the per-layer ratio a little below its own knee, not at it. ## Cost and noise The first pass needs no retraining, so a 30-layer network at four ratios is 120 evaluations, each one forward pass over a subset. Keeping it honest matters more than keeping it large: - Use the **same** held-out subset for every point. Shared sampling noise then cancels when you compare curves against each other. - Size the subset so that the metric's own run-to-run variability is well below the drops you care about. A 0.2-point difference measured on two thousand examples is noise, not sensitivity, and ranking layers on noise produces a confidently wrong budget. - Never measure on training data — a memorized training metric hides exactly the capacity loss you are trying to detect. - Coarse first. Sweep four ratios across all layers, then refine only near the knees of the layers you intend to cut hard. ## What the sweep does not tell you Two limits are worth saying out loud in an interview, because they are where the naive version of this method fails. First, it is a **marginal** measurement. Each curve answers "what happens if only this layer is cut". When you cut ten layers at once, the perturbations propagate and interact, and the joint drop is generally worse than the sum of the individual drops — layers that were quietly compensating for each other can no longer do so. The sweep gives you a ranking and a starting allocation; the chosen joint configuration still has to be evaluated as a whole. Second, it measures the network **immediately after the cut**, before any adaptation. Part of the instant drop is recoverable disturbance rather than lost information, so the instant ranking and the ranking after a short recovery fine-tune are not the same ordering. ## What good practice looks like Sweep once, read the profile, propose an uneven allocation, cut, retrain, and re-measure the layers you cut hardest. Treat the profile as a hypothesis about where the redundancy lives, and record it with the model — when the architecture or the data changes, the profile changes with it.
- How many evaluations does the sweep cost, and how do you keep it affordable?Layers times ratios, each a single forward evaluation with no retraining, so a 30-layer network at four ratios is 120 runs. Keep it cheap by evaluating on a fixed stratified held-out subset rather than the full validation set, sweeping coarse ratios first, and refining only around the knees of the layers you actually plan to cut hard.
- A layer shows no accuracy drop at 90 percent sparsity. What do you conclude?That the layer is over-parameterized for the function it computes, so it is where the compression budget should come from. It does not mean the layer is dead weight that can be removed entirely: the remaining 10 percent of its weights are still doing the work, and deleting the layer is a different, much larger perturbation.
- How do you avoid mistaking evaluation noise for sensitivity?Fix the held-out subset and reuse it for every point so sampling noise is common to all curves, then establish the baseline metric's variability first. Any per-layer drop smaller than that variability is not evidence. If the interesting differences are within noise, enlarge the subset rather than ranking layers on it.
It is load-testing a bridge one member at a time: you weaken a single strut, see how much the deck sags, then restore it, and only afterwards decide which struts can be made thinner.
saying these in an interview costs you the question
- Ranks layers by parameter count instead of measured accuracy drop
- Prunes every layer to the same ratio and calls it a sweep
- Measures the drop on training data
- Treats a drop smaller than evaluation noise as real sensitivity
- Assumes independent per-layer drops add up when layers are cut jointly