skip to content

Pruning and Quantized Training

Cutting parameters out of a network that already trained, by magnitude or by whole channel, and retraining it under simulated low precision. Interviewers probe why sparser rarely means faster.

on this pageshow

explore

questions

17

How do you measure per-layer sensitivity before pruning a trained network?

level: middleimportance: must knowfreq 52%

answer

  1. one layer at a time
  2. hold every other layer dense
  3. drop versus ratio, per layer
  4. find each layer's knee
  5. marginal effect, not joint

basics

~20 s

Prune one layer at a time to a fixed ratio, leave every other layer dense, and record the held-out drop. Repeating over several ratios gives a per-layer sensitivity curve that ranks which layers absorb cuts.

solid answer

~50 s

I run a one-layer-at-a-time sweep. For each layer I zero its least important weights to 30, 50, 70 and 90 percent while every other layer stays dense, evaluate on a fixed held-out subset, and restore the layer before moving on. Plotting validation drop against sparsity gives one curve per layer: some stay flat until 80 or 90 percent, others fall off a cliff at 30 percent. The knee of each curve is the ratio I am willing to spend there, and the flat layers are where the compression budget should come from. Two caveats I state up front: the sweep measures a *marginal* effect, so the drops do not simply add up once several layers are cut together, and any drop smaller than the evaluation noise on that subset is not a signal. It costs layers times ratios forward evaluations and no retraining, so it is cheap.

go deeper

for a junior

Be ready to state the procedure: cut one layer, keep the rest dense, measure on held-out data, restore, repeat. Know that the result is a curve of accuracy drop against compression ratio, one curve per layer.

for a middle

Explain why isolating one layer at a time is what makes the number interpretable, how to read a knee off the curve, and why the sweep is cheap: it is forward evaluations only, no retraining in the first pass.

for a senior

Show the operational judgment: fix the evaluation subset, establish metric noise before trusting small drops, and say plainly that a marginal sweep does not predict the joint drop, so the chosen configuration is still validated as a whole.

for a principal

Own the question of how much measurement is worth buying. Argue when a full sweep pays for itself against a two-tier default, and treat the sensitivity profile as an artifact that must be re-derived whenever the architecture or data distribution moves.

## The problem the sweep solves A uniform compression ratio treats every layer as equally expendable, and networks are not built that way. Some layers are heavily over-parameterized for the function they compute and lose nothing at 90 percent sparsity; others are narrow bottlenecks where removing a third of the weights destroys the model. A *sensitivity sweep* is the cheap measurement that turns that guesswork into a per-layer number before you commit to any cut. ## The procedure Start from the trained network and a fixed held-out subset of data that was never used for training. 1. Pick a layer. Rank its weights by whatever criterion the final recipe will use (typically magnitude) and zero the lowest-ranked fraction, say 50 percent. 2. Leave every other layer untouched at full density. This is the whole point: you are isolating one layer's contribution. 3. Evaluate the task metric on the held-out subset. Record the drop from the uncompressed baseline. 4. Restore the layer's original weights. 5. Repeat for a ladder of ratios, commonly 30 / 50 / 70 / 90 percent, and for every layer in the network. The output is one curve per layer: metric drop on the vertical axis, compression ratio on the horizontal. Overlay them and the network's structure becomes visible at a glance. ## Reading the curves A **flat** curve, where accuracy is unchanged at 90 percent, says the layer is redundant relative to its job. These layers are the donors: the size budget you want to save should come from them first. Flat does not mean the layer is useless — it means its function survives a low-rank, sparse approximation of its weights, not that you can delete it. A **steep** curve, where the metric falls at 30 percent, marks a fragile layer. Fragility is not proportional to size. Layers with few parameters per unit of function are typically the fragile ones — a depthwise convolution holds only a small kernel per channel, so there is nothing redundant to remove, while the wide pointwise layer next to it holds most of the parameters and absorbs heavy pruning. Boundary layers behave the same way for a different reason: their errors have no downstream layer to correct them. The **knee** — the ratio where the curve turns from flat to steep — is the practical per-layer allowance. In practice you set the per-layer ratio a little below its own knee, not at it. ## Cost and noise The first pass needs no retraining, so a 30-layer network at four ratios is 120 evaluations, each one forward pass over a subset. Keeping it honest matters more than keeping it large: - Use the **same** held-out subset for every point. Shared sampling noise then cancels when you compare curves against each other. - Size the subset so that the metric's own run-to-run variability is well below the drops you care about. A 0.2-point difference measured on two thousand examples is noise, not sensitivity, and ranking layers on noise produces a confidently wrong budget. - Never measure on training data — a memorized training metric hides exactly the capacity loss you are trying to detect. - Coarse first. Sweep four ratios across all layers, then refine only near the knees of the layers you intend to cut hard. ## What the sweep does not tell you Two limits are worth saying out loud in an interview, because they are where the naive version of this method fails. First, it is a **marginal** measurement. Each curve answers "what happens if only this layer is cut". When you cut ten layers at once, the perturbations propagate and interact, and the joint drop is generally worse than the sum of the individual drops — layers that were quietly compensating for each other can no longer do so. The sweep gives you a ranking and a starting allocation; the chosen joint configuration still has to be evaluated as a whole. Second, it measures the network **immediately after the cut**, before any adaptation. Part of the instant drop is recoverable disturbance rather than lost information, so the instant ranking and the ranking after a short recovery fine-tune are not the same ordering. ## What good practice looks like Sweep once, read the profile, propose an uneven allocation, cut, retrain, and re-measure the layers you cut hardest. Treat the profile as a hypothesis about where the redundancy lives, and record it with the model — when the architecture or the data changes, the profile changes with it.

  • How many evaluations does the sweep cost, and how do you keep it affordable?
    Layers times ratios, each a single forward evaluation with no retraining, so a 30-layer network at four ratios is 120 runs. Keep it cheap by evaluating on a fixed stratified held-out subset rather than the full validation set, sweeping coarse ratios first, and refining only around the knees of the layers you actually plan to cut hard.
  • A layer shows no accuracy drop at 90 percent sparsity. What do you conclude?
    That the layer is over-parameterized for the function it computes, so it is where the compression budget should come from. It does not mean the layer is dead weight that can be removed entirely: the remaining 10 percent of its weights are still doing the work, and deleting the layer is a different, much larger perturbation.
  • How do you avoid mistaking evaluation noise for sensitivity?
    Fix the held-out subset and reuse it for every point so sampling noise is common to all curves, then establish the baseline metric's variability first. Any per-layer drop smaller than that variability is not evidence. If the interesting differences are within noise, enlarge the subset rather than ranking layers on it.

It is load-testing a bridge one member at a time: you weaken a single strut, see how much the deck sags, then restore it, and only afterwards decide which struts can be made thinner.

saying these in an interview costs you the question

  • Ranks layers by parameter count instead of measured accuracy drop
  • Prunes every layer to the same ratio and calls it a sweep
  • Measures the drop on training data
  • Treats a drop smaller than evaluation noise as real sensitivity
  • Assumes independent per-layer drops add up when layers are cut jointly

context

open as a page

In unstructured magnitude pruning, which weights are zeroed and why must the model be retrained?

level: middleimportance: must knowfreq 68%

basics

~20 s

Unstructured magnitude pruning zeroes the individual weights with the smallest absolute values, wherever they sit in the tensor, using |w| as a cheap proxy for importance. Retraining with those weights held at zero lets the survivors re-fit what was lost.

open as a page

In quantization-aware training, what does a fake-quantization node do in the forward pass?

level: middleimportance: must knowfreq 55%

basics

~20 s

A fake-quantization node clips a tensor, rounds it onto the low-precision grid, then scales it back to floating point. Values stay float but carry real rounding and clipping error, so the network trains against the error deployment imposes.

open as a page

What makes structured channel pruning of a convolutional network actually reduce latency?

level: middleimportance: must knowfreq 62%

basics

~20 s

Deleting whole filters leaves a smaller dense layer, so every multiply shrinks into a shape the hardware already runs fast, with no special kernel. Fewer operations and fewer bytes moved — but accuracy falls faster per removed parameter.

open as a page

A 90-percent unstructured-sparse model runs no faster than the dense one on CPU — why?

level: seniorimportance: must knowfreq 57%

basics

~20 s

Because the zeros are scattered. Dense matrix kernels multiply every element regardless of value, and a sparse kernel that skips zeros must load an index for each surviving value, whose irregular memory access usually costs more than the multiplies it saves.

open as a page

Why do compression recipes exempt the input stem and final classifier from aggressive cuts?

level: middleimportance: should knowfreq 46%

basics

~20 s

Both sit at the network's boundary, where an error has no later layer to correct it, and the input stem holds a negligible share of the parameters. Exempting them costs almost no budget while protecting real accuracy.

open as a page

When should sparsity be ramped up gradually during training rather than cut once after training?

level: middleimportance: should knowfreq 45%

basics

~20 s

Ramp gradually when the target sparsity is high enough that a single cut would remove more than the network can recover from. One-shot pruning of the finished model plus a fine-tune is cheaper and adequate at modest sparsity.

open as a page

Why does quantization-aware training need a surrogate gradient for the rounding step?

level: middleimportance: should knowfreq 48%

basics

~20 s

Rounding is a staircase: its derivative is zero almost everywhere, so exact backpropagation would send zero gradient to every quantized weight. The surrogate pretends the rounding step is the identity and passes the incoming gradient straight back.

open as a page

How do you rank a conv layer's channels for pruning: filter norm or activation-based importance?

level: middleimportance: should knowfreq 52%

basics

~20 s

Filter norm ranks channels by the L1 or L2 magnitude of their weights: cheap and data-free, but it assumes magnitude means influence. Activation and first-order Taylor scores use a calibration batch to estimate how much the loss moves when a channel goes.

open as a page

Why judge layer pruning sensitivity after recovery retraining rather than right after the cut?

level: seniorimportance: should knowfreq 38%

basics

~20 s

The instant drop mixes real information loss with recoverable disturbance such as stale normalization statistics. Layers that crater immediately can return to baseline after a short fine-tune, so only the post-recovery ranking predicts the shipped model.

open as a page

What constraint does 2:4 semi-structured sparsity place on a weight matrix, and what does it cost?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Every contiguous group of four weights along the reduction dimension may keep at most two nonzeros, fixing sparsity at exactly 50 percent. The cost is local selection: within an important group you must drop two weights even if all four matter.

open as a page

After you delete an output filter from a conv layer, what else must be removed to keep the network valid?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Deleting a filter also removes its bias, the following normalization layer's scale, shift and running statistics for that channel, and the matching input slice of every consumer. Layers joined by a residual sum must drop identical indices.

open as a page

When is quantization-aware training worth a day of retraining over simply converting the model?

level: principalimportance: should knowfreq 38%

basics

~20 s

When the converted model misses a product threshold on the real task metric, cheaper fixes are exhausted, and the deployment target is stable enough that a second training pipeline is re-run rarely. A weekly refresh makes it a bad trade.

open as a page

When does a learned clipping threshold beat a min-max calibration range in QAT?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Whenever a layer's activations have a long outlier tail. A min-max range stretches to the largest value ever seen, so ordinary activations collapse into a few levels; a trainable threshold clips the tail and buys back resolution.

open as a page

How do you assign per-layer bit-widths across a network under one total size budget?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Spend bits where the sensitivity curve is steep and take them from where it is flat, equalizing the accuracy lost per bit saved until the size budget is met. Then validate the joint configuration, because per-layer curves were measured independently.

open as a page

A researcher proposes iterative magnitude pruning with rewinding to ship a smaller model — how do you evaluate it?

level: principalimportance: nice to knowfreq 26%

basics

~10 s

Treat the lottery-ticket result as a claim about trainability, not a deployment technique. Finding the mask means training the dense network repeatedly, and the artifact is scattered sparsity that ships no faster than dense.

open as a page

To hit 30 fps on a camera stream, would you channel-prune your trained detector or train a narrower one?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

The answer turns on what you still have: pruning plus fine-tuning is the cheap path, and the only one available without the original data and recipe. For a large cut, a purpose-built narrow architecture trained to convergence usually wins.

open as a page