In unstructured magnitude pruning, which weights are zeroed and why must the model be retrained?
answer
- rank by size, not by structure
- |w| stands in for importance
- a mask, not a smaller tensor
- zeros must be re-applied every update
- survivors have to re-fit the gap
basics
~20 sUnstructured magnitude pruning zeroes the individual weights with the smallest absolute values, wherever they sit in the tensor, using |w| as a cheap proxy for importance. Retraining with those weights held at zero lets the survivors re-fit what was lost.
solid answer
~50 sIt ranks individual weights by absolute value and zeroes the smallest fraction, wherever they fall — no whole row, filter or head has to go, which is why it is called unstructured. The rationale is first-order: a weight contributes `w * x` to its pre-activation, so removing a tiny `|w|` perturbs the layer's output least, provided the inputs it multiplies are on a comparable scale. What you keep is a binary mask, not a smaller tensor; during retraining the mask is reapplied after every update so pruned weights stay at zero and their gradients cannot revive them. Retraining is not optional at any serious sparsity — cutting weights shifts every downstream activation, and only continued training lets the surviving weights absorb that error. The threshold can be computed per tensor or once globally across the network, and the global choice is what can empty a layer whose weights are naturally small.
go deeper
Be ready to say what is removed and what is not: individual weights chosen by absolute value, anywhere in the tensor, with the shape unchanged and a mask recording the choice.
You are expected to explain the w * x argument for why magnitude is a reasonable proxy, and to describe the mechanics of retraining under a mask, including reapplying it after every update.
Show the operational judgment: sweep sparsity to find the knee, guard against a global threshold emptying a layer, and confirm the mask actually held after training rather than trusting the target number.
Own the framing that magnitude is the cheapest point on a saliency-cost curve. Be able to argue when a more expensive importance criterion, or none at all, is the right investment for the model in front of you.
## The operation Magnitude pruning takes a trained network and sets a chosen fraction of its individual weights to zero, selecting them by absolute value: sort the candidate weights by `|w|`, pick a threshold, zero everything below it. It is *unstructured* because the surviving nonzeros can land anywhere — no constraint says a whole row, column, filter or attention head must go together. Three things are produced: the pruned weight tensor, a binary **mask** of the same shape (1 = kept, 0 = pruned), and a **sparsity level** (the fraction of zeros). The tensor's *shape does not change*. A pruned layer still holds an array of the original dimensions, most of whose entries happen to be zero. That single fact explains most of the disappointment people report later about speed. ## Why absolute value at all A weight enters its layer as one term of a sum: the pre-activation is `z = sum_i w_i * x_i + b`. Deleting weight `w_i` changes `z` by exactly `w_i * x_i`. If the inputs `x_i` are on roughly the same scale, then the size of the perturbation is governed by `|w_i|`, and zeroing the smallest ones is the least-damaging single-weight edit you can make for a given number of edits. That is the whole argument, and its assumption is visible in the statement: **magnitude is a proxy for saliency, not a measurement of it.** A weight with a small value that multiplies a systematically large activation matters more than its magnitude suggests, and classical saliency methods instead score weights by how much the loss would rise if they were removed, estimated from curvature. Magnitude survives because it costs nothing to compute and works well enough on over-parameterised networks. The proxy is also only comparable *within a comparable scale*. Two layers can carry weights an order of magnitude apart simply because a normalization layer downstream rescales their outputs; the raw numbers are then not on the same footing. ## Global versus per-tensor thresholds There are two ways to turn a target sparsity into a threshold. - **Per-tensor:** compute a separate threshold inside each weight tensor, so every layer ends at the same sparsity. Nothing can be wiped out, but every layer is treated as equally redundant, which they are not. - **Global:** pool all weights in the network, take one threshold, and let layers land wherever they land. This lets naturally redundant layers give up more — and it has a sharp failure mode. A layer whose weights are all small in absolute terms, for whatever scale reason, can fall almost entirely below a single global threshold and be zeroed out, which severs the network. Small layers, the first layer, and the output projection are the usual casualties. In practice a global threshold is combined with explicit floors: a minimum keep-ratio per tensor, and layers exempted outright. ## Why retraining is not optional A pruned network is a perturbed network. Every zeroed weight changes a pre-activation, that change propagates through a nonlinearity, and errors accumulate layer by layer. At low sparsity the damage may be small enough that accuracy barely moves; on a typical over-parameterised model, accuracy on a held-out set holds up through roughly half the weights and then starts to bend, with a knee — a sparsity beyond which accuracy falls off a cliff — that is empirical and model-specific. Sweeping 50, 80, 90, 95 and 99 percent global sparsity and plotting accuracy is how you find it, and the shape is usually flat, flat, slightly down, down, catastrophic. Retraining (equivalently, fine-tuning under the mask) recovers most of what a cut costs because the surviving weights are free to change. The network does not need the specific weights you removed; it needs *a* function, and the remaining parameters can often express a very similar one once gradient descent is allowed to rearrange them. Mechanically, retraining a pruned network means applying the mask after every optimizer update — or equivalently masking the gradients — so that pruned entries never drift away from zero. If you skip that, weight decay and momentum will quietly repopulate the zeros and your sparsity evaporates. One consequence: in plain magnitude pruning the mask is **monotone**. A weight zeroed at 40 percent sparsity does not come back. Methods that deliberately re-grow connections during training exist and change the mask over time; they are a different algorithm, not a property of magnitude pruning. ## What you should expect to have at the end A network of the original shape, most of whose weights are zero, matching or nearly matching the dense model's accuracy after retraining, that compresses very well on disk. What you do **not** automatically have is a faster model — the zeros are scattered, and whether anything can exploit that is a separate question from whether the accuracy survived.
- Why is absolute weight value only a proxy for importance, and when does it mislead?Removing weight `w_i` changes the pre-activation by `w_i * x_i`, so magnitude only bounds the damage when the activations it multiplies are on a comparable scale. A small weight feeding a systematically large activation matters more than its size suggests, and raw magnitudes are not comparable across layers whose outputs are rescaled by normalization. Curvature-based saliency scores the actual loss increase instead, at much higher cost.
- Can a pruned weight come back during retraining?Not in plain magnitude pruning. The mask is reapplied after every update, so pruned entries stay at zero and sparsity only increases. If you forget to reapply it, weight decay and momentum will move those entries off zero and your sparsity silently disappears. Separate methods do allow re-growth — typically choosing candidates by gradient magnitude — but then the mask itself is being learned, which is a different algorithm.
- What does a single global threshold do that per-tensor thresholds do not?A global threshold pools every weight in the network and lets each layer land at whatever sparsity its own magnitudes imply, so genuinely redundant layers give up more than rigid ones. The price is that a layer whose weights are all small in absolute terms can fall entirely below the threshold and be zeroed out, cutting the network in half. Per-tensor thresholds cannot do that, but treat every layer as equally prunable.
It is closer to muting faders on a mixing desk than to removing channels: the desk is the same size, some channels are silent, and you re-balance the rest by ear afterwards.
saying these in an interview costs you the question
- Says the tensors get smaller, so inference is automatically faster
- Believes pruned weights are deleted rather than masked to zero
- Reports accuracy at high sparsity with no retraining step at all
- Treats a small weight as proof the connection is unused
- Forgets to reapply the mask, so weight decay revives the zeros