skip to content

How do you rank a conv layer's channels for pruning: filter norm or activation-based importance?

level: middleimportance: should knowfreq 52%

answer

  1. cheap and data-free, or data-driven
  2. does the objective get a vote?
  3. activation times gradient
  4. normalization makes filter norm unidentifiable
  5. duplicates fool every per-channel score

basics

~20 s

Filter norm ranks channels by the L1 or L2 magnitude of their weights: cheap and data-free, but it assumes magnitude means influence. Activation and first-order Taylor scores use a calibration batch to estimate how much the loss moves when a channel goes.

solid answer

~50 s

There are three families. The **weight-norm** criterion ranks each output filter by the L1 or L2 norm of its weights and drops the smallest — no data needed, one pass over the parameters. **Activation** criteria run a calibration batch and score a channel by its mean absolute output, or by how often it is zero after a rectified activation, which at least reflects the data. **First-order Taylor** scores approximate the change in loss from zeroing the channel as the average of `|a * dL/da|` over the batch, so the objective, not just signal strength, decides. They disagree in instructive places: a filter with a tiny norm feeding a normalization layer with a large scale is not small at all, and under batch normalization the filter's norm is not even identifiable. Two channels that duplicate each other score high on all three criteria despite being jointly redundant.

go deeper

for a junior

Know that channels are ranked by some importance score and the lowest are removed, and that the simplest score is the magnitude of the filter's weights.

for a middle

Explain each criterion's mechanics: what the norm measures, what a calibration batch gives you, and why activation times gradient estimates the loss change from removing a channel.

for a senior

Demonstrate judgment about failure cases — normalization making the norm unidentifiable, calibration data that does not match deployment, and re-scoring between rounds instead of trusting one ranking.

for a principal

Own the choice between estimating importance after training and learning it during training with a sparsity penalty, and the redundancy problem no per-channel score can see.

## The question every criterion is answering Once you have decided to remove some of a layer's output channels, you must decide *which*. Formally you want the subset whose removal raises the loss least. You cannot evaluate that exactly — the number of subsets is astronomically large — so every practical criterion is a cheap per-channel proxy for it. ## Weight-norm criteria Score output filter *j* by `||W_j||_1` or `||W_j||_2`, the sum or root-sum-of-squares of the weights that produce channel *j*. Rank ascending and drop the smallest. This needs no data, no forward pass and no gradients — one sweep of the parameters — which is why it is the default in almost every quick experiment. The assumption underneath is that a filter with small weights produces a small output and therefore contributes little. That assumption fails in two ways that come up in interviews. First, **scale invariance under normalization**. If a batch normalization layer follows the convolution, it subtracts each channel's mean and divides by its standard deviation, both measured from the data, then applies a learned per-channel scale and shift. Multiply the filter by any constant `c` and the normalized output is unchanged apart from the epsilon in the denominator. The filter's norm is therefore not identifiable: it can be inflated or shrunk by an optimization path that changes nothing about the function. Ranking by it is close to ranking by noise. In such a network the *scale parameter* of the normalization layer is the meaningful magnitude. Second, **input scale**. A small filter sitting on a large-magnitude input can dominate a large filter sitting on a nearly dead one. Norm looks only at weights and never at what flows through them. Raw norms are also not comparable between layers — different layers live at different scales — so a norm criterion is applied *within* a layer, or after normalizing each layer's scores, never as one global sort of raw values. ## Activation criteria Run a calibration batch that resembles deployment data and record statistics of each channel's output feature map: the mean absolute activation, its variance, or the average proportion of entries that are zero after a rectified activation. A channel that is almost always zero, or almost always constant, carries little information forward and is a good candidate. This fixes the input-scale blind spot and, because normalization has already been applied by the time you read the activation, it also sidesteps the identifiability problem. What it still misses is *usefulness*: a channel can be large and stable and still be something the downstream layers ignore. ## Gradient-based (first-order Taylor) criteria Take the loss as a function of a channel's activation `a` and expand around the current value. Removing the channel means driving `a` to zero, and the first-order estimate of the resulting change in loss is `-a * dL/da`. Scoring channels by the average of `|a * dL/da|` over a calibration batch — activation times gradient — therefore estimates directly what you care about: how much the loss moves if this channel goes. This is the only one of the three that consults the objective. A channel with a big activation that the loss is insensitive to gets a small gradient and a low score; a modest channel the loss leans on hard scores high. The costs are that you need labelled data and a backward pass, and that the first-order estimate is only valid for a small perturbation — zeroing an entire channel is not small, so the estimate degrades as you prune harder. Re-scoring between rounds mitigates this. ## Learned importance instead of estimated importance A fourth option changes the training run rather than the analysis. **Network slimming** attaches an L1 penalty to the per-channel scale parameters of the normalization layers during training. The penalty pushes scales that the task does not need toward zero while the task loss holds up the ones it does; afterwards you drop the channels whose scale collapsed and fine-tune. The attraction is that the importance signal is optimized jointly with the task rather than guessed after the fact, and the threshold sits on a single interpretable number per channel. The cost is that you must control the training run, and the penalty strength becomes another thing to tune. ## Where the rankings disagree, and what to do Build the two rankings and look at the disagreements — that is the diagnostic an interviewer wants to hear. Typical patterns: norm and activation disagree wherever normalization scales vary a lot across channels; activation and Taylor disagree wherever a strong signal is being ignored downstream. None of the per-channel criteria handles **redundancy**. Two near-duplicate channels each score high on all three, yet the pair is jointly worth roughly one channel. Catching that needs a set-level criterion — correlation between channel outputs, or choosing the subset that best reconstructs the layer's original output — which is more expensive but noticeably better at aggressive ratios. Practical rules: score within a layer or normalize before comparing across layers; use a calibration set drawn from the deployment distribution, not a convenient shuffled slice; re-score after each round rather than trusting one ranking all the way down; and judge the criterion by accuracy *after* recovery training, since a criterion that looks worse immediately after the cut may recover better.

  • How does network slimming obtain channel importance during training rather than after it?
    It adds an L1 penalty on the per-channel scale parameters of the normalization layers to the training objective. Channels the task does not need have their scales driven toward zero, while useful ones are held up by the task loss. Afterwards you drop the channels whose scale collapsed and fine-tune. The importance signal is optimized jointly with the task, and the threshold sits on one interpretable number per channel — at the cost of owning the training run and tuning the penalty strength.
  • Why is a filter's L2 norm a poor score when batch normalization follows the convolution?
    Normalization divides each channel by its measured standard deviation, so multiplying the filter by any constant leaves the output essentially unchanged. The norm is therefore not identifiable — it can drift during training without the function changing — and ranking by it is close to ranking by noise. Use the normalization layer's learned scale, an activation statistic, or a gradient-based score instead.
  • Two channels are near duplicates of each other. Which criterion notices?
    None of the per-channel ones. Norm, activation and Taylor all evaluate a channel in isolation, so a duplicated pair scores high twice even though the pair is jointly worth about one channel. Detecting it requires a set-level criterion: correlation between channel outputs, or picking the subset of channels that best reconstructs the layer's original output. It is more expensive but pays off at aggressive ratios.

saying these in an interview costs you the question

  • Sorts raw norms across layers without normalizing them
  • Says a near-zero filter norm always means a dead channel
  • Ignores the normalization scale that follows the convolution
  • Scores once and prunes all the way down without re-scoring
  • Assumes the highest-activation channel is the most useful one

context