skip to content

After you delete an output filter from a conv layer, what else must be removed to keep the network valid?

level: seniorimportance: should knowfreq 44%

answer

  1. the edit has a fan-out
  2. normalization keeps four per-channel vectors
  3. elementwise addition constrains the indices
  4. prune the group, not the layer
  5. flatten costs H times W columns

basics

~20 s

Deleting a filter also removes its bias, the following normalization layer's scale, shift and running statistics for that channel, and the matching input slice of every consumer. Layers joined by a residual sum must drop identical indices.

solid answer

~50 s

Removing output filter `j` is a graph edit with a fan-out. Locally you delete the filter's weights, its bias entry, and the per-channel scale, shift, running mean and running variance of the normalization that follows. Downstream, every layer that consumes this feature map loses its `j`-th input slice — and if two layers consume it, both do. Then come the coupling constraints: a residual addition is elementwise, so all tensors entering that sum must have identical channel counts *and* identical surviving indices, which makes the whole residual group a single prunable unit rather than a set of independent layers. If the feature map is flattened into a fully connected layer, one channel costs that layer every column belonging to that channel's spatial positions. In practice you build a dependency graph, prune per group, and prefer the unconstrained interior channels of a block.

go deeper

for a junior

Remember that removing an output channel also removes the matching input channel in the next layer — the two dimensions are the same number and must stay consistent.

for a middle

Be able to list everything attached to one filter: its weights, its bias, and the four per-channel vectors of the normalization that follows, plus every consumer's input slice.

for a senior

Show you have hit this in practice — dependency groups formed by residual sums, the flatten index arithmetic, and the silent mis-indexing bugs that recovery training hides.

for a principal

Own the design decision of pruning at group granularity: which dimensions in the architecture are free, which are coupled, and whether the tooling should refuse cuts it cannot verify.

## Pruning is a graph edit, and edits propagate The seductive mental model is that pruning is a per-layer decision: score this layer's channels, keep the top ones, move on. That model breaks the first time you try to run the pruned network, because a channel index is shared by everything that touches the tensor. **Directly attached to the filter.** Output filter `j` owns a slice of the weight tensor and one bias entry. If a normalization layer follows the convolution, it holds four per-channel vectors — learned scale, learned shift, running mean and running variance — and entry `j` of each must go. Forgetting the running statistics is a classic bug: with luck it raises a shape error, and without luck an off-by-one reindexing quietly applies channel `k`'s statistics to channel `k+1`'s activations, which shows up as an unexplained accuracy collapse that recovery training only partly hides. **Consumers.** Every layer reading this feature map has an input dimension equal to its channel count, so each consumer's kernels lose their `j`-th input slice. Multi-branch architectures have several consumers of the same tensor, and every one of them must be edited consistently. ## Coupling groups: where independence disappears **Residual sums.** An addition is elementwise. Two tensors can only be added if they have the same channel count *and* if channel `i` of one means the same position as channel `i` of the other. So if a block's output is added to a skip path, the block's final layer and whatever produced the skip tensor cannot choose different survivors — they must drop the same index set. In a deep network these constraints chain: a whole stage of blocks whose outputs are added together forms one coupled group with a single shared channel dimension. This is why practitioners usually prune the *interior* channels of a block, which nothing else consumes and which are therefore free, and leave the block's output width alone or prune the entire group with a merged score (for example the sum or maximum of each branch's per-index scores). **Concatenation.** Concatenating branches is friendlier: each branch keeps its own channels, so branches may prune independently. What you must track is the offset bookkeeping — the consumer's input slices for branch two shift once branch one gets narrower. **Flatten into a fully connected layer.** If the last feature map is `C x H x W` and is flattened, the first fully connected layer has `C * H * W` input columns. Removing one channel removes not one column but `H * W` of them — every column belonging to that channel's spatial positions. Get the index arithmetic wrong here and the model still runs, silently permuting features. ## The same constraint outside convolutions Attention heads are the transformer's version of the same story. A head is not one parameter block but four coupled slices: its share of the query projection, of the key projection, of the value projection, and the matching block of rows in the output projection that maps its result back into the residual stream. Removing a head means removing that group together; you cannot delete its value slice and keep its output rows. Pruning heads from a trained encoder with, say, twelve heads per layer is attractive precisely because published studies find that many layers keep most of their accuracy on only one or two surviving heads — the redundancy is real, but the removal is group-shaped. Removing an entire residual block is the coarsest group of all: the block's output must be shape-compatible with the identity path, which is exactly why a block can be dropped wholesale without disturbing anything around it. ## How this is done in practice Build an explicit dependency graph over the network before pruning anything. Nodes are the tensors; edges record producer, consumers, and the operations (addition, concatenation, flatten) that constrain them. Then group the parameters that must be pruned together, score at the level of the *group* rather than the individual layer, and apply one index set per group. What falls out of this analysis is a list of unconstrained dimensions — the ones you can cut freely — and constrained ones that cost you either a group-wide decision or nothing at all. ## Failure modes worth naming A shape error at load time is the good outcome; it is loud. The dangerous outcomes are silent: mis-indexed normalization statistics, an unshifted offset after a concatenation, or a flatten whose column removal was done per channel instead of per channel-position. All three produce a network that runs, produces plausible-looking outputs, and trains back to *nearly* the accuracy you expected — hiding the bug behind recovery training. Verify by checking that the pruned network and the original agree on a batch when the pruned channels are instead zeroed in the original: the outputs should match closely.

  • A residual branch and the skip path rank different channels as least important. What do you do?
    They belong to one coupled group, so a single index set has to serve both. Either merge the scores — sum or average each index's importance across the branches, then prune the merged ranking — or leave the group's shared channel dimension untouched and take the cut from the block's interior channels, which nothing else consumes. The second option is usually preferred because it is unconstrained and costs less accuracy for the same width reduction.
  • How does the same grouping constraint appear when pruning attention heads?
    A head is four coupled slices: its portion of the query, key and value projections, plus the matching block of rows in the output projection that writes it back into the residual stream. They must be removed together; deleting the value slice while keeping the output rows leaves parameters multiplying nothing. The residual stream's own width is untouched, which is what makes head removal a clean, self-contained group.
  • What breaks if you drop the filter but keep the following normalization layer's running statistics for that channel?
    Best case, the shapes disagree and it fails loudly at load time. Worst case, an off-by-one reindexing applies channel k's mean and variance to channel k plus one's activations. The network still runs and still trains, so recovery fine-tuning masks the damage as a mysteriously larger accuracy loss. Verify instead by comparing the pruned model's outputs against the original with those channels zeroed.

Removing a lane from a bridge is not a decision the bridge deck makes alone. Every ramp that feeds it and every lane it merges into has to agree on which lane disappeared.

saying these in an interview costs you the question

  • Prunes a layer in isolation and ignores its consumers
  • Thinks branches of a residual sum can prune different indices
  • Forgets the normalization scale, shift and running statistics
  • Removes one column per channel from the layer after a flatten
  • Assumes every channel in the network is independently prunable

context