Why do compression recipes exempt the input stem and final classifier from aggressive cuts?
answer
- cheap to exempt, expensive to cut
- stem is a rounding error in size
- logits have nothing downstream
- check the head's parameter share
- a prior, not a law
basics
~20 sBoth sit at the network's boundary, where an error has no later layer to correct it, and the input stem holds a negligible share of the parameters. Exempting them costs almost no budget while protecting real accuracy.
solid answer
~50 sIt is a cost-benefit asymmetry. The stem holds very few weights — a first convolution over three input channels is a rounding error in a network of tens of millions — yet every downstream feature is built on what it produces, so damage there propagates through the whole network. The final classifier converts features into logits, and the decision margin between the top two classes is often small, so quantization or pruning error there lands directly on the prediction with nothing downstream to absorb it. Keeping them dense, or at 8 bit while the body goes to 4, costs a fraction of a percent of the budget and recovers a disproportionate share of the accuracy. I still treat it as a prior, not a law: with many classes or a wide embedding the head is not small, so I check its parameter share and confirm the exemption against my own sweep.
go deeper
Recall the rule and one reason for it: those two layers touch the raw input and the final logits, and the input stem is tiny, so leaving them alone costs very little size.
Give both halves of the argument — negligible parameter share on one side, error with nothing downstream to correct it on the other — and explain why the classifier's small decision margin makes it especially exposed.
Demonstrate that you verify rather than inherit the rule: compute the head's share, confirm against your own sensitivity curves, and report which layers were exempted whenever you quote a compression ratio.
Own where the exemption list stops. Set the policy that exemptions are budgeted and justified per model family, so the practice does not drift into keeping half the network dense and calling the result compressed.
## The rule and why it exists Almost every practical compression recipe carries the same exception: the first layer and the last layer are not cut as hard as the middle of the network. In pruning that means leaving them dense or at a much lower sparsity; in low-precision training it means keeping them at 8 bit while the body goes to 4. The reason is an asymmetry between what those layers cost you in budget and what they cost you in accuracy. ## The parameter-share argument Compression budgets are measured in parameters or bits, and the input stem holds almost none of them. A convolution over a three-channel image produces a modest number of output channels from a small kernel — on the order of ten thousand weights in a network whose total is tens of millions. Exempting it changes the achieved compression ratio by well under a percent. Whatever accuracy it protects is therefore nearly free, which is what makes the exemption an easy call rather than a tradeoff. The final classifier deserves the same arithmetic rather than the same reflex. In a network that ends in a projection from a feature vector to a modest number of classes, the head is also small and the exemption is nearly free. But the head's size is the feature width times the number of outputs, so with thousands of classes, or a wide embedding, or a dense output over a large vocabulary, it can be a substantial share of the model — and then exempting it genuinely spends budget you wanted elsewhere. The discipline is to compute the share before applying the rule. ## The error-propagation argument The cost side is about position in the network, not size. The stem sees the **raw input**. Its outputs are the only representation everything else has access to, so a distortion introduced there is inherited by every later layer and compounds as it travels. There is also less redundancy to exploit: the layer maps a handful of input channels through a small kernel, so there is not much slack in its weights to begin with. The final layer emits the **logits themselves**. Every other layer's error can, in principle, be partly compensated by the layers after it — the network's later stages have seen this kind of perturbation during any recovery fine-tuning and can adapt. The classifier has nothing after it. The gap between the top two logits is frequently small, so a perturbation of similar magnitude flips the predicted class outright. The last layer's output distribution also tends to have a different dynamic range from hidden activations, which makes a single global quantization scale a poor fit for it. ## Fragility is about redundancy, not depth The boundary rule is one instance of a broader pattern: layers with little redundancy relative to their function are the fragile ones. A depthwise convolution, which applies one small kernel per channel with no mixing, has very few parameters for the work it does, and sensitivity sweeps regularly show it collapsing at ratios the neighbouring wide pointwise layer shrugs off — even though the pointwise layer holds far more of the model's parameters. Reading "first and last" as the complete list of fragile layers misses that. ## How to use the rule honestly Treat it as a strong prior that saves you sweep time, not as a substitute for measurement: - Compute the parameter share of the stem and the head first. If the head is large, decide deliberately whether the exemption is worth its budget, or exempt it partially — for example a higher bit-width rather than full precision. - Confirm with the sensitivity curves. If your own sweep shows the head tolerating a heavy cut on your task, the prior loses to the measurement. - Do not extend the exemption silently. "Keep the first and last layers dense" quietly becomes "keep the first three blocks and the whole head dense", and at that point the exemptions, not the compression, decide the model's size. - Report the exemption when you report the compression ratio. A model described as 90 percent sparse with an unstated dense head is not comparable to one that is 90 percent sparse throughout. ## What an interviewer is checking They want to hear the two halves of the argument — negligible budget cost, outsized accuracy cost — rather than the folk version, "the first and last layers are special". The candidates who stand out add the caveat: the arithmetic that makes the exemption free for a convolutional stem does not automatically hold for a wide output head, and the underlying principle is redundancy, not position.
- When is exempting the final classifier actually expensive?When the head is wide: its size is feature width times number of outputs, so with thousands of classes or a large output vocabulary it can be a meaningful fraction of the model. Then the exemption spends real budget, and the honest options are a partial exemption — a higher bit-width rather than full precision — or accepting the cut and recovering it with retraining.
- In a depthwise-separable network for chest X-ray triage, which layers do the sweeps usually flag as fragile?The depthwise convolutions, not the pointwise ones. A depthwise layer applies one small kernel per channel and holds very few parameters for the work it does, so there is no redundancy to remove; the wide pointwise layers hold most of the parameters and absorb heavy pruning. It is a reminder that fragility tracks redundancy, not depth or position.
- Would you apply this rule without running your own sweep?As a starting allocation, yes — it is a well-supported prior and it saves sweep time. But I would still measure, because it is task-dependent and the rule says nothing about which middle layers are fragile. I would also report the exemption alongside the compression ratio, since a model that is sparse everywhere except a dense head is not the same model.
saying these in an interview costs you the question
- Says the first and last layers are simply more important
- Applies the exemption without checking the head's parameter share
- Assumes only the boundary layers can be fragile
- Reports a compression ratio while hiding the exempted layers
- Widens the exemption until it dominates the model's size