A frozen backbone's features drift between epochs even though no weight gets a gradient — why?
answer
- freezing is two switches, not one
- gradients versus training mode
- some state updates on the forward pass
- running mean and variance are not parameters
- same batch twice should give identical features
basics
~20 sNormalization layers hold running mean and variance estimates that are state, not learned parameters. Blocking gradients does not stop them: in training mode the forward pass keeps re-estimating them from the new domain's batches, so the outputs move.
solid answer
~50 sFreezing is two separate switches and people flip only one. The first is whether a parameter receives a gradient and gets updated. The second is whether the module runs in training or inference mode. A batch-statistic normalization layer in training mode normalizes with the current batch's statistics and folds them into its stored running estimates on every forward pass, none of which involves a gradient. So a 'frozen' backbone left in training mode gives different features for the same image across epochs, and its stored statistics migrate toward the target domain. The head is then fitting a moving representation, and previously cached embeddings no longer match. The fix is to put the backbone in inference mode as well as excluding it from updates, then verify by pushing one batch through twice and checking the outputs are identical.
go deeper
Know that a network behaves differently in training mode than in inference mode, and that switching modes is a separate action from deciding which weights get updated. Being able to name the two switches is enough here.
Explain that a batch-statistic normalization layer keeps running mean and variance as state maintained by the forward pass, not as learned parameters, so excluding it from gradient updates leaves those estimates free to move.
Describe the symptoms you would actually see — a noisy loss curve on a supposedly fixed representation, a jump when evaluation mode kicks in, cached embeddings that stop matching — and give the two-forward-passes check plus the fix.
Own the convention so nobody rediscovers this. Make feature extraction pin both switches in one place, and add a test that asserts a fixed backbone returns identical outputs for the same batch twice, so the bug fails the build instead of quietly costing accuracy.
## Two switches, not one "Freezing a backbone" is shorthand for two independent things: 1. **No update.** The parameter is excluded from whatever the optimiser updates, or simply receives no gradient. This is the switch people think about. 2. **Inference mode.** Certain layers behave differently during training than during inference. This switch is orthogonal, and forgetting it is the classic transfer-learning bug. ## Parameters versus running state A batch-statistic normalization layer such as BatchNorm holds two kinds of quantity. Its **scale and shift** are learned parameters, updated by gradients like any weight. Its **running mean and running variance** are not learned at all: they are state that the *forward pass* maintains, an exponential moving average of the per-batch statistics the layer has seen, blended in with a momentum coefficient. Backpropagation never touches them. That is exactly why freezing by gradient has no effect on them. Push a batch of radiographs through a backbone pretrained on natural photos, and the layer records radiograph statistics — inside a single epoch the stored estimates have visibly migrated toward the new domain, with every learned weight in the network untouched. On top of that, while it is in training mode the layer does not even *use* its stored estimates for the forward computation: it normalizes each activation using the current mini-batch's statistics. So the feature you get for an image depends on which other images happened to share its batch. ## What it looks like in a real run - **Non-reproducibility.** The same image yields different feature vectors on epoch 1 and epoch 5. Your head is chasing a representation that is quietly moving underneath it, which shows up as a loss curve that is noisier than a frozen-feature run has any right to be. - **A train/evaluation discontinuity.** During training the layer uses batch statistics; at evaluation it switches to the stored running estimates. If those estimates have drifted, training-time and evaluation-time behaviour diverge, and validation accuracy can jump or collapse in a way that has nothing to do with learning. - **Incomparable cached embeddings.** The whole appeal of a fixed backbone is precomputing embeddings once — for a nearest-neighbour index, a retrieval store, or a linear probe. If some extractions ran in training mode and others in inference mode, the vectors live in slightly different spaces and every distance you compute is wrong. Nothing crashes; the numbers are just worse than they should be. - **Dropout is the same trap.** Dropout is also mode-dependent. A backbone with dropout left in training mode randomly zeroes activations, so the head is fitting noisy features and, again, no gradient is involved. ## The one-minute check Run the identical batch through the backbone twice, with no update step in between, and compare the outputs. Identical means the backbone really is fixed. Differing means either the normalization statistics are live, or dropout is active, or both. A second check: read one layer's stored running mean, do a forward pass, read it again — if it changed, the module is in training mode. ## The fix Put the backbone into inference mode *and* exclude its parameters from updates, and re-apply the mode every time the training loop re-enters a training phase — a common bug is code that flips the whole model back to training mode at the top of each epoch, silently undoing the setting for the frozen part. If you also want the learned scale and shift held fixed, make sure they are genuinely excluded from the optimised parameter set, not merely small. ## When the drift is not a bug Letting normalization statistics re-estimate on target-domain data is sometimes a deliberate adaptation choice, and it can help when input statistics really did shift. The point is that it must be a decision you made and measured, not a side effect of a mode flag. Keep the two switches distinct in your head, and say which one you are setting and why. ## Why this matters for the freeze decision The reason to freeze is a *fixed*, cheap, reusable representation: stable features, no stored activations, cacheable embeddings. If the representation is not actually fixed, you have paid the accuracy cost of freezing without collecting any of its benefits. The same applies to the frozen portion of a partially frozen network — the lower stages you left alone can still be drifting. A useful architectural note: layer normalization and group normalization compute their statistics from the current activations of each individual sample, at both training and inference time. They keep no running estimates, so a backbone built on them simply does not have this failure mode. Only batch-statistic normalization does.
- How would you prove in one minute that this is what is happening?Push the identical batch through the backbone twice with no update step in between and compare the outputs. Identical outputs mean the backbone is genuinely fixed; differing outputs mean normalization statistics or dropout are still live. As a second check, read a normalization layer's stored running mean before and after a single forward pass and see whether it moved.
- Which normalization choices avoid this failure entirely?Layer normalization and group normalization compute statistics from each sample's own activations, identically at training and inference time, so they hold no running estimates that can drift. Only batch-statistic normalization keeps an exponential moving average of per-batch statistics, which is the state that migrates onto a new domain. If you intend to freeze a backbone, that difference is worth knowing up front.
- Besides normalization, what else changes when a frozen module is left in training mode?Dropout stays active, so the frozen backbone emits randomly zeroed activations and the head fits noise it will never see at inference; stochastic depth behaves the same way. Neither involves a gradient, so neither is affected by excluding the backbone from updates. Both are governed by the same training-versus-inference mode switch.
Bolting a machine's dials in place does not stop its gauges from re-zeroing themselves every time you switch it on. The settings are frozen; the readings are not.
saying these in an interview costs you the question
- Freezing the weights means the module cannot change at all
- Running statistics are learned by backpropagation
- Turning off gradients also puts the layer in inference mode
- The same input always gives the same features during training
- Cached embeddings are safe to mix regardless of extraction mode