How do you decide how many pretrained blocks to reuse versus retrain for a new target task?
answer
- the split point is a hyperparameter
- sweep k, plot target performance
- two effects, not one
- layers are jointly optimised neighbours
- fine-tuning repairs the broken interaction
basics
~10 sSweep the split point: transfer the first k blocks, retrain the rest, and read target performance against k. Two effects fight — depth makes features source-specific, and an arbitrary split breaks co-adapted layers.
solid answer
~50 sTreat the split point as a hyperparameter and measure it. Take the pretrained network, keep blocks 1 to k, reinitialise everything above, retrain on the target set and plot target validation performance against k. The curve is shaped by two distinct effects. First, specificity: as k grows, more of what you transferred is tuned to the source categories rather than to your task. Second, fragile co-adaptation: layers in a trained network are jointly optimised, and splitting them at an arbitrary boundary while freezing the lower half destroys interactions that neither half can rebuild alone — which is why the frozen-transfer curve tends to dip in the middle of the stack. Fine-tuning the transferred blocks rather than freezing them largely repairs the co-adaptation loss, so it usually dominates when target data allows it. Priors before you run the sweep: small target sets and near source domains favour transferring more and freezing more; large target sets and distant domains favour retraining deeper.
go deeper
Know that you do not have to reuse the whole network: you can keep the lower layers and retrain the upper ones, and how many you keep is a choice, not a fixed recipe.
Explain the sweep and the two mechanisms behind its shape — specificity increasing with depth, and the interaction between adjacent layers being broken by an arbitrary split.
Demonstrate you have operated this: staged unfreezing, per-depth learning rates, genuinely frozen normalisation behaviour, and a fair comparison protocol across split points on a small target set.
Own the cost side of the tradeoff — frozen prefixes let one precomputed feature store serve many downstream tasks — and decide when a team should standardise a split policy rather than sweep per project.
## The decision, stated precisely A pretrained network is a stack of blocks. You are choosing two things, and it helps to separate them: - **How much to transfer**: which prefix of the stack, blocks 1 to k, is initialised from the source model rather than randomly. - **How much to freeze**: of that prefix, which blocks are held fixed during target training and which are allowed to keep learning. Agents and candidates often collapse these into one knob. They are different: you can transfer the entire backbone and still fine-tune all of it, or transfer everything and freeze everything but the head. ## The sweep The honest way to answer is empirical. For k spanning the depth of the network, build a model that keeps the first k blocks from the source model, reinitialises the rest plus a target-sized head, trains on the target set, and reports target validation performance. Run each k in two variants — transferred blocks frozen, and transferred blocks fine-tuned — and plot both curves against k. It is a modest sweep on a small target set and it replaces an argument with a graph. ## What the curves show, and why Two mechanisms act at once, and reading the sweep means separating them. **Specificity rises with depth.** The first blocks encode edges, colours and textures that any image task needs; the last blocks encode configurations that separate the source's own categories. So the deeper you transfer, the more of what you brought along is about a task you do not care about. This effect penalises large k, and it penalises it harder the further the target domain sits from the source domain. **Co-adaptation is fragile.** Within a trained network, adjacent layers are jointly optimised: a layer's units are tuned to the specific, sometimes idiosyncratic, features the layer below produces. If you freeze blocks 1 to k and randomly reinitialise everything above, the newly trained upper half has to relearn how to talk to a fixed lower half, and gradient descent from a random start often cannot recover the interaction that existed before. The signature of this is a dip in the frozen-transfer curve at *middle* values of k — layers deep enough to be entangled with what is above them, but not yet so specific that specificity explains the drop. Transferring the whole co-adapted group, or fine-tuning the transferred blocks so they can re-adapt, largely removes the dip. Because the two effects have different shapes — one monotonic in depth, one worst in the middle — the frozen curve can be non-monotonic while the fine-tuned curve is much flatter and usually higher. ## Priors to bring before you measure - **Target dataset size.** With few target labels, retraining deep layers overfits: you cannot re-estimate millions of parameters from a small set. Transfer more, freeze more, fit little. With a large target set, freezing wastes capacity and fine-tuning deeper wins. - **Domain distance.** Retail-shelf photographs sit close to a natural-image source corpus, so deep features stay useful. Spectrograms, overhead imagery or medical scans diverge earlier in the stack, so you keep the low-level filters and retrain more of the top. - **Label-space relatedness.** A source whose categories demand the same discriminations as your target — the same kinds of texture, part or motion distinctions — extends the useful depth. - **Compute and serving cost.** Freezing a prefix means you can precompute those activations once for the whole target set and train only the top, which is dramatically cheaper and worth real accuracy in some settings. ## Operational details people get wrong - **Normalisation layers are not fully frozen by freezing weights.** Layers that maintain running input statistics keep updating those statistics from target batches even when their learnable parameters are fixed. If you intend a block to be genuinely fixed, you must put it in inference mode as well, otherwise its behaviour drifts as target data flows through it, and your 'frozen' baseline is not the baseline you think it is. - **Staged unfreezing.** A common recipe is to train the head with everything frozen, then progressively unfreeze from the top down at a reduced learning rate. This is a cheap approximation of the sweep and reduces the risk of early gradients damaging the transferred features. - **Per-depth learning rates.** Rather than a binary frozen/trainable split, you can assign smaller learning rates to lower blocks and larger ones nearer the head, which makes the split continuous instead of a hard cut. - **Compare fairly.** Each k must get its own tuned learning rate and its own early-stopping decision; a sweep where every configuration shares one learning rate mostly measures which k happened to suit that learning rate. ## How to answer it in the room Say that it is a hyperparameter, describe the sweep in one sentence, then name the two competing effects — depth-increasing specificity and split-induced co-adaptation loss — and state that fine-tuning rather than freezing repairs the second. Finish with the two priors that decide the starting guess: how much target data you have and how far your domain is from the source's. That sequence demonstrates you have actually run this rather than repeated a recipe.
- Your frozen-transfer curve dips in the middle of the stack and recovers higher up — what explains that?That shape is the co-adaptation effect rather than specificity. Splitting the network at a mid-stack boundary and randomly reinitialising above forces the new upper half to relearn how to consume a fixed lower half, and it often cannot recover the joint solution. Transfer the whole entangled group, or let the transferred blocks fine-tune, and the dip largely disappears. A monotone decline that only appears near the top is the specificity effect instead.
- Your target set has a few hundred labelled images. How does that change the split?It pushes you hard toward transferring the whole backbone and freezing most of it. A few hundred examples cannot re-estimate deep layers without overfitting, so you fit a small head, or the head plus the last block, on largely fixed features. Freezing also lets you precompute activations once and iterate cheaply. If validation says the domain is too far for frozen deep features, unfreeze from the top with a small learning rate rather than opening the whole stack.
- How do you make sure the sweep itself is a fair comparison across values of k?Give each configuration its own learning-rate search and its own early stopping, because the best step size for a mostly frozen model differs from the best for a mostly retrained one. Hold the target split, augmentation and epoch budget fixed across configurations, run each on more than one seed if the target set is small, and put normalisation layers in inference mode wherever you claim a block is frozen, so the frozen arm really is frozen.
saying these in an interview costs you the question
- Treats freeze-everything-but-the-head as the only option
- Assumes deeper always means better transfer
- Ignores that frozen normalisation layers still update running statistics
- Confuses how much to transfer with how much to freeze
- Compares split points under one shared learning rate