With 800 labelled chest radiographs and a natural-image backbone, do you freeze or fine-tune?
answer
- two axes, not one
- labelled examples versus domain distance
- trainable parameters must fit the data
- early generic, late domain-specific
- small and far: unfreeze the top only
basics
~20 sFine-tune part of it. 800 examples cannot safely update a whole backbone, but radiographs sit too far from natural photos for frozen top-layer features to work, so unfreeze the last stage and train a new head.
solid answer
~40 sTwo axes decide it: how much labelled target data you have, and how far the target domain is from the pretraining domain. 800 examples is small, so the number of trainable parameters has to stay small or the model just memorises. But radiographs are far from natural photos, so the late, domain-specific layers encode the wrong things and a fully frozen backbone underfits. That puts this case in the hardest cell: replace the head, keep the generic early layers frozen, and unfreeze only the top stage, with strong augmentation and early stopping. Contrast that with 200k in-domain marketplace product photos, where you fine-tune everything because the data supports it. In practice I still run the frozen-feature baseline first — it is cheap and it sets the bar every heavier option must beat.
go deeper
Be ready to name the two options — train only a new head, or update the whole network — and to say that fewer labelled examples means unfreezing fewer layers. Knowing that the head is always trained is the baseline expectation.
Explain both axes and place a scenario in the right cell out loud. An interviewer expects you to say why early layers transfer further than late ones, and why trainable parameter count has to track the amount of labelled data.
Show a process: run the cheap frozen baseline first, read the train/validation gap to decide how much to unfreeze, cut at a stage boundary, and treat a small validation split as noisy enough to need cross-validation.
Own the default across a portfolio of small tasks. One shared frozen backbone with cached embeddings and many tiny heads is cheap to serve and retrain; a fine-tuned model per task buys accuracy and multiplies serving, storage and retraining cost. Say which you would standardise on and what would make you switch.
## The decision, framed properly A pretrained network splits into a **backbone** — everything up to the pooled feature vector — and a **head**, the small layer mapping that vector to your task's outputs. The head is always new and always trained, because a new task means a new label set. The real question is what happens to the backbone: leave every weight fixed (*feature extraction*), update every weight (*full fine-tuning*), or fix the lower part and update the top (*partial freezing*). Two variables decide it, and an interviewer wants to hear both: 1. **How many labelled target examples you have.** Every weight you unfreeze is a parameter fitted from your data. A mid-sized image backbone holds tens of millions of weights; 800 labelled radiographs cannot constrain that many degrees of freedom, so the network fits the training set and generalises badly. The working rule is that trainable parameter count should scale with labelled example count. 2. **How far the target domain is from the pretraining domain.** Early layers learn generic structure — oriented edges, corners, textures, intensity transitions — which is present in essentially any image, radiographs included. Later layers compose those into domain-specific parts: animal faces, wheels, text. On a far domain those late features describe things that are not in your images, so freezing them asks the classifier to separate your classes using coordinates that do not describe them. ## The four cells - **Large data, near domain** — 200,000 in-domain marketplace product photos against a natural-image backbone. Fine-tune everything. The data supports it, the pretrained weights act as a good initialisation rather than a constraint, and freezing leaves accuracy on the table. - **Small data, near domain** — a few thousand photos of the same kind of object. Freeze the backbone and train the head, or unfreeze only the last block. Fast, hard to overfit, and the frozen features are already close to right. - **Large data, far domain** — hundreds of thousands of radiographs. Fine-tune everything; with that much in-domain data the pretrained weights are mostly a faster starting point. - **Small data, far domain** — the 800-radiograph case, and the one interviewers pick. Keep the generic early layers, replace the head, unfreeze only the top stage so the domain-specific part can be relearned inside a parameter budget the data can support. Lean on augmentation, weight decay and early stopping, and remember that a validation split carved out of 800 examples is itself noisy — cross-validation buys you a more trustworthy comparison. ## Where to cut Cut at a structural boundary — the end of a stage or a residual block — never in the middle of one, so a block's convolutions and its normalization move together. In a five-stage convolutional backbone, *freeze stages 1-3, tune stages 4-5* is the common split, and it lands roughly where features stop being generic. You can locate the boundary empirically: pool the activations at the output of each stage, train a cheap linear classifier on each pooled representation, and see where target accuracy stops improving with depth. If accuracy peaks at stage 3 and then falls, the layers above encode the source domain — those are exactly the ones worth unfreezing. ## Run the cheap baseline first Feature extraction is the fastest experiment available: with the backbone fixed you push each image through once, store the resulting vector, and train heads on the stored vectors in seconds. That gives a number within an hour, tells you whether the pretrained representation is useful for your task at all, and sets the bar every heavier option must clear. ## Costs that follow the choice Freezing does more than skip weight updates. A frozen sub-network stores no intermediate activations for backpropagation and computes no gradients through them, so memory drops and epochs get faster, and cached embeddings can be reused across many head experiments. Full fine-tuning gives all of that up and produces a full model copy per task, which matters when you serve many tasks from one platform. ## Reading the result Compare options on the same held-out split. A wide gap between training and validation accuracy after full fine-tuning is the signature of too many trainable parameters for the data — unfreeze less. A frozen baseline that underfits *both* training and validation is the opposite signal: the representation is wrong for your domain, so unfreeze more. Two adjacent questions have their own answers elsewhere — how small the step size must become once you unfreeze, and what the model forgets about its source task. One practical warning: if you decide to freeze, make sure the backbone is actually fixed. Normalization layers left in training mode keep updating their running statistics even when no weight receives a gradient, and then your 'frozen' features are quietly moving.
- If you partially freeze, where exactly do you cut the network?At a stage or residual-block boundary, so a block's convolutions and its normalization stay together — typically freeze the first three stages and tune the last two. The cut should land where features stop being generic. You can find that point empirically by pooling each stage's output, training a cheap linear classifier on each, and seeing at which depth target accuracy stops improving.
- Your frozen-feature baseline beats full fine-tuning on validation — what does that tell you?Almost always that you unfroze more parameters than the labelled data can support. Check the training-versus-validation gap: if training accuracy is near perfect and validation is not, it is overfitting, and the fix is to unfreeze fewer layers or regularise harder. It is not evidence the pretrained weights are bad — that is a different diagnosis.
- You now have 200k in-domain product photos instead — what changes?Move to full fine-tuning. With that much in-domain labelled data the backbone's weights are an initialisation, not a constraint you must protect, and freezing would leave accuracy unclaimed. The frozen baseline is still worth running once as a reference point, but you should expect full fine-tuning to beat it clearly here.
It is like adapting to a new dialect of a language you already speak. With years of exposure you can shift your whole grammar; with a handful of overheard sentences you should only swap vocabulary, or you end up speaking something nobody understands.
saying these in an interview costs you the question
- Always fine-tune everything; more trainable weights is always better
- Freezing is purely a compute optimisation, not a generalisation choice
- Domain distance is irrelevant if the pretrained model is large enough
- 800 examples is plenty to update tens of millions of weights
- Picks freeze or fine-tune without asking about data size or domain