skip to content

Why train a fine-tune's randomly initialised head with the pretrained backbone frozen first?

level: middleimportance: should knowfreq 52%

answer

  1. the new head knows nothing yet
  2. random predictions mean large loss gradients
  3. that gradient flows down into the backbone
  4. freeze first, unfreeze when head loss flattens

basics

~20 s

A fresh head predicts almost randomly, so its loss gradient is large and, seen from the backbone, close to noise. Freezing the backbone for the first epochs lets the head reach sensible outputs before any pretrained weight moves.

solid answer

~50 s

At step zero the head's weights are random, so its predictions are arbitrary and the gradient at the logits, `dL/dz = p - y` for softmax with cross-entropy, is large. That signal is backpropagated into the backbone through random head weights, so the direction it induces there carries no information about the task; it is a big push in an arbitrary direction applied to weights whose exact values are the whole point of using a pretrained model. Head warmup removes the risk instead of shrinking it: hold the backbone fixed for the first epoch or two, train only the head at a comparatively large rate, and unfreeze once the head's loss flattens. By then the gradients reaching the backbone are driven by a head that already predicts reasonably, so they point somewhere useful. This is a different thing from a learning-rate warmup schedule: here what warms up is the head, not the step size.

go deeper

for a junior

Be ready to say what happens first in a fine-tune: the new head is random, so you train it with the pretrained layers held fixed before letting anything else move.

for a middle

Explain the mechanism. A random head makes the logit gradient large, backpropagation carries it into the backbone through random head weights, so the direction is uninformative, and freezing removes that update entirely.

for a senior

Show operating judgment: unfreeze on the head-loss curve rather than a copied step count, know that staged unfreezing buys safety with epochs, and be able to say when a short warmup plus full unfreezing is the better trade.

for a principal

Own the recipe your teams inherit. Decide whether head warmup is a mandatory step in the standard fine-tuning path, how many staged epochs the default budget funds, and what evidence would justify making the staged schedule optional.

## The problem: an untrained head attached to a trained body Transfer starts by replacing the pretrained output layer with a new one sized for the target label set, and that new layer is randomly initialised. The backbone below it is the opposite: every weight carries information earned over a long pretraining run. The first fine-tuning steps therefore mix two components in very different states, and the untrained one drives the gradient. Concretely, with softmax and cross-entropy the gradient at the logits is `dL/dz = p - y`, where `p` is the predicted distribution and `y` the one-hot label. A random head produces `p` far from `y`, so this quantity is large. Backpropagation carries it into the backbone through the head's random weight matrix, which means the direction the backbone is pushed in is determined by numbers that encode nothing. The result is a large update in an essentially arbitrary direction applied to the most valuable part of the model — the exact mechanism behind features being wrecked in the first few hundred steps. ## The fix: warm the head, not the step size Head warmup makes the risk structurally impossible for a while. Freeze every pretrained parameter, train only the new head, and let it run for a short spell — commonly one to two epochs on a modest target set. With the features fixed, the head's problem is small and well behaved: it is fitting a shallow classifier on top of static representations, so it converges quickly. Once its loss flattens, unfreeze the backbone and continue at a small backbone rate. The gradients now entering the backbone come from a head that already predicts sensibly, so they encode real task signal instead of noise. Two details matter in practice. First, the head can afford a much larger rate than the backbone will later use, precisely because nothing valuable is downstream of it. Second, the signal to unfreeze is the head's own loss curve levelling off — not a fixed step count copied from someone else's recipe. On a very small target set, watch that the head does not start overfitting during a long warmup; if it does, warm up for less time. Do not confuse this with a learning-rate warmup schedule, which ramps the step size up from near zero over the first updates. Those are different mechanisms answering overlapping concerns: one changes *which parameters* are trainable, the other changes *how big the steps are*. ## Extending warmup into gradual unfreezing Head warmup is the first move of a longer protocol. Gradual unfreezing continues it: after the head is warm, release the topmost pretrained block for one epoch, then the next block down for the following epoch, and so on until the whole network is trainable. The ordering follows what the layers hold. Top blocks encode the most task-specific abstractions and are the ones that most need to change; bottom blocks encode generic structure that usually transfers as is. Releasing top-down means the layers that must move are trained longest, and at any point in the schedule only a small part of the network can drift, so a bad batch damages less. A worked contrast: a four-class legal-document topic classifier built on a pretrained text encoder. Unfreezing everything on epoch one puts the entire encoder under gradients driven by a head that has seen the target labels once, and the run is fragile — its first-epoch validation number is often below what a frozen encoder gives. Warming the head for two epochs and then releasing one block per epoch keeps every intermediate checkpoint at or above that reference and reaches a higher final score on the same budget. ## The cost, and when to skip it Gradual unfreezing is not free. Each staged epoch is an epoch not spent training the full network, so on a fixed budget the schedule trades adaptation time for safety, and it adds a decision — how many blocks, how many epochs each — to your tuning surface. When the target domain is close to the source, the dataset is reasonably large, and the backbone rate is already small, unfreezing everything after a short head warmup usually performs just as well and is simpler to run. The staged version earns its cost when the target set is small, the domain is further away, or the run has already shown you a first-epoch collapse. What almost never pays is skipping the head warmup itself. It costs one or two cheap epochs — cheap because the frozen backbone means no gradients through most of the network — and it removes the single most common way a fine-tune destroys the thing it started from.

  • How do you decide when the head is warm enough to unfreeze the backbone?
    Watch the head's own loss curve: once it flattens over an epoch, further head-only training adds little and the gradients it sends down are as informative as they will get. On a modest target set that is usually one to two epochs. On a very small set, also watch for the head starting to overfit during warmup, which is a reason to unfreeze sooner rather than to keep going.
  • After the head is warm, why release one block per epoch instead of unfreezing everything?
    Because top blocks hold the most task-specific features and need the most change, while lower blocks hold generic structure that transfers as is. Releasing top-down trains the layers that must move for longest, and at any moment only a small part of the network can drift, so a bad batch damages less. It costs staged epochs, so it earns its keep mainly on small target sets or distant domains.
  • Does head warmup remove the need for a small backbone rate afterwards?
    No. Warmup fixes the direction problem — gradients now carry task signal instead of noise from a random head — but not the distance problem. The pretrained weights still sit near a good solution and still only need small corrections, so the backbone rate stays well below the pretraining rate once it is unfrozen. Warmup lets you use that small rate confidently, not skip it.

You do not let a brand-new driver take the wheel of the team's only car on day one. They practise in the lot first, and only then get to steer something valuable.

saying these in an interview costs you the question

  • The head trains fine either way, warmup is just tradition
  • Head warmup and learning-rate warmup are the same thing
  • Unfreeze on a fixed step count regardless of the loss curve
  • Gradual unfreezing is always worth the extra epochs
  • Warmup means you can then use the pretraining rate

context