skip to content

When retargeting a 1000-class pretrained backbone to 12 classes, what do you do with its head?

level: juniorimportance: must knowfreq 70%

answer

  1. the last layer is indexed by class
  2. 1000 rows, none of them yours
  3. keep the features, drop the mapping
  4. a random layer on a trained stack
  5. reinitialise, then warm it up

basics

~20 s

Delete the source classifier and attach a freshly initialised 12-output layer on top of the same features. The source label space has no meaning for the new task, so its output weights go, while everything below is kept.

solid answer

~50 s

The final layer is a linear map from the penultimate feature vector to 1000 logits, with one weight vector per source category. None of those 1000 categories is one of your 12, so those weight vectors are meaningless for the target and the whole layer is discarded; you attach a new linear layer with 12 outputs, randomly initialised, and keep the feature extractor beneath it. The practical consequence is that at step zero you have a random head sitting on a well-trained backbone, so the first gradients are large and noisy and can wreck the pretrained features — which is why people warm the head up first, or use a smaller learning rate for the backbone than for the head. One nuance worth stating: discarding the label space is not the same as discarding what those labels taught. The penultimate features still encode every distinction the 1000 categories forced the network to learn.

go deeper

for a junior

Be ready to describe the surgery concretely: cut off the source classifier, attach a new randomly initialised layer with as many outputs as your task has classes, keep everything underneath.

for a middle

Explain why the source rows are meaningless rather than merely unused, and why a random head next to a trained backbone calls for head warm-up or a reduced backbone learning rate.

for a senior

Show the operational habits: verify the input preprocessing matches the source model, stage the unfreezing, and be able to say what you would inspect if the retargeted model underperforms early.

for a principal

Frame the choice between reusing the penultimate representation and consuming source outputs as a black box, and argue which one leaves the organisation less coupled to a source taxonomy it does not control.

## What the head actually is At the top of a supervised classifier sits a linear layer: it takes the penultimate feature vector — call it a d-dimensional description of the input — and produces one score per source class. Its weight matrix has one row per source category, so with 1000 source categories the matrix is 1000 by d, plus a 1000-entry bias. A softmax turns those scores into probabilities and cross-entropy trains them. That matrix is indexed by the *source* label space. Row 417 means whatever source category 417 was. There is no sense in which row 417 corresponds to one of your 12 retail-shelf classes, and there is no reordering or subsetting that makes it correspond. That is why the entire layer is removed rather than trimmed. ## The surgery 1. Cut the network at the penultimate representation — the output of the last pooling or the last hidden block, whichever your architecture exposes as the feature vector. 2. Attach a new linear layer with 12 outputs and d inputs, initialised randomly (a small-variance initialisation appropriate to the layer's fan-in) with zero or small bias. 3. Keep everything below unchanged, and decide separately how much of it you will freeze or retrain. Nothing constrains the new head to be smaller than the old one. A target with 1200 classes is just as valid as one with 12; the head's width is a free choice and what limits you is whether the d-dimensional feature vector carries enough information to separate the target classes, not the arithmetic of 12 versus 1000. ## Why the random head is a hazard At initialisation the new head predicts roughly a uniform distribution over 12 classes, so the loss is near log 12 and the gradients flowing out of it are large. Those gradients propagate backwards into a backbone whose weights are, by assumption, already good. A few large steps early in training can undo a substantial part of what pretraining bought — the effect is sometimes described as the random head washing out the pretrained features. The standard mitigations follow directly: - **Warm up the head.** Train only the new layer for a short phase with the backbone held fixed, so the head becomes sensible before any gradient is allowed to reshape the features. - **Use a smaller learning rate below.** Give the backbone a fraction of the head's learning rate, so early noise perturbs it only slightly. - **Warm up the learning rate itself.** Ramping the step size from near zero over the first fraction of training limits the damage from the initial large loss. ## What is discarded and what is not A candidate who says 'we throw away everything the source labels taught us' has the accounting wrong. Two different things are being discarded: - **The source label space**: the identity and semantics of the 1000 categories, the class prior over them, and the linear boundaries between them. Gone, and irrecoverable from the truncated network. - **What learning those labels produced**: the representation. Fully kept. The reason the penultimate vector is useful is precisely that separating 1000 diverse categories forced the network to encode textures, parts and configurations at a level of detail no small target set could have taught it. ## The alternative worth mentioning Instead of cutting the head off, you can keep the full source model as a fixed black box and feed its 1000 output scores into a target model as an input vector. This is a legitimate technique — the source logits form a compact, semantically rich descriptor, and on a very small target set they can be competitive. But it is usually the weaker option: a 1000-dimensional vector of source-category decisions is a lossy summary that has already collapsed everything the source model knew into source-taxonomy terms, whereas the penultimate feature vector is typically higher-dimensional and not yet committed to any label space. It also permanently ties the target model to the source taxonomy. Knowing both options and why the default is head replacement is the level of nuance an interviewer is listening for. ## Related surgery you should expect to be asked about The input side needs the same care: the target images must be preprocessed the way the source model was trained to expect — same resizing convention, same channel ordering, same normalisation statistics — or the pretrained filters see a distribution they were never trained on. Head replacement gets all the attention, but a silently mismatched input pipeline produces the same symptom, a backbone that appears not to transfer.

  • Could you keep the 1000 source scores as input features instead of throwing the head away?
    Yes, and it is a real technique: freeze the whole source model and feed its 1000 outputs into a small target model. It sometimes wins on tiny target sets because the descriptor is compact and semantically rich. Usually it loses to using the penultimate features, because the source head has already collapsed everything into decisions about source categories, and it permanently ties your target model to that taxonomy.
  • Why can a randomly initialised head damage a well-pretrained backbone?
    At step zero the head is uninformative, so the loss is high and the gradients it emits are large. Those gradients flow into backbone weights that were already near a good solution, and a handful of oversized steps can destroy features that took a million labelled images to learn. Warming up the head with the backbone frozen, or giving the backbone a much smaller learning rate, keeps the early phase from doing that damage.
  • Does anything break if the target has more classes than the source?
    No. The head's width is independent of the source's; a 1200-class target on a 1000-class backbone is ordinary. What actually limits you is whether the penultimate feature vector is expressive enough to separate that many target classes, and whether you have enough target labels per class to fit the new layer. If not, the fix is retraining more of the backbone or enlarging the feature stage, not resizing the source head.

saying these in an interview costs you the question

  • Wants to keep the source head and remap 12 of its 1000 outputs
  • Thinks the target must have fewer classes than the source
  • Trains a random head at full learning rate into the backbone
  • Says removing the head discards what the source labels taught
  • Forgets to match the source model's input preprocessing

context