skip to content

Transfer Learning and Embeddings

You will learn why pretrained representations transfer, how to choose between freezing layers and full fine-tuning, and what catastrophic forgetting costs you. Interviewers probe this because nearly all industrial DL starts from a pretrained model rather than training from scratch.

on this pageshow

explore

questions

page 1 of 2

What is catastrophic forgetting when a pretrained network is fine-tuned on a new task?

level: juniorimportance: must knowfreq 58%

answer

  1. learning B costs you A
  2. the loss has no term for old data
  3. shared weights, myopic optimizer
  4. new head, new gradients, drifted features
  5. old task metrics nobody is watching

basics

~20 s

Catastrophic forgetting is the sharp drop in a network's performance on its original task after it is trained on a new one. Fine-tuning minimises only the new task's loss, so the shared weights drift away from the old solution.

solid answer

~50 s

Catastrophic forgetting is what happens when you keep training a network on task B and its accuracy on task A collapses, often within a few hundred steps. The cause is not mysterious: gradient descent is minimising a loss computed only on task B's data, and nothing in that objective references task A. Every parameter is shared, so the optimizer is free to move weights out of the region where task A's loss was low as long as that lowers task B's loss. A warehouse pick-and-place policy fine-tuned for a new suction gripper can end up unable to run the parallel-jaw task it shipped with, even though nobody changed the parallel-jaw code. It is worst when the two tasks share a backbone but differ in output space or input statistics, and when you train for many epochs at a step size large enough to move the pretrained features rather than just the head.

go deeper

for a junior

Be able to state the definition in one sentence and name the cause: training minimises only the new task's loss, and the weights are shared. Know that the old task's metrics have to be measured or the regression is invisible.

for a middle

Explain the optimization view — no term in the objective references the old data, so nothing resists moving out of the old low-loss region. Distinguish representation drift from a re-aimed output head, and say how you would tell them apart.

for a senior

Show that you expect it and instrument for it. Talk about which factors set the severity — task distance, how many layers move, how long you train, how many adaptation stages have stacked up — and about the retained evaluation that catches it before release.

for a principal

Own the position that any repeatedly adapted model needs a forgetting budget as an explicit release criterion, not a debugging afterthought. Be ready to argue when accepting the loss on a retired capability is the right call versus paying to retain it.

## The phenomenon A network is trained on task A until it performs well. Training then continues on task B — new labels, new data, or simply a newer slice of the same stream. Performance on B rises as expected. Performance on A, which nobody is measuring because A's data is no longer in the loop, falls hard. That collapse is **catastrophic forgetting**: the loss of previously acquired capability caused by learning something new in the same set of weights. The word *catastrophic* is doing real work. Forgetting in a neural network is not the slow decay a person experiences; it can take a network from 0.91 recall to below 0.5 on the old task inside a single epoch of the new one. ## Why it happens Think about what the optimizer is actually asked to do. During fine-tuning the objective is the average loss over the **new** task's batches. There is no term in that objective that mentions the old data, so there is no force at all resisting movement away from the old solution. The old solution was one point in a large region of weight space where task A's loss is low; the new gradient points wherever task B's loss falls fastest, and that direction is essentially unrelated to A's low-loss region. Three structural facts make it worse: - **The parameters are shared.** A single backbone encodes features for both tasks. There is no partition of weights reserved for A, so any update that helps B may overwrite a feature A depended on. - **The objective is myopic.** Stochastic gradient descent optimises the current batch's loss. It has no memory of the loss surface it came from, and no notion that some directions are cheap for B but expensive for A. - **The output layer usually changes.** Adapting to a new label set typically means a fresh, randomly initialised head. Early in training that head produces large, badly aimed gradients that propagate into the backbone, and if the old head was discarded the old task's predictions are gone by construction, not merely degraded. ## Two distinguishable failures It is worth separating **representation drift** from **head drift**, because they are diagnosed and repaired differently. - *Head drift*: the backbone still encodes everything the old task needed, but the classifier on top has been re-aimed at the new labels. In a class-incremental setup, where each stage introduces new classes and only those classes appear as positives, the final layer picks up a strong recency bias — logits for recently seen classes systematically outrank older ones. - *Representation drift*: the features themselves have moved, so the information the old task needed is no longer linearly available. A cheap way to tell them apart is to freeze the current backbone and fit a fresh linear classifier on old-task data. If that probe recovers most of the old performance, the features survived and the head was the casualty. If it does not, the representation itself has been overwritten. ## Where the severity comes from Forgetting is not a constant; it scales with how far training is allowed to move the weights and how far apart the tasks are. - **Distance between tasks.** Two tasks that need similar features interfere less. A new suction-gripper skill that reuses the same visual features as the old parallel-jaw skill damages less than one that needs an entirely different notion of graspability. - **How much of the network moves.** The more layers are unfrozen and the larger the steps, the more the pretrained features can be displaced. Lower layers, which encode generic structure, typically drift less than the task-specific upper layers, partly because their gradients are smaller and partly because generic features remain useful for the new task. - **How long training runs.** Forgetting compounds with steps taken on the new distribution, so a long fine-tune on a small new dataset is a common way to destroy a good backbone. - **Sequence length.** Under repeated sequential adaptation — five stages, then ten — the earliest stage is usually the worst hit, because it has had the most subsequent updates to survive. ## What it is not Two confusions are worth naming. First, forgetting is **not** overfitting. An overfit model has memorised the training set of the task it is being trained on; a model that has forgotten may generalise perfectly well on the new task while having lost a different one. They can occur together, but neither implies the other. Second, forgetting is **not** literal erasure. The weights still carry a great deal of old-task structure, which is why old performance is often partially recoverable from a small amount of old data far faster than it was originally learned. That partial recoverability is exactly what the standard mitigations exploit: keeping a slice of old data in the training stream, or adding an explicit penalty that makes the parameters the old task relied on expensive to move. ## The practical consequence Because the new-task metrics look healthy the whole time, forgetting is invisible unless you deliberately keep measuring the old task. Any pipeline that adapts a model repeatedly needs a retained evaluation set from the original task, scored on the same cadence as the new-task validation set, or the regression ships.

  • Is catastrophic forgetting the same thing as overfitting the new task?
    No. Overfitting means the model memorises its current training set and generalises poorly on that same task. Forgetting means performance on a different, earlier task collapses, and it can happen while the new task generalises well. They have different causes and different fixes, though a long, aggressive fine-tune on a small dataset tends to produce both at once.
  • Which parts of a network tend to forget most, and why?
    The task-specific upper layers and the output head. Lower layers encode generic structure that remains useful for the new task, so their gradients are smaller and they drift less. The head is often replaced outright for a new label set, and a freshly initialised head also sends large early gradients back into the backbone, which is what damages the features.
  • If the old task's data is gone, is the old capability permanently lost?
    Usually not entirely. The weights still carry much of the old structure, so a modest amount of old data restores old performance far faster than the original training did. That is evidence the capability was buried rather than erased. Without any old data at all you are limited to methods that constrained the parameters before the new training began.

Overwriting a whiteboard for a new meeting: nothing forbids using the space where yesterday's diagram was, because the only thing being optimised is today's diagram.

saying these in an interview costs you the question

  • Calls it overfitting to the new task
  • Thinks the network literally erases stored memories
  • Believes a low new-task loss proves nothing was lost
  • Assumes only the output head can be affected
  • Says it only happens with a bad learning rate

context

open as a page

In unsupervised domain adaptation, what is covariate shift and why does it hurt a trained network?

level: juniorimportance: must knowfreq 66%

basics

~20 s

Covariate shift means the input distribution moves from source to target while the rule mapping input to label is unchanged. The network was fitted where source data lived, so target inputs land where its decision boundary was never pinned down.

open as a page

Why is the fine-tuning learning rate for a pretrained network far smaller than its pretraining rate?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Pretraining already puts the weights in a good region, so fine-tuning only needs small corrections. A large step moves every weight far enough to destroy the learned features, and a small target dataset cannot rebuild them.

open as a page

How do you replace a pretrained model's 1000-class head for a 7-class task?

level: juniorimportance: must knowfreq 66%

basics

~20 s

Discard the old output layer and attach a new, randomly initialised one with seven outputs, keeping the layers beneath it. The old head maps to the wrong label set, so none of its weights are reusable.

open as a page

Why does face verification use an embedding with a distance threshold instead of an N-way classifier?

level: juniorimportance: must knowfreq 66%

basics

~20 s

A classifier can only recognise identities it was trained on, so every new person means retraining. An embedding model learns a distance where same-person pairs land close together, so a new identity is enrolled by storing one vector.

open as a page

When retargeting a 1000-class pretrained backbone to 12 classes, what do you do with its head?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Delete the source classifier and attach a freshly initialised 12-output layer on top of the same features. The source label space has no meaning for the new task, so its output weights go, while everything below is kept.

open as a page

What does the InfoNCE contrastive loss optimise, and what does its temperature control?

level: middleimportance: must knowfreq 70%

basics

~20 s

InfoNCE is a softmax cross-entropy over similarities: an anchor must pick its own positive view out of a pool of negatives. The temperature divides those similarities before the softmax, so a small temperature concentrates the gradient on the hardest negatives.

open as a page

With 800 labelled chest radiographs and a natural-image backbone, do you freeze or fine-tune?

level: middleimportance: must knowfreq 72%

basics

~20 s

Fine-tune part of it. 800 examples cannot safely update a whole backbone, but radiographs sit too far from natural photos for frozen top-layer features to work, so unfreeze the last stage and train a new head.

open as a page

What does a linear probe on a frozen encoder's features actually measure?

level: middleimportance: must knowfreq 62%

basics

~20 s

A linear probe trains only a linear classifier on features a frozen network already produces. Its accuracy measures how linearly separable the target classes are in that fixed feature space; nothing inside the encoder is changed or learned.

open as a page

Why does masked-image pretraining mask around 75% of patches when masked text masks only 15%?

level: middleimportance: must knowfreq 62%

basics

~20 s

Images are highly redundant, so a lightly masked patch can be interpolated from its neighbours and the pretext teaches nothing. Text tokens are far denser in information, so hiding even a small fraction already forces real inference about meaning and structure.

open as a page

What does the margin in a triplet loss enforce, and when is a triplet's loss exactly zero?

level: middleimportance: must knowfreq 58%

basics

~20 s

The margin demands that the negative sit farther from the anchor than the positive by at least that gap, not merely farther. Any triplet already satisfying the gap has loss exactly zero and contributes no gradient at all.

open as a page

How would you prove a pretrained backbone helped rather than caused negative transfer?

level: middleimportance: must knowfreq 52%

basics

~20 s

Negative transfer means a pretrained start leaves you worse off than starting cold. The only proof is a control run: the same architecture trained from scratch on the same target data, under a matched budget and equal tuning effort.

open as a page

Why reuse a trained image classifier's penultimate activations as an embedding rather than its logits?

level: middleimportance: must knowfreq 72%

basics

~20 s

The penultimate vector is the general-purpose feature that the final linear layer scores. Logits are that same vector already projected onto the training classes, so they discard every distinction the label set happened to ignore.

open as a page

Why does skip-gram training use negative sampling instead of a full softmax?

level: middleimportance: must knowfreq 54%

basics

~20 s

A full softmax normalises over the whole vocabulary, so every update touches every output vector. Negative sampling instead scores the real pair up and a few sampled noise pairs down, making cost independent of vocabulary size.

open as a page

How do skip-gram and CBOW differ, and which wins on a small, rare-word-heavy corpus?

level: middleimportance: must knowfreq 62%

basics

~20 s

Skip-gram predicts each context word from the centre word; CBOW predicts the centre word from the averaged context. Skip-gram creates more separate updates per rare word, so it wins on small rare-word-heavy corpora, while CBOW trains faster.

open as a page

Why do features from a supervised ImageNet backbone transfer to a task with different classes?

level: middleimportance: must knowfreq 76%

basics

~10 s

Early layers learn generic edges, colours and textures that nearly any image task needs, and only the deepest layers specialise to the source categories. Transfer keeps the generic stack and replaces the specialised end.

open as a page

After masked-reconstruction pretraining, why do you throw away the reconstruction decoder?

level: juniorimportance: should knowfreq 51%

basics

~20 s

The decoder exists only to turn the hidden input into a training signal. What you want afterwards is the encoder's representation, so the decoder is discarded and a fresh head for the real task is attached in its place.

open as a page

Why does a static word vector give 'bank' one vector for both of its meanings?

level: juniorimportance: should knowfreq 58%

basics

~10 s

The model stores one row per word type, keyed by spelling. River-bank and money-bank occurrences both pull on that same row, so the result is a single compromise vector sitting between two unrelated neighbourhoods.

open as a page

How does replaying a small buffer of source examples reduce catastrophic forgetting?

level: middleimportance: should knowfreq 44%

basics

~20 s

Interleaving a small sample of the original task's data into every batch puts an old-task term back into each gradient. The optimizer can no longer lower the new loss at unlimited cost to the old one.

open as a page

How do BYOL and SimSiam avoid representational collapse without using any negative pairs?

level: middleimportance: should knowfreq 46%

basics

~20 s

They break the symmetry that makes a constant output optimal: one branch carries an extra predictor network, the other is a stop-gradient target. Remove the stop-gradient and training reaches the trivial solution where every input maps to the same vector.

open as a page

In domain adaptation, why does re-estimating normalization statistics on target data help?

level: middleimportance: should knowfreq 45%

basics

~20 s

A network that normalizes with dataset-level means and variances carries constants measured on source data. Recomputing them from unlabelled target batches re-centres and re-scales every layer's activations into the range later layers expect, with no labels and no gradient steps.

open as a page

Why train a fine-tune's randomly initialised head with the pretrained backbone frozen first?

level: middleimportance: should knowfreq 52%

basics

~20 s

A fresh head predicts almost randomly, so its loss gradient is large and, seen from the backbone, close to noise. Freezing the backbone for the first epochs lets the head reach sensible outputs before any pretrained weight moves.

open as a page

A linear probe on frozen features scores near chance while a two-layer head succeeds - what does that prove?

level: middleimportance: should knowfreq 38%

basics

~20 s

It proves the property is present in the frozen features but not exposed along any single hyperplane. A low linear-probe score bounds linear decodability, not the information itself, so it is never evidence that a feature is absent.

open as a page

What does L2-normalising a penultimate embedding before cosine comparison actually discard?

level: middleimportance: should knowfreq 55%

basics

~20 s

It discards the vector's magnitude and keeps only its direction, which is exactly what cosine similarity compares. That magnitude usually tracks how typical or confident the input was, so keep it as a separate stored scalar instead of losing it.

open as a page

How do you detect catastrophic forgetting while adapting a model to a new task?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Keep a frozen evaluation set from the original task and score it on the same cadence as the new task's validation set. Forgetting shows up as a falling retained-source score, so make it a release gate.

open as a page

In contrastive self-supervision, how does the augmentation set decide what the encoder ignores?

level: seniorimportance: should knowfreq 52%

basics

~20 s

A positive pair is one input under two augmentations, so the encoder learns to discard whatever those augmentations change. Colour jitter buys hue invariance on street photos and ruins a dermatology encoder, where hue is the diagnostic signal.

open as a page

Under domain shift, why is a network's confidence a poor filter for target pseudo-labels?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Softmax confidence is not a calibrated probability, and shift makes it worse: a network stays confident where it is wrong. A fixed high threshold then keeps the most source-like, easiest examples, skewing the pseudo-label set toward classes the model already handles.

open as a page

A frozen backbone's features drift between epochs even though no weight gets a gradient — why?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Normalization layers hold running mean and variance estimates that are state, not learned parameters. Blocking gradients does not stop them: in training mode the forward pass keeps re-estimating them from the new domain's batches, so the outputs move.

open as a page

A linear probe hits 68% where full fine-tuning hits 79% - what does that gap tell you?

level: seniorimportance: should knowfreq 45%

basics

~20 s

An eleven-point gap says the target task needs structure the frozen features do not expose linearly - the encoder either discarded it during pretraining or holds it in a form only updated weights surface. The gap sizes the mismatch, not its cause.

open as a page

In a rotation or jigsaw pretext, how do you detect that a shortcut solved the task?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The tell is a pretext that is solved suspiciously well while the pretrained encoder transfers no better than random initialisation. Confirm it by ablation: destroy the suspected low-level cue in the input and see whether pretext accuracy collapses.

open as a page

showing 1–30 of 44