skip to content

Transfer Learning and Embeddings

You will learn why pretrained representations transfer, how to choose between freezing layers and full fine-tuning, and what catastrophic forgetting costs you. Interviewers probe this because nearly all industrial DL starts from a pretrained model rather than training from scratch.

on this pageshow

explore

questions

page 2 of 2

Your triplet model mines only the hardest negatives and the loss stalls exactly at the margin — why?

level: seniorimportance: should knowfreq 44%

basics

~20 s

The encoder has collapsed: it maps every input to nearly the same point, so all distances are zero and every triplet's loss equals the margin. Hardest-negative mining causes it because the hardest negatives are mostly mislabels and near-duplicates.

open as a page

Negative transfer can come from a domain gap or a source-task mismatch — how do you tell which?

level: seniorimportance: should knowfreq 38%

basics

~20 s

A domain gap means the target inputs look statistically unlike the source data. A source-task mismatch means the source objective learned invariance to the very signal the target needs. Ask what the source loss rewarded, and what it deliberately ignored.

open as a page

How do you decide how many pretrained blocks to reuse versus retrain for a new target task?

level: seniorimportance: should knowfreq 54%

basics

~10 s

Sweep the split point: transfer the first k blocks, retrain the rest, and read target performance against k. Two effects fight — depth makes features source-specific, and an arbitrary split breaks co-adapted layers.

open as a page

How do you choose 64, 512 or 2048 as the embedding width for 80M stored listings?

level: principalimportance: should knowfreq 44%

basics

~20 s

Width sets a capacity ceiling, not quality, while storage and per-comparison cost grow linearly with it. Train at the widest width you can afford, plot the downstream metric against width, and ship the knee of that curve rather than the widest option.

open as a page

How does elastic weight consolidation protect a network against catastrophic forgetting?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Elastic weight consolidation adds a quadratic penalty pulling every parameter toward its old-task value, scaled by that parameter's Fisher information. Weights the old task depended on become stiff, while unimportant ones stay free to fit the new task.

open as a page

How does a gradient-reversal layer make a network's features domain-invariant?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

A domain classifier trained to tell source from target sits on the shared features. The gradient-reversal layer passes activations through unchanged but negates the gradient going back, so the feature extractor learns to defeat it while the task head keeps features useful.

open as a page

In layer-wise discriminative fine-tuning, why does the bottom block get a much smaller rate than the head?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Depth decides how much a layer must change. Bottom blocks hold generic features that transfer almost unchanged and should barely move, while upper blocks are task-specific and the head is random, so rates rise geometrically from bottom to top.

open as a page

Your 80M stored 2048-dimension embeddings show only about 30 large singular values — what happened?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

The embeddings occupy a roughly 30-dimensional subspace of a 2048-dimensional space — dimensional collapse. The training objective and data diversity, not the layer width, decided that, so the remaining coordinates cost storage and carry almost no information.

open as a page

Your pretrained word vectors have no entry for misspellings or rare surnames — what fixes it?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

Switch to vectors that represent a word as the sum of its character n-gram vectors. An unseen surname or typo still shares n-grams with trained words, so a vector is composed for it rather than a shared placeholder.

open as a page

Would you supply contrastive negatives with SimCLR's large batch or MoCo's momentum queue on a small cluster?

level: principalimportance: nice to knowfreq 36%

basics

~20 s

Both need many negatives but buy them differently. SimCLR's negatives are the rest of the batch, so a 4096-sample batch and its memory are the price. MoCo decouples negatives from batch size with a momentum-encoder queue, so modest hardware suffices.

open as a page

How do you make linear-probe comparisons between self-supervised encoders trustworthy?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Fix everything except the encoder: one layer and pooling convention, one probe family with an identical tuning budget, identical preprocessing and splits, and a reported label-budget curve. Add a parameter-free nearest-neighbour check as a cross-test, and freeze the protocol before results arrive.

open as a page

How would you set the decision threshold for speaker verification when each user enrols from three utterances?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Sweep the distance on trials from speakers held out of training, then pick the point where the false-accept rate matches what the security policy allows and the false-reject rate stays tolerable. The threshold is a policy choice.

open as a page

When does a smaller model trained on more in-domain data beat fine-tuning a large pretrained backbone?

level: principalimportance: nice to knowfreq 29%

basics

~20 s

When target labels are plentiful and the domain is far from the source. Pretraining supplies a prior worth most when data is scarce; with enough in-domain examples a small model learns better-suited features and costs less to serve.

open as a page

Your target is a forklift near-miss video classifier — which supervised pretraining corpus do you pick?

level: principalimportance: nice to knowfreq 34%

basics

~20 s

Pick the source whose labels force the same discriminations your target needs. For near-miss events that means a large human-action video corpus: separating hundreds of action classes demands motion features a still-image corpus never builds.

open as a page

showing 31–44 of 44