Transfer Learning and Embeddings
You will learn why pretrained representations transfer, how to choose between freezing layers and full fine-tuning, and what catastrophic forgetting costs you. Interviewers probe this because nearly all industrial DL starts from a pretrained model rather than training from scratch.
on this pageshowhide
explore
- Learning Transferable Features19 questions
- Supervised Pretraining4 questions
- Contrastive Self-Supervision4 questions
- Masked Reconstruction Pretexts3 questions
- Metric Learning4 questions
- Word Vector Pretraining4 questions
- Adapting a Pretrained Model17 questions
- Freeze or Fine-Tune3 questions
- Fine-Tuning Step Sizes3 questions
- Catastrophic Forgetting4 questions
- Negative Transfer3 questions
- Domain Adaptation4 questions
- Representation Quality8 questions
- Penultimate-Layer Embeddings4 questions
- Linear Probing4 questions
questions
page 2 of 2Your triplet model mines only the hardest negatives and the loss stalls exactly at the margin — why?
basics
~20 sThe encoder has collapsed: it maps every input to nearly the same point, so all distances are zero and every triplet's loss equals the margin. Hardest-negative mining causes it because the hardest negatives are mostly mislabels and near-duplicates.
Negative transfer can come from a domain gap or a source-task mismatch — how do you tell which?
basics
~20 sA domain gap means the target inputs look statistically unlike the source data. A source-task mismatch means the source objective learned invariance to the very signal the target needs. Ask what the source loss rewarded, and what it deliberately ignored.
How do you decide how many pretrained blocks to reuse versus retrain for a new target task?
basics
~10 sSweep the split point: transfer the first k blocks, retrain the rest, and read target performance against k. Two effects fight — depth makes features source-specific, and an arbitrary split breaks co-adapted layers.
How do you choose 64, 512 or 2048 as the embedding width for 80M stored listings?
basics
~20 sWidth sets a capacity ceiling, not quality, while storage and per-comparison cost grow linearly with it. Train at the widest width you can afford, plot the downstream metric against width, and ship the knee of that curve rather than the widest option.
How does elastic weight consolidation protect a network against catastrophic forgetting?
basics
~20 sElastic weight consolidation adds a quadratic penalty pulling every parameter toward its old-task value, scaled by that parameter's Fisher information. Weights the old task depended on become stiff, while unimportant ones stay free to fit the new task.
How does a gradient-reversal layer make a network's features domain-invariant?
basics
~20 sA domain classifier trained to tell source from target sits on the shared features. The gradient-reversal layer passes activations through unchanged but negates the gradient going back, so the feature extractor learns to defeat it while the task head keeps features useful.
In layer-wise discriminative fine-tuning, why does the bottom block get a much smaller rate than the head?
basics
~20 sDepth decides how much a layer must change. Bottom blocks hold generic features that transfer almost unchanged and should barely move, while upper blocks are task-specific and the head is random, so rates rise geometrically from bottom to top.
Your 80M stored 2048-dimension embeddings show only about 30 large singular values — what happened?
basics
~20 sThe embeddings occupy a roughly 30-dimensional subspace of a 2048-dimensional space — dimensional collapse. The training objective and data diversity, not the layer width, decided that, so the remaining coordinates cost storage and carry almost no information.
Your pretrained word vectors have no entry for misspellings or rare surnames — what fixes it?
basics
~20 sSwitch to vectors that represent a word as the sum of its character n-gram vectors. An unseen surname or typo still shares n-grams with trained words, so a vector is composed for it rather than a shared placeholder.
Would you supply contrastive negatives with SimCLR's large batch or MoCo's momentum queue on a small cluster?
basics
~20 sBoth need many negatives but buy them differently. SimCLR's negatives are the rest of the batch, so a 4096-sample batch and its memory are the price. MoCo decouples negatives from batch size with a momentum-encoder queue, so modest hardware suffices.
How do you make linear-probe comparisons between self-supervised encoders trustworthy?
basics
~20 sFix everything except the encoder: one layer and pooling convention, one probe family with an identical tuning budget, identical preprocessing and splits, and a reported label-budget curve. Add a parameter-free nearest-neighbour check as a cross-test, and freeze the protocol before results arrive.
How would you set the decision threshold for speaker verification when each user enrols from three utterances?
basics
~20 sSweep the distance on trials from speakers held out of training, then pick the point where the false-accept rate matches what the security policy allows and the false-reject rate stays tolerable. The threshold is a policy choice.
When does a smaller model trained on more in-domain data beat fine-tuning a large pretrained backbone?
basics
~20 sWhen target labels are plentiful and the domain is far from the source. Pretraining supplies a prior worth most when data is scarce; with enough in-domain examples a small model learns better-suited features and costs less to serve.
Your target is a forklift near-miss video classifier — which supervised pretraining corpus do you pick?
basics
~20 sPick the source whose labels force the same discriminations your target needs. For near-miss events that means a large human-action video corpus: separating hundreds of action classes demands motion features a still-image corpus never builds.
showing 31–44 of 44