skip to content

Learning Transferable Features

Where a reusable representation comes from: label-supervised pretraining, contrastive and masked pretext objectives, and metric losses. The objective decides what actually transfers.

on this pageshow

explore

questions

19

Why does face verification use an embedding with a distance threshold instead of an N-way classifier?

level: juniorimportance: must knowfreq 66%

answer

  1. open set, not a fixed roster
  2. new person joins on Monday
  3. compare two inputs, do not label one
  4. identities live in a store, not the weights

basics

~20 s

A classifier can only recognise identities it was trained on, so every new person means retraining. An embedding model learns a distance where same-person pairs land close together, so a new identity is enrolled by storing one vector.

solid answer

~50 s

Verification is an open-set problem: the people the system must handle at test time are mostly not in the training set. An N-way softmax head hard-codes a fixed label list, needs many labelled examples per class, and has to be retrained and redeployed every time someone joins or leaves. Metric learning instead trains an encoder — usually a Siamese setup where the same shared-weight network embeds both inputs — so that distance in embedding space means identity: `d(f(x1), f(x2))` is small for the same person and large for different people. Enrolment then means running one photo through the encoder and storing the vector; the decision is `d < tau` for a threshold tuned on held-out identities. The class list lives in a store, not in the weights, so adding a person costs one forward pass instead of a training run.

go deeper

for a junior

Be ready to state the core contrast in one line: a classifier predicts a fixed list of identities, an embedding model measures similarity so a new identity is enrolled by storing a vector.

for a middle

Explain the Siamese setup with shared weights, why sharing is what makes distances comparable, and what enrolment and verification actually do at inference time.

for a senior

Show you know the operational consequences: threshold tuning on held-out identities, rejecting unenrolled people in 1:N search, and re-checking the threshold when the capture conditions change.

for a principal

Own the framing decision — whether identity belongs in the weights or in a store — and the downstream cost of each: retraining cadence and redeploy risk versus an embedding store, a threshold policy and drift monitoring.

## The problem shape Two tasks look similar but are not. **Classification** assigns an input to one of a fixed, known set of classes: this photo is one of the 1,000 employees the model was trained on. **Verification** answers a yes/no question about a *pair*: are these two samples the same person? **Identification** is the 1:N version — find which enrolled person, if any, this sample matches. Verification and identification are *open-set*: the identities seen in production were mostly never seen in training. An N-way softmax classifier is a closed-set device by construction. Its final layer has one output unit per training identity, and those units are the only answers it can ever produce. Three consequences follow: 1. **New identities require surgery.** Adding a person means adding an output unit, collecting labelled examples of them, and retraining (at minimum, the head). In a system where people join weekly, that is a training pipeline in the critical path of onboarding. 2. **It needs many examples per class.** Softmax training learns a decision region per class; one or two photos per person is not enough to fit one. 3. **It cannot say 'nobody'.** Softmax outputs sum to one over the known identities, so a stranger is confidently assigned to whichever known identity they most resemble. Bolting on an 'unknown' class does not fix this, because 'not any of these people' is not a coherent class to sample from. ## What metric learning does instead Metric learning trains an encoder `f` so that a *distance* carries the semantics: `d(f(x1), f(x2))` is small when `x1` and `x2` are the same person and large otherwise. The architecture that expresses this is the **Siamese network**: two (or three) copies of the *same* network with *shared* weights, each embedding one input, with the loss defined on the resulting distances rather than on any label output. Weight sharing is not an optimisation trick — it is what makes distances meaningful. Two independently parameterised branches would map inputs into two different spaces, where comparing coordinates is nonsense, and would break the symmetry `d(x, y) = d(y, x)`. The losses that shape this space are pairwise **contrastive** loss (pull labelled same-pairs together, push different-pairs apart until they exceed a margin) and **triplet** loss (make the negative farther from the anchor than the positive by at least a margin). What matters for this question is what they produce: a space where a single global threshold separates same from different. ## What deployment looks like - **Enrolment**: run the new person's sample(s) through the encoder, L2-normalise, store the vector (or the average of a few) under their id. No gradient step, no retraining, no redeploy. - **Verification**: embed the probe, compute the distance to the claimed identity's stored vector, accept if it is below the threshold. - **Identification**: embed the probe, find the nearest stored vector, and *still* apply the threshold so that an unknown person is rejected rather than matched to the closest employee. - **Removal**: delete a row. The identity list has moved out of the weights and into data. That is the whole point, and it is why the same design shows up for speaker verification, signature verification, and any 'is this the same entity?' problem with a churning population. ## The honest trade-offs Metric learning is not free. Training is harder: the loss depends on *which* pairs or triplets you feed it, most randomly drawn ones quickly become uninformative, and a badly mined batch can drive the encoder into a degenerate solution. You also inherit a threshold that must be tuned on held-out identities and re-checked when the population, camera, or microphone changes, whereas a classifier just gives you an argmax. So the classifier is still right when the label set is genuinely fixed and small, you have plenty of examples per class, and nobody will ever be added: predicting which of five machine parts is in an image needs no distance function. In practice a classification-trained network's features are also a strong initialisation for the encoder — but the metric objective is what actually shapes the distance you are going to threshold, and a classifier's raw features do not come with a calibrated one. ## Interview framing The crisp answer is one sentence about the label space: a classifier stores the identities in its weights; a metric model stores them in a database, and only the notion of similarity lives in the weights. Everything else — enrolment cost, rejecting strangers, one-shot use — falls out of that.

  • What is the difference between verification and identification with the same embedding model?
    Verification is 1:1 — embed the probe, compare it to the one claimed identity's stored vector, accept or reject against a threshold. Identification is 1:N — compare the probe against every enrolled vector and take the nearest. Identification must still apply a threshold, otherwise an unenrolled stranger is always matched to whoever happens to be closest.
  • Why do the two branches of a Siamese network share weights?
    Because both inputs must be mapped by the same function into one space; otherwise the two embeddings live in different coordinate systems and their distance means nothing. Sharing also halves the parameters and guarantees the symmetry d(x, y) = d(y, x), which any sane distance-based decision rule assumes.
  • When would you still prefer a plain classifier over a metric-learned embedding?
    When the label set is fixed, small and fully known in advance, with plenty of examples per class and no enrolment requirement. Softmax training is simpler, converges faster, needs no pair or triplet sampling, and gives a probability you can calibrate directly. Reach for metric learning when the class list churns or classes have one or two examples.

A classifier is a guest list of known names; an embedding model is the desk clerk who compares any face to whatever photo is on file, including one filed this morning.

saying these in an interview costs you the question

  • Says you just retrain the classifier whenever a new person joins
  • Claims a Siamese network's branches have separate weights
  • Believes a softmax head can reject an unknown person as-is
  • Assumes any classifier's features already give a usable distance threshold

context

open as a page

When retargeting a 1000-class pretrained backbone to 12 classes, what do you do with its head?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Delete the source classifier and attach a freshly initialised 12-output layer on top of the same features. The source label space has no meaning for the new task, so its output weights go, while everything below is kept.

open as a page

What does the InfoNCE contrastive loss optimise, and what does its temperature control?

level: middleimportance: must knowfreq 70%

basics

~20 s

InfoNCE is a softmax cross-entropy over similarities: an anchor must pick its own positive view out of a pool of negatives. The temperature divides those similarities before the softmax, so a small temperature concentrates the gradient on the hardest negatives.

open as a page

Why does masked-image pretraining mask around 75% of patches when masked text masks only 15%?

level: middleimportance: must knowfreq 62%

basics

~20 s

Images are highly redundant, so a lightly masked patch can be interpolated from its neighbours and the pretext teaches nothing. Text tokens are far denser in information, so hiding even a small fraction already forces real inference about meaning and structure.

open as a page

What does the margin in a triplet loss enforce, and when is a triplet's loss exactly zero?

level: middleimportance: must knowfreq 58%

basics

~20 s

The margin demands that the negative sit farther from the anchor than the positive by at least that gap, not merely farther. Any triplet already satisfying the gap has loss exactly zero and contributes no gradient at all.

open as a page

Why does skip-gram training use negative sampling instead of a full softmax?

level: middleimportance: must knowfreq 54%

basics

~20 s

A full softmax normalises over the whole vocabulary, so every update touches every output vector. Negative sampling instead scores the real pair up and a few sampled noise pairs down, making cost independent of vocabulary size.

open as a page

How do skip-gram and CBOW differ, and which wins on a small, rare-word-heavy corpus?

level: middleimportance: must knowfreq 62%

basics

~20 s

Skip-gram predicts each context word from the centre word; CBOW predicts the centre word from the averaged context. Skip-gram creates more separate updates per rare word, so it wins on small rare-word-heavy corpora, while CBOW trains faster.

open as a page

Why do features from a supervised ImageNet backbone transfer to a task with different classes?

level: middleimportance: must knowfreq 76%

basics

~10 s

Early layers learn generic edges, colours and textures that nearly any image task needs, and only the deepest layers specialise to the source categories. Transfer keeps the generic stack and replaces the specialised end.

open as a page

After masked-reconstruction pretraining, why do you throw away the reconstruction decoder?

level: juniorimportance: should knowfreq 51%

basics

~20 s

The decoder exists only to turn the hidden input into a training signal. What you want afterwards is the encoder's representation, so the decoder is discarded and a fresh head for the real task is attached in its place.

open as a page

Why does a static word vector give 'bank' one vector for both of its meanings?

level: juniorimportance: should knowfreq 58%

basics

~10 s

The model stores one row per word type, keyed by spelling. River-bank and money-bank occurrences both pull on that same row, so the result is a single compromise vector sitting between two unrelated neighbourhoods.

open as a page

How do BYOL and SimSiam avoid representational collapse without using any negative pairs?

level: middleimportance: should knowfreq 46%

basics

~20 s

They break the symmetry that makes a constant output optimal: one branch carries an extra predictor network, the other is a stop-gradient target. Remove the stop-gradient and training reaches the trivial solution where every input maps to the same vector.

open as a page

In contrastive self-supervision, how does the augmentation set decide what the encoder ignores?

level: seniorimportance: should knowfreq 52%

basics

~20 s

A positive pair is one input under two augmentations, so the encoder learns to discard whatever those augmentations change. Colour jitter buys hue invariance on street photos and ruins a dermatology encoder, where hue is the diagnostic signal.

open as a page

In a rotation or jigsaw pretext, how do you detect that a shortcut solved the task?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The tell is a pretext that is solved suspiciously well while the pretrained encoder transfers no better than random initialisation. Confirm it by ablation: destroy the suspected low-level cue in the input and see whether pretext accuracy collapses.

open as a page

Your triplet model mines only the hardest negatives and the loss stalls exactly at the margin — why?

level: seniorimportance: should knowfreq 44%

basics

~20 s

The encoder has collapsed: it maps every input to nearly the same point, so all distances are zero and every triplet's loss equals the margin. Hardest-negative mining causes it because the hardest negatives are mostly mislabels and near-duplicates.

open as a page

How do you decide how many pretrained blocks to reuse versus retrain for a new target task?

level: seniorimportance: should knowfreq 54%

basics

~10 s

Sweep the split point: transfer the first k blocks, retrain the rest, and read target performance against k. Two effects fight — depth makes features source-specific, and an arbitrary split breaks co-adapted layers.

open as a page

Your pretrained word vectors have no entry for misspellings or rare surnames — what fixes it?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

Switch to vectors that represent a word as the sum of its character n-gram vectors. An unseen surname or typo still shares n-grams with trained words, so a vector is composed for it rather than a shared placeholder.

open as a page

Would you supply contrastive negatives with SimCLR's large batch or MoCo's momentum queue on a small cluster?

level: principalimportance: nice to knowfreq 36%

basics

~20 s

Both need many negatives but buy them differently. SimCLR's negatives are the rest of the batch, so a 4096-sample batch and its memory are the price. MoCo decouples negatives from batch size with a momentum-encoder queue, so modest hardware suffices.

open as a page

How would you set the decision threshold for speaker verification when each user enrols from three utterances?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Sweep the distance on trials from speakers held out of training, then pick the point where the false-accept rate matches what the security policy allows and the false-reject rate stays tolerable. The threshold is a policy choice.

open as a page

Your target is a forklift near-miss video classifier — which supervised pretraining corpus do you pick?

level: principalimportance: nice to knowfreq 34%

basics

~20 s

Pick the source whose labels force the same discriminations your target needs. For near-miss events that means a large human-action video corpus: separating hundreds of action classes demands motion features a still-image corpus never builds.

open as a page