Why does face verification use an embedding with a distance threshold instead of an N-way classifier?
answer
- open set, not a fixed roster
- new person joins on Monday
- compare two inputs, do not label one
- identities live in a store, not the weights
basics
~20 sA classifier can only recognise identities it was trained on, so every new person means retraining. An embedding model learns a distance where same-person pairs land close together, so a new identity is enrolled by storing one vector.
solid answer
~50 sVerification is an open-set problem: the people the system must handle at test time are mostly not in the training set. An N-way softmax head hard-codes a fixed label list, needs many labelled examples per class, and has to be retrained and redeployed every time someone joins or leaves. Metric learning instead trains an encoder — usually a Siamese setup where the same shared-weight network embeds both inputs — so that distance in embedding space means identity: `d(f(x1), f(x2))` is small for the same person and large for different people. Enrolment then means running one photo through the encoder and storing the vector; the decision is `d < tau` for a threshold tuned on held-out identities. The class list lives in a store, not in the weights, so adding a person costs one forward pass instead of a training run.
go deeper
Be ready to state the core contrast in one line: a classifier predicts a fixed list of identities, an embedding model measures similarity so a new identity is enrolled by storing a vector.
Explain the Siamese setup with shared weights, why sharing is what makes distances comparable, and what enrolment and verification actually do at inference time.
Show you know the operational consequences: threshold tuning on held-out identities, rejecting unenrolled people in 1:N search, and re-checking the threshold when the capture conditions change.
Own the framing decision — whether identity belongs in the weights or in a store — and the downstream cost of each: retraining cadence and redeploy risk versus an embedding store, a threshold policy and drift monitoring.
## The problem shape Two tasks look similar but are not. **Classification** assigns an input to one of a fixed, known set of classes: this photo is one of the 1,000 employees the model was trained on. **Verification** answers a yes/no question about a *pair*: are these two samples the same person? **Identification** is the 1:N version — find which enrolled person, if any, this sample matches. Verification and identification are *open-set*: the identities seen in production were mostly never seen in training. An N-way softmax classifier is a closed-set device by construction. Its final layer has one output unit per training identity, and those units are the only answers it can ever produce. Three consequences follow: 1. **New identities require surgery.** Adding a person means adding an output unit, collecting labelled examples of them, and retraining (at minimum, the head). In a system where people join weekly, that is a training pipeline in the critical path of onboarding. 2. **It needs many examples per class.** Softmax training learns a decision region per class; one or two photos per person is not enough to fit one. 3. **It cannot say 'nobody'.** Softmax outputs sum to one over the known identities, so a stranger is confidently assigned to whichever known identity they most resemble. Bolting on an 'unknown' class does not fix this, because 'not any of these people' is not a coherent class to sample from. ## What metric learning does instead Metric learning trains an encoder `f` so that a *distance* carries the semantics: `d(f(x1), f(x2))` is small when `x1` and `x2` are the same person and large otherwise. The architecture that expresses this is the **Siamese network**: two (or three) copies of the *same* network with *shared* weights, each embedding one input, with the loss defined on the resulting distances rather than on any label output. Weight sharing is not an optimisation trick — it is what makes distances meaningful. Two independently parameterised branches would map inputs into two different spaces, where comparing coordinates is nonsense, and would break the symmetry `d(x, y) = d(y, x)`. The losses that shape this space are pairwise **contrastive** loss (pull labelled same-pairs together, push different-pairs apart until they exceed a margin) and **triplet** loss (make the negative farther from the anchor than the positive by at least a margin). What matters for this question is what they produce: a space where a single global threshold separates same from different. ## What deployment looks like - **Enrolment**: run the new person's sample(s) through the encoder, L2-normalise, store the vector (or the average of a few) under their id. No gradient step, no retraining, no redeploy. - **Verification**: embed the probe, compute the distance to the claimed identity's stored vector, accept if it is below the threshold. - **Identification**: embed the probe, find the nearest stored vector, and *still* apply the threshold so that an unknown person is rejected rather than matched to the closest employee. - **Removal**: delete a row. The identity list has moved out of the weights and into data. That is the whole point, and it is why the same design shows up for speaker verification, signature verification, and any 'is this the same entity?' problem with a churning population. ## The honest trade-offs Metric learning is not free. Training is harder: the loss depends on *which* pairs or triplets you feed it, most randomly drawn ones quickly become uninformative, and a badly mined batch can drive the encoder into a degenerate solution. You also inherit a threshold that must be tuned on held-out identities and re-checked when the population, camera, or microphone changes, whereas a classifier just gives you an argmax. So the classifier is still right when the label set is genuinely fixed and small, you have plenty of examples per class, and nobody will ever be added: predicting which of five machine parts is in an image needs no distance function. In practice a classification-trained network's features are also a strong initialisation for the encoder — but the metric objective is what actually shapes the distance you are going to threshold, and a classifier's raw features do not come with a calibrated one. ## Interview framing The crisp answer is one sentence about the label space: a classifier stores the identities in its weights; a metric model stores them in a database, and only the notion of similarity lives in the weights. Everything else — enrolment cost, rejecting strangers, one-shot use — falls out of that.
- What is the difference between verification and identification with the same embedding model?Verification is 1:1 — embed the probe, compare it to the one claimed identity's stored vector, accept or reject against a threshold. Identification is 1:N — compare the probe against every enrolled vector and take the nearest. Identification must still apply a threshold, otherwise an unenrolled stranger is always matched to whoever happens to be closest.
- Why do the two branches of a Siamese network share weights?Because both inputs must be mapped by the same function into one space; otherwise the two embeddings live in different coordinate systems and their distance means nothing. Sharing also halves the parameters and guarantees the symmetry d(x, y) = d(y, x), which any sane distance-based decision rule assumes.
- When would you still prefer a plain classifier over a metric-learned embedding?When the label set is fixed, small and fully known in advance, with plenty of examples per class and no enrolment requirement. Softmax training is simpler, converges faster, needs no pair or triplet sampling, and gives a probability you can calibrate directly. Reach for metric learning when the class list churns or classes have one or two examples.
A classifier is a guest list of known names; an embedding model is the desk clerk who compares any face to whatever photo is on file, including one filed this morning.
saying these in an interview costs you the question
- Says you just retrain the classifier whenever a new person joins
- Claims a Siamese network's branches have separate weights
- Believes a softmax head can reject an unknown person as-is
- Assumes any classifier's features already give a usable distance threshold