What does L2-normalising a penultimate embedding before cosine comparison actually discard?
answer
- direction kept, length thrown away
- unit vectors make dot product cosine
- squared distance becomes two minus twice cosine
- the norm tracked confidence and typicality
- store the norm as a separate column
basics
~20 sIt discards the vector's magnitude and keeps only its direction, which is exactly what cosine similarity compares. That magnitude usually tracks how typical or confident the input was, so keep it as a separate stored scalar instead of losing it.
solid answer
~50 sCosine similarity is `dot(a,b) / (|a| * |b|)`, so scaling every vector to unit length makes a plain inner product equal to cosine, and makes squared Euclidean distance equal `2 - 2*cos` — a monotone function of cosine, so both give the same ranking. What you throw away is the norm. In a network trained with softmax cross-entropy the class score is `w_c . h`, which grows with `|h|`, and empirically the feature norm tends to be larger for typical, well-represented, confident inputs and smaller for blurry, ambiguous or out-of-distribution ones. For marketplace listing dedup that is the right trade — I want two photos of the same bicycle to match regardless of image quality — but I store the pre-normalisation norm as its own column and use it to triage junk uploads, rather than smuggling it back into the vector.
go deeper
Know the formula for cosine similarity and that dividing each vector by its length is what makes an inner product equal to it. Be able to say normalisation keeps direction and drops length.
Derive that squared Euclidean distance on unit vectors equals two minus twice the cosine, and explain why the feature norm of a softmax-trained network tends to grow with confidence.
Show the operational move: normalise for comparison, persist the raw norm as a triage and drift-monitoring column, and recognise a rectified non-negative embedding from a compressed similarity range.
Frame normalisation as a policy decision about what the similarity contract ignores, and make sure quality and typicality are handled by an explicit signal rather than leaking into ranking by accident.
## What normalisation does mechanically L2 normalisation replaces each embedding `h` by `h / |h|`, where `|h| = sqrt(sum of squares of the components)`. Every vector then sits on the unit sphere, and three identities follow. First, `dot(a,b)` on unit vectors *is* the cosine similarity, because the denominator `|a| * |b|` equals one. Any downstream machinery that computes inner products therefore behaves as a cosine comparison without knowing it. Second, squared Euclidean distance becomes `|a - b|^2 = |a|^2 + |b|^2 - 2*dot(a,b) = 2 - 2*cos`. So on normalised vectors, Euclidean distance and cosine similarity induce **the same ranking**; only the threshold scale differs. Arguments about which of the two to use are moot once you have normalised. Third, similarities become bounded in `[-1, 1]` and, for a set of vectors from the same encoder, comparable across datasets. That is why a normalised space is where you calibrate a near-duplicate threshold: an unnormalised inner product ranks partly by magnitude, so a handful of long vectors score high against nearly everything and the threshold means something different in every region of the space. ## What the norm was carrying The norm is not noise. In a classifier trained with softmax cross-entropy, the score for class `c` is `w_c . h = |w_c| * |h| * cos(angle)`. Increasing `|h|` scales all logits up together, which sharpens the softmax and lowers the loss on examples the network already gets right. The optimisation therefore has a standing incentive to grow the norm of features it classifies confidently. The empirical regularity that follows is well documented and worth being able to state: **feature norm tends to be larger for typical, frequent, cleanly-classified inputs and smaller for atypical, ambiguous, corrupted or out-of-distribution ones.** It is a tendency, not a law, and it is encoder-specific — you should verify it on your own data before relying on it — but it is strong enough that feature norm is used as a cheap out-of-distribution and confidence signal. So normalising is a deliberate act of forgetting. On the used-goods marketplace, that is usually what you want: a crisp studio photo of a bicycle and a dim handheld photo of the same bicycle should match on content, and if the dim one has a smaller norm, an unnormalised comparison would quietly rank it below unrelated but confidently-classified images. Direction-only comparison removes photo quality from the similarity. ## Keeping the signal you just discarded The right pattern is *normalise for comparison, retain the norm as metadata*. Store `|h|` as a separate scalar column next to the unit vector. It gives you, for free, a quality and typicality score you can use to route low-norm uploads to review, to suppress them from being chosen as the canonical listing in a duplicate cluster, or to monitor drift: a sustained downward shift in the norm distribution of incoming items is an early sign that production traffic has moved away from what the backbone was trained on. What you must **not** do is append the norm as an extra coordinate of the vector you compare. That reintroduces magnitude into the distance through the back door, in an uncontrolled way — the appended coordinate has whatever scale the norms happen to have, which is usually far larger than the unit-vector coordinates, so it silently dominates every comparison. ## The rectified-activation wrinkle If the penultimate vector comes after a rectified activation it is non-negative, so every vector lies in a single orthant of the space. Cosine similarity is then confined to `[0, 1]` and, in practice, bunched high — unrelated items may sit around 0.6 and duplicates around 0.9, leaving a narrow usable band. Normalisation does not fix this, because it is a property of direction, not length. Mean-centring the set first — subtract the mean embedding computed over a large sample, then normalise — spreads the similarity distribution back out and makes the threshold far easier to set. If you do centre, the mean vector becomes part of the encoder contract and must be versioned and reused identically at query time. ## When not to normalise If the vector feeds a downstream trained model rather than a similarity comparison, normalisation is a modelling choice rather than a requirement, and you may prefer to hand the model both the direction and the norm and let it decide. And if the magnitude *is* the signal — you are ranking by confidence, detecting anomalies, or scoring image quality — normalising first destroys precisely the thing you came for.
- After normalising, does it matter whether you rank by cosine or by Euclidean distance?No. For unit vectors `|a - b|^2 = 2 - 2*cos`, which is strictly decreasing in cosine, so the two produce identical orderings and identical neighbour sets. Only the numeric threshold changes, and it converts exactly between the two. Any observed difference in results means something else differs — usually that one path is not actually normalising.
- How do you keep the typicality signal after normalising?Persist the pre-normalisation norm as a separate scalar column alongside the unit vector, and use it as a triage or monitoring feature: low-norm items are atypical, low-quality or off-distribution. Do not append it as an extra dimension of the compared vector — its scale dwarfs the unit coordinates and would silently dominate every distance.
- Your cosine similarities all sit between 0.6 and 0.95. What is going on?Most likely the penultimate vector is post-rectification, so it is non-negative and every vector occupies one orthant, forcing cosines high and into a narrow band. Mean-centre the embeddings over a large sample before normalising to spread the distribution out, and version that mean vector as part of the encoder contract so query-time extraction matches.
saying these in an interview costs you the question
- Thinking normalisation changes which vectors are nearest neighbours in cosine terms
- Treating the norm as meaningless scale with no information
- Appending the norm as an extra vector dimension
- Claiming Euclidean and cosine rank differently on unit vectors
- Normalising when the magnitude is the very signal being ranked