skip to content

Why do cosine, dot product and L2 rank identically once vectors are unit-normalized?

level: middleimportance: must knowfreq 62%

answer

  1. the norms both become 1
  2. one identity links all three functions
  3. squared distance written from the cosine
  4. monotone means same order, different numbers

basics

~20 s

Unit-normalized vectors have length 1, so the dot product equals the cosine, and squared Euclidean distance equals 2 minus twice that cosine. All three are monotone functions of one quantity, so the ordering is identical even though the numbers differ.

solid answer

~50 s

Normalizing makes every vector length 1. The dot product a·b equals ‖a‖‖b‖cos, so with both norms at 1 it *is* the cosine. Squared Euclidean distance expands to ‖a‖² + ‖b‖² - 2a·b, which collapses to 2 - 2cos. Since that is strictly decreasing in cosine, sorting ascending by L2 gives exactly the reverse-free equivalent of sorting descending by cosine — same order, different numbers. Practically this means normalizing at write time removes the metric decision: you can then pick whichever function your index computes fastest, usually the dot product, since it skips the division. Two caveats matter. The *values* are not interchangeable, so a threshold calibrated on cosine cannot be pasted into a distance comparison. And normalization is lossy: it discards the magnitude, which in some models carries a real prior you may have wanted.

code

python · 13 lines
python
import numpy as np

a = np.array([3.0, 4.0])
b = np.array([1.0, 0.0])

an = a / np.linalg.norm(a)
bn = b / np.linalg.norm(b)

cos = float(an @ bn)
print(cos)                              # 0.6  -> cosine
print(float(an @ bn))                   # 0.6  -> dot product, identical
print(float(np.sum((an - bn) ** 2)))    # 0.8  -> squared L2
print(2 - 2 * cos)                      # 0.8  -> same value

go deeper

for a junior

Know that normalizing means rescaling every vector to length 1, and that after doing so cosine and dot product give the same number while Euclidean distance orders results the same way in reverse.

for a middle

Derive it on the whiteboard: a·b = ‖a‖‖b‖cos with unit norms, and ‖a-b‖² = 2 - 2cos. Then say plainly that monotone means same ranking, not same values.

for a senior

Show the operational consequence — normalize on write, compare with the cheapest function, convert every calibrated threshold through the identity, and renormalize after any averaging or combination step.

for a principal

Frame normalization as a modelling commitment rather than hygiene: decide per embedding space whether magnitude carries a learned prior worth keeping, and make that decision explicit and documented rather than a default in a helper function.

## The algebra, in full Start from the definition of the dot product in terms of the angle: a·b = ‖a‖ · ‖b‖ · cos(θ) Unit-normalizing means replacing each vector v with v/‖v‖, so that ‖a‖ = ‖b‖ = 1. Substituting gives a·b = cos(θ). The dot product and the cosine are then literally the same number, not merely correlated. Now expand squared Euclidean distance: ‖a - b‖² = (a - b)·(a - b) = ‖a‖² + ‖b‖² - 2(a·b) With unit norms this becomes 2 - 2cos(θ), so the distance itself is √(2 - 2cos θ). Both are strictly decreasing functions of the cosine: as similarity rises from -1 to 1, squared distance falls from 4 to 0. A strictly monotone transform never reorders anything, so the top-k under any of the three functions contains the same items in the same relative order — provided you sort each in its correct direction. ## What this buys you operationally Because the three coincide, normalizing on write turns the metric choice into an implementation detail. Systems typically normalize once at indexing time and once per query, then compare with the dot product, since it is the cheapest of the three: no square roots, no per-comparison division, and it is the operation hardware and libraries optimize hardest. The chosen metric still has to be declared consistently at index build and query time — the equivalence tells you the *ranking* is the same, not that a mismatched configuration is harmless. ## What normalization throws away Normalization is a projection onto the unit sphere and it is not reversible: the magnitude is gone. Whether that matters depends entirely on the model. For most text-embedding models trained with a cosine objective, the magnitude was never a supervised quantity and discarding it is free. For models trained with an inner-product objective, magnitude can encode a learned prior — a per-item bias, a confidence, a popularity term — and normalizing silently deletes a signal the training put there on purpose. Deciding to normalize is therefore a modelling decision, not a hygiene step. ## The asymmetric-normalization case A nuance worth having ready: suppose documents are normalized but queries are not. For the *dot product*, the ranking is unaffected. Scoring is q·d over a fixed q, and scaling q by a constant scales every candidate score by the same constant, which cannot reorder them. For *Euclidean distance*, the ranking absolutely does change — ‖q - d‖² = ‖q‖² + 1 - 2(q·d), and although ‖q‖² is constant, the cross term is now weighted by the query's length relative to a fixed 1, so the geometry shifts. This is why "we normalize documents, that's enough" is correct under one metric and wrong under another. The reverse asymmetry — normalized queries against unnormalized documents — breaks dot-product ranking too, because each candidate carries its own different length into the score. Only the query side is free. ## Scores are not portable, only order is The most common practical trap here is treating the equivalence as if it made the values interchangeable. It does not. A minimum-similarity cut of 0.8 on cosine corresponds to a squared distance of 2 - 2(0.8) = 0.4, that is, a distance of about 0.632. Pasting 0.8 into a distance comparison keeps items out to a far larger radius, which quietly loosens the filter to near-uselessness. When you migrate a system from one metric to the other, convert every constant through the identity rather than reusing it. ## Operations that break unit norm Normalization is a property of the vectors as stored, and ordinary post-processing destroys it. Averaging several chunk vectors into one document vector produces something with length below 1 unless the chunks were identical. Summing vectors, applying a weighted combination, adding a small perturbation, or dimensionality-truncating a vector all leave the unit sphere. Any pipeline that manipulates vectors after embedding needs an explicit renormalization step at the end, otherwise the system drifts back into the unnormalized regime where the three metrics diverge again — and it does so gradually, without an error.

  • If documents are normalized but query vectors are not, does dot-product ranking change?
    No. The query's length is a single constant multiplying every candidate score, and a positive constant factor cannot reorder a list. Euclidean ranking, however, does change, because the query norm enters the cross term against fixed unit-length documents. So "we normalize the documents" is a sufficient answer under dot product and an insufficient one under L2.
  • What do you lose by normalizing, and when does that loss matter?
    You lose the magnitude permanently. That is free for models trained with a cosine objective, where length was never supervised. It is costly for models trained with an inner-product objective, where length can encode a learned prior such as item quality or confidence. If you still need that signal after normalizing, you have to reintroduce it as an explicit, tunable term rather than hoping it survived.
  • You average several chunk embeddings into one document vector. Anything to watch?
    The mean of unit vectors is not itself a unit vector — it is shorter, and how much shorter depends on how much the chunks disagree. Left unnormalized, coherent documents get systematically longer vectors than diverse ones, which then biases any dot-product ranking. Renormalize after any averaging, summing, weighting or truncation step.
  • Can you convert a cosine threshold of 0.8 into an equivalent Euclidean cut-off?
    Yes, through the identity: squared distance is 2 - 2 x 0.8 = 0.4, so the distance cut-off is about 0.632. The ordering is preserved between the metrics but the numbers are not, so every calibrated constant has to be converted rather than copied when you switch metrics.

saying these in an interview costs you the question

  • Says normalized cosine and L2 produce the same scores, not just the same ranking
  • Reuses a cosine threshold directly as a distance threshold
  • Claims cosine and dot product are the same function regardless of normalization
  • Treats normalization as lossless housekeeping rather than a modelling decision
  • Forgets to renormalize after averaging or combining vectors

context