skip to content

Similarity Metrics

Comparing two vectors: cosine similarity, dot product, and Euclidean distance, when each is the right choice, and why cosine and dot product coincide once vectors are normalized. A standard interview question, usually phrased as 'which metric would you use here and why'.

on this pageshow

questions

4

In vector search, what does cosine similarity measure, and when is it preferred over Euclidean distance?

level: juniorimportance: must knowfreq 80%

answer

  1. angle versus straight-line gap
  2. vector length is divided out
  3. the scale runs -1 to 1
  4. match the model's training objective

basics

~20 s

Cosine similarity measures the angle between two vectors and ignores their lengths, ranging from -1 to 1. Prefer it over Euclidean distance when only direction is meaningful and vector length reflects something you do not want scored, such as text length.

solid answer

~50 s

Cosine similarity is the dot product of two vectors divided by the product of their lengths, so it is the cosine of the angle between them: 1 means they point the same way, 0 means they are perpendicular, -1 means opposite. Dividing out both lengths makes it scale-invariant — doubling a vector does not change its cosine with anything. Euclidean (L2) distance is the straight-line gap between the two points, so it reacts to direction *and* length: two vectors can point identically and still be far apart if one is much longer. For text embeddings, vector length usually reflects incidental properties like text length or token frequency rather than meaning, so cosine is the common default. The decisive rule, though, is to compare with the metric the embedding model was trained under — most sentence-embedding models are trained with a cosine objective, and deviating from that is the real mistake.

go deeper

for a junior

Be able to state that cosine compares the angle and ignores length while Euclidean measures the straight-line gap, and that cosine runs from -1 to 1 with higher meaning more similar.

for a middle

Explain why length is usually noise for text embeddings, and that the deciding factor is the objective the embedding model was trained under rather than a general preference for one metric.

for a senior

Show you would verify the metric end to end — index build, query path, evaluation harness — and name a case like face verification where an absolute distance, not just an ordering, is the product.

for a principal

Own the consequence for the platform: one canonical metric per embedding space, documented alongside the model version, so teams cannot silently mix conventions as models are swapped.

## The two functions Cosine similarity between vectors `a` and `b` is their dot product divided by the product of their lengths: cos(a, b) = (Σ aᵢbᵢ) / (√Σaᵢ² · √Σbᵢ²) Geometrically this is the cosine of the angle between the two arrows. It equals 1 when they point in exactly the same direction, 0 when they are perpendicular, and -1 when they point in exactly opposite directions. Because the length of each vector is divided out, multiplying either vector by any positive constant leaves the value unchanged. Cosine is a *similarity*: higher is better. Euclidean distance, also called L2 distance, is the ordinary straight-line gap between the two points: L2(a, b) = √Σ(aᵢ - bᵢ)² It runs from 0 (identical vectors) upward with no ceiling, and it is a *distance*: lower is better. It responds to both the angle and the lengths. Two vectors pointing in exactly the same direction have cosine 1 but a Euclidean distance of |‖a‖ - ‖b‖|, which can be large. ## Why direction usually carries the meaning in text embeddings An embedding model maps text into a vector space where semantically related texts are placed near one another. What "near" means was fixed by how the model was trained. Most modern text-embedding models are trained with a contrastive objective that scores pairs by cosine similarity: matching pairs are pushed toward a high cosine, mismatched pairs toward a low one. Nothing in that objective constrains how *long* a vector is, so length ends up encoding whatever happens to fall out of the architecture — typically things correlated with token count, token frequency, or how strongly the model reacts to the input. Those are rarely the ranking signal you want. Cosine strips them out and compares meaning-as-direction only. This is why a search over documents of wildly different lengths — a two-line note and a ten-page report — should usually not have the report win simply because its vector came out bigger. ## Where Euclidean is genuinely the right call Three cases come up in interviews. First, when the embedding space was trained so that absolute distance is meaningful. Face-verification embeddings are the classic example: models trained with a margin-based loss are designed so that two images of the same person land within a fixed radius of each other, and the deployed system compares an L2 distance to a fixed acceptance threshold. Here the distance value itself, not just the ordering, is the product. Second, when the vectors are not embeddings at all but genuine coordinates — positions, physical measurements, sensor readings — where length is a real quantity and discarding it would be throwing away data. Third, and most importantly in practice, it does not matter at all when the vectors are unit-normalized: with all lengths equal to 1, Euclidean distance is a strictly decreasing function of cosine, so the two produce exactly the same ordering. Many systems normalize on write precisely so this choice stops mattering. ## The rule that actually decides The interview-grade answer is not "cosine is better for text." It is: use the metric the model was trained with, and if you do not know what that was, check the model's documentation rather than guessing. A model trained with a cosine objective and queried with raw Euclidean distance will still return plausible-looking results — plausible enough that nobody notices — while quietly losing recall on exactly the cases the length variation touches. ## Practical edge cases - Cosine is undefined for a zero vector (division by zero). Empty or whitespace-only inputs can produce degenerate vectors; guard the input rather than the metric. - Cosine of exactly -1 essentially never appears with text encoders; encoders do not place opposites at 180 degrees. Unrelated text tends to land somewhere near perpendicular, so the interesting part of the scale is compressed toward the positive end, and a raw cosine value is not a probability. - Cosine similarity is not a metric in the mathematical sense (it does not satisfy the triangle inequality); some index structures that assume a true metric want the distance form instead. - Whatever you choose, use the same metric everywhere: at index build time, at query time, in your evaluation harness, and in any cached scores. Mixed metrics across those stages is a common and silent source of bad results.

  • What if the model you are using was trained with a dot-product objective rather than a cosine one?
    Then use the dot product. In such models the vector length is part of the learned signal — it can encode a prior like item quality or confidence — and normalizing it away discards information the training put there deliberately. The general rule holds: the query-time metric should match the objective the model was optimized under, otherwise you are scoring pairs with a function the model never saw.
  • Would you ever see a cosine similarity of -1 between two text embeddings?
    Practically never. Encoders do not place antonyms or contradictions at 180 degrees; "hot" and "cold" are used in similar contexts and sit close together. Unrelated text typically lands near perpendicular, so real cosines cluster in a band rather than spanning the full range. That is why a raw cosine value should be interpreted relative to other scores in the same system, not read as an absolute confidence.
  • Your corpus mixes one-line FAQ entries with long PDF sections. Does the metric choice alone fix the length problem?
    No. Cosine removes the raw magnitude effect, but a long section still dilutes its topic across many sentences, so its direction genuinely drifts from any single query. The real fix is chunking to comparable granularity before embedding; the metric choice only removes the crude length-times-score artefact on top of it.

saying these in an interview costs you the question

  • Says cosine similarity ranges from 0 to 1
  • Treats cosine as a distance where lower is better
  • Claims cosine and Euclidean always produce identical rankings
  • Assumes a longer document vector means a more relevant document
  • Picks the metric by habit without checking the model's training objective

context

open as a page

Why do cosine, dot product and L2 rank identically once vectors are unit-normalized?

level: middleimportance: must knowfreq 62%

basics

~20 s

Unit-normalized vectors have length 1, so the dot product equals the cosine, and squared Euclidean distance equals 2 minus twice that cosine. All three are monotone functions of one quantity, so the ordering is identical even though the numbers differ.

open as a page

When one vector store returns a distance and another a similarity score, what bug follows?

level: middleimportance: should knowfreq 38%

basics

~20 s

Distance is better when lower and similarity is better when higher. Code that sorts both the same way returns the worst matches from one of them, and any minimum-score threshold inverts into a keep-only-the-worst filter. Nothing crashes; results merely become quietly wrong.

open as a page

In a dot-product recommender, popular titles dominate results — why, and is that a bug?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Dot-product scores grow with vector length, and items with abundant training interactions often end up with longer vectors, so popular titles outrank better-matching niche ones. Whether that is a defect depends on whether popularity is a signal you deliberately want in the score.

open as a page