Why does OpenAI recommend cosine similarity for comparing its embeddings?
answer
- Direction matters, magnitude does not
- Returned vectors already have length 1
- Dot product is the same thing, cheaper
- Same ranking as Euclidean distance
- No universal relevance threshold
basics
~20 sOpenAI embeddings are returned normalised to unit length, so cosine similarity reduces to a plain dot product and gives exactly the same ranking as Euclidean distance. Cosine is recommended because it is the cheapest and least error-prone of the equivalent choices.
solid answer
~50 sCosine similarity measures the angle between two vectors and ignores their magnitude, which is the right notion of "similar meaning" for text embeddings. The specific reason it is OpenAI's recommendation is that its embeddings come back L2-normalised — every returned vector already has length 1. For unit vectors, cosine similarity is algebraically identical to the dot product, so you can skip the norm divisions entirely, and the ranking produced by Euclidean distance is identical to the ranking produced by cosine. In other words the metric choice does not change your top-k; it only changes how much arithmetic you do. The practical caveats matter more than the maths. Scores are only comparable within one model and one dimension width, absolute values cluster in a narrow positive band rather than spanning -1 to 1, and if you truncate vectors client-side you must re-normalise or the equivalence quietly breaks.
go deeper
Say that cosine similarity compares direction rather than length, that higher means more similar, and that it is the standard way to compare two OpenAI embeddings.
Explain that the vectors are returned unit-normalised, so cosine equals the dot product and ranks identically to Euclidean distance — the metric choice is about cost, not results.
Show the operational judgment: no portable relevance threshold, scores incomparable across models and widths, and re-normalisation required after averaging or truncating vectors.
Set the house rule that similarity scores are ranking signals, never confidence figures exposed to users or hard-coded as thresholds, and that any transform of a stored vector re-normalises before write.
## The three candidate metrics Given two vectors a and b: - **Cosine similarity** = (a · b) / (‖a‖‖b‖) — the cosine of the angle between them, in [-1, 1]. Magnitude-invariant. - **Dot product** = a · b — the same numerator without the normalisation. Sensitive to magnitude. - **Euclidean (L2) distance** = ‖a − b‖ — straight-line distance. Sensitive to magnitude. For text embeddings, the semantic signal lives in *direction*, not length. Two documents about the same subject should point the same way regardless of how long they are. That is the general argument for cosine over raw dot product in embedding search. ## The specific reason for OpenAI's models OpenAI's embedding endpoints return vectors already normalised to unit L2 length. That single fact collapses the distinction: - With ‖a‖ = ‖b‖ = 1, the cosine formula's denominator is 1, so **cosine similarity equals the dot product exactly**. - Expanding the squared Euclidean distance gives ‖a − b‖² = ‖a‖² + ‖b‖² − 2(a · b) = 2 − 2(a · b). Distance is a strictly decreasing function of the dot product, so **ranking by smallest Euclidean distance and ranking by largest cosine similarity produce the identical ordering**. So the honest answer to "which metric should I use" is: for these embeddings it does not change your results. Cosine is recommended because it is the conventional default, its scores are interpretable on a fixed scale, and computing it as a bare dot product is the cheapest option — you skip two norm computations per comparison, which matters when you are scanning millions of vectors. ## Configuring a vector store Most vector databases ask you to pick a distance metric at index creation. For OpenAI embeddings, choose cosine — or the store's "inner product" / "dot" metric, which for unit vectors is the same thing and is often marginally faster because the engine skips normalisation. What you must not do is mix: if the store normalises on ingest but you feed it vectors you truncated yourself without re-normalising, or you build with one metric and query assuming another, results degrade in ways that look like a bad model rather than a bad configuration. ## Reading the scores A classic trap. Cosine similarity is mathematically bounded by [-1, 1], but real OpenAI text embeddings rarely produce negative or near-zero scores between natural-language strings. Unrelated sentences commonly land somewhere in the 0.1–0.4 region and closely related ones in the 0.5–0.9 region, with the exact bands depending on the model and the domain. Consequences: - **There is no universal relevance threshold.** A hard-coded "0.8 means relevant" cutoff copied from a blog post will silently under- or over-retrieve on your corpus. Calibrate against a labelled sample, or avoid absolute thresholds and take top-k plus a reranking step. - **Scores are not probabilities** and should not be shown to users as confidence percentages. - **Scores are not comparable across models or widths.** A 0.62 from `text-embedding-3-small` and a 0.62 from `text-embedding-3-large` mean different things, and comparing vectors of different widths is not defined at all. ## Where the equivalence breaks The cosine/dot equivalence holds only while your vectors are unit length. Two ways to lose that: 1. **Client-side truncation.** Slicing a vector to its first N coordinates yields something shorter than unit length, by an amount that varies per vector. Dot product then partly measures how much norm survived truncation rather than semantic similarity — which distorts ranking, not just scale. Re-normalise after any manual truncation. 2. **Post-processing.** Averaging several chunk embeddings to represent a document, or adding vectors together, produces a non-unit result. Re-normalise before storing. In both cases cosine similarity still behaves correctly because it normalises internally; it is the dot-product shortcut that becomes wrong. That is an argument for configuring cosine rather than inner product in a store where you cannot guarantee ingest hygiene. ## What an interviewer is checking That you know the vectors are normalised and can say what follows from it, rather than reciting "cosine is standard for embeddings" as folklore. The follow-on judgment — no magic threshold, scores incomparable across models, re-normalise anything you transform — is what separates someone who has run a retrieval system from someone who has read about one.
- If cosine and Euclidean give the same ranking here, why does the choice ever matter?Cost and configuration hygiene. Cosine on unit vectors reduces to a dot product, which is the cheapest operation and what most ANN indexes optimise for. Euclidean does extra work for the same ordering. The choice also matters the moment your vectors stop being unit length — after averaging or manual truncation — because only cosine self-corrects.
- A colleague hard-codes 0.8 cosine similarity as the relevance cutoff. What is wrong with that?Score distributions are corpus- and model-specific; OpenAI similarities cluster in a narrow positive band rather than spreading over [-1, 1]. A threshold tuned elsewhere will drop relevant chunks or admit noise on your data. Retrieve top-k instead, calibrate any cutoff against a labelled sample, and prefer a reranking stage over an absolute score gate.
- You average the embeddings of ten chunks to represent a whole document. What must you do next?Re-normalise the average to unit length before storing it. The mean of ten unit vectors is generally shorter than 1, and by a different amount for each document depending on how coherent its chunks are. Left uncorrected, a dot-product index would systematically rank internally-consistent documents higher for reasons unrelated to the query.
saying these in an interview costs you the question
- Thinks cosine and Euclidean give different top-k for unit vectors
- Treats similarity scores as probabilities or confidence
- Copies a fixed relevance threshold across models or corpora
- Compares scores from two different embedding models
- Assumes averaged or truncated vectors stay unit length