In the vector space model, why is cosine similarity used instead of a raw dot product?
answer
- magnitude versus direction
- a padded document should not win
- divide by both vector lengths
- equals a dot product after normalising
- the angle, not the size
basics
~20 sA raw dot product grows with vector magnitude, so long documents win simply for containing more terms. Cosine divides by both vector lengths, comparing the direction of the vectors — the mix of terms — rather than their size.
solid answer
~50 sThe vector space model puts documents and queries in the same space, one dimension per vocabulary term, with a TF-IDF weight in each coordinate. Similarity is then a geometric question. The **dot product** sums the products of matching coordinates, so it rewards magnitude: a document three times longer with the same term proportions has roughly three times the weights and three times the dot product, without being any more relevant. **Cosine** divides the dot product by the Euclidean norm of both vectors, which is exactly equivalent to taking the dot product of the L2-normalised vectors. What survives is the angle between them — the *composition* of the document, not its size. That is why engines that pre-normalise their vectors can use a plain dot product and still be computing cosine. Cosine's flat length correction is itself imperfect, which motivated pivoted length normalisation and later BM25's tunable length parameter.
code
python · 18 linesimport math
def dot(a, b):
return sum(a.get(t, 0.0) * w for t, w in b.items())
def norm(v):
return math.sqrt(sum(w * w for w in v.values()))
def cosine(q, d):
denom = norm(q) * norm(d)
return dot(q, d) / denom if denom else 0.0
# doc_long has identical term proportions to doc_short, just twice the weights
doc_short = {"coffee": 1.0, "roast": 2.0}
doc_long = {"coffee": 2.0, "roast": 4.0}
query = {"coffee": 1.0, "roast": 1.0}
# dot: 3.0 vs 6.0 -> long wins on size alone
# cosine: identical -> same direction, same scorego deeper
Know that documents and queries become vectors of term weights and that similarity is the angle between them. Recall that cosine ignores vector length while a raw dot product does not.
Derive the bias concretely: show that doubling a document's weights doubles its dot product but leaves its cosine unchanged, and state that cosine equals the dot product of L2-normalised vectors.
Discuss the engineering consequences — normalise once at index time and score with a dot product, drop the query norm since it cannot reorder — and name cosine's over-penalty on long documents as the reason pivoted normalisation exists.
Reason about when a geometric bag-of-words similarity is the right abstraction at all, versus a probabilistic scorer or a learned relevance model, and what each choice implies for tuning surface and explainability.
## The model The vector space model represents each document as a point in a space with one dimension per distinct term in the vocabulary. Coordinate i of document d holds the weight of term i in d — classically its TF-IDF weight, zero for terms the document does not contain. The query is embedded in the same space as a very sparse vector. Retrieval becomes: find the document vectors closest to the query vector. The vocabulary is huge and the vectors are almost entirely zeros, which is precisely why an inverted index is efficient — only dimensions the query touches need to be visited at all. ## Dot product and its bias The dot product of query q and document d is the sum over terms of q_i x d_i. It is the natural first choice: it accumulates evidence from every term the two share, weighted by how strongly each holds that term. Its flaw is scale sensitivity. Suppose document A is a 300-word article about coffee roasting and document B is the same article concatenated with itself ten times. B contains each term ten times as often, so its weights are roughly ten times larger and its dot product with any query is roughly ten times larger. B is not more relevant — it is the same document, padded. More realistically, a long reference manual mentioning a term in passing accumulates a bigger dot product than a short, focused page about exactly that term. Magnitude, in this space, mostly encodes verbosity. That is not the signal you want ranking your results. ## What cosine does Cosine similarity is ``` cos(q, d) = (q . d) / (||q|| * ||d||) ``` where ||v|| is the Euclidean (L2) norm, the square root of the sum of squared weights. Dividing by both norms removes magnitude entirely, leaving the cosine of the angle between the vectors. Values run from 0 (no shared terms, orthogonal) up to 1 (identical direction). The document that wins is the one whose *proportional mix* of terms best matches the query's. Two practical consequences follow. First, cosine similarity is identical to the dot product of L2-normalised vectors, so systems that normalise once at index time can score with a plain dot product and still be computing cosine — which matters, because a dot product is far cheaper in an inner loop. Second, the query norm ||q|| is the same for every document in a single query, so it does not change the ranking and can be dropped when you only need an ordering, not a comparable absolute number. ## Cosine versus Euclidean distance Euclidean distance also has a magnitude problem, and a worse one: two documents about the same topic but of very different lengths sit far apart in Euclidean terms even though they point the same way. For normalised vectors the two measures become monotonically related — ranking by cosine descending and by Euclidean distance ascending give the same order — so the distinction collapses once you normalise. Unnormalised, cosine is the safer default for text. ## Where cosine normalisation goes wrong Cosine imposes a *fixed* correction: divide by the vector norm, always, regardless of collection. Empirical work on real collections found this over-corrects. Long documents genuinely are relevant more often than a strict proportional correction predicts — they cover more ground and satisfy more information needs — yet cosine pushes them down hard. The response was **pivoted length normalisation**: identify the length at which the probability of relevance crosses the probability of retrieval, and tilt the normalisation curve around that pivot so short documents are not artificially favoured. This is the direct ancestor of BM25's length handling, where a single parameter interpolates between no length normalisation at all and full normalisation against the collection's average document length. BM25 also drops the strict geometric framing: it is not computing an angle, it is summing per-term contributions from a probabilistic model, with length appearing inside the saturation denominator. ## Beyond term vectors The same geometry reappears with dense embeddings, where a model maps text to a few hundred or a few thousand dense dimensions rather than one per vocabulary term. Cosine remains the standard similarity there for the same reason — direction encodes meaning, magnitude often encodes uninteresting artefacts of the encoder — and normalising once at index time so a dot product suffices is the usual optimisation. ## What interviewers listen for That you can state the bias concretely ("the dot product rewards length"), know cosine equals the dot product of normalised vectors, and can say why the query norm is irrelevant to ranking. A strong answer adds that cosine's flat correction over-penalises long documents, which is exactly the gap pivoted normalisation and BM25's length parameter were designed to close.
- If documents are L2-normalised at index time, does using a dot product change the ranking?No. Cosine is by definition the dot product of the normalised vectors, so normalising once at index time and scoring with a plain dot product produces exactly the same order. It is a standard optimisation because a dot product avoids computing norms in the scoring loop. The catch is that the normalisation must be redone whenever a document's weights change.
- Why can the query vector's norm be dropped from the cosine calculation?Because it is constant across every document scored for that query, so dividing by it scales all scores identically and cannot reorder them. Drop it and you save a computation per document. Keep it only when you need scores that are comparable across different queries, which is rarer than people expect and fragile in any case.
- What criticism of cosine length normalisation led to pivoted length normalisation?Cosine applies the same proportional correction everywhere and was found to over-penalise long documents: on real collections long documents are relevant more often than the strict correction implies, because they cover more ground. Pivoted normalisation tilts the correction curve around a pivot length so longer documents are not pushed down as hard, and BM25's tunable length parameter generalises the same idea.
Two recipes can call for the same ingredients in the same proportions, one scaled for a family and one for a banquet. The dot product notices the banquet; cosine notices they are the same dish.
saying these in an interview costs you the question
- Says the dot product is length-independent
- Believes cosine and Euclidean distance rank differently on normalised vectors
- Cannot state that cosine equals a normalised dot product
- Thinks cosine similarity ranges from minus one to one for TF-IDF vectors
- Assumes cosine normalisation fully solves the document-length problem