How does cosine similarity differ from Euclidean distance between two vectors?
answer
- angle versus gap
- one is scale invariant
- dot product over both lengths
- same word mix, different document length
- unit vectors make the two agree
basics
~20 sCosine similarity measures only the angle between two vectors; Euclidean distance also reacts to their magnitudes. Two term-count vectors with the same word mix but different document lengths score cosine 1.0 while sitting far apart in Euclidean distance.
solid answer
~50 sCosine similarity is `cos(theta) = (a . b) / (||a|| * ||b||)` — the dot product divided by both lengths, so it depends only on direction. It runs from 1 (same direction) through 0 (orthogonal) to -1 (opposite direction). Euclidean distance is `||a - b||`, which grows when either vector gets longer even if the direction is unchanged. Take raw term counts `a = (1, 2, 1)` for a short document and `b = (3, 6, 3)` for a longer one with the same word mix: cosine similarity is exactly 1.0, but the Euclidean distance is `||(2, 4, 2)|| = sqrt(24)`, about 4.9. Use cosine when only composition or direction is meaningful and magnitude is an artefact of size; use Euclidean when magnitude carries real information. If both vectors are first scaled to unit length the two agree, since `||a - b||^2 = 2 - 2 * cos(theta)`.
go deeper
Be ready to write both formulas from memory and say in one line what each ignores. Knowing that cosine looks at angle and Euclidean looks at position is the answer most screens are checking for.
Expect to prove the scale invariance rather than assert it, and to show that on unit-length vectors the squared distance equals 2 minus twice the cosine, so the two rankings coincide.
Show judgment about when magnitude is signal and when it is an artefact of size, and flag the trap that cosine removes per-vector length but never fixes coordinates measured in mismatched units.
Own the framing decision: whether magnitude should influence similarity at all is a modelling choice with downstream consequences, and it should be stated explicitly rather than inherited from whatever the first implementation happened to use.
## The two quantities Given two real vectors `a` and `b` of the same dimension, there are two natural ways to ask how alike they are. **Euclidean distance** is the length of the difference: `d(a, b) = ||a - b|| = sqrt(sum over i of (a_i - b_i)^2)` It answers *how far apart are these two points*. Smaller is more similar; the minimum is 0, and there is no upper bound. **Cosine similarity** is the dot product normalised by both lengths: `cos(theta) = (a . b) / (||a|| * ||b||)`, where `a . b = sum over i of a_i * b_i` It answers *do these two arrows point the same way*. Larger is more similar; for real vectors it always lies in `[-1, 1]`, with 1 for the same direction, 0 for perpendicular vectors, and -1 for exactly opposite directions. It is undefined when either vector is the zero vector, because you would divide by zero. ## Where they disagree The difference is entirely about magnitude. Cosine similarity is **scale invariant**: multiply `a` by any positive number `c` and both `a . b` and `||a||` scale by `c`, so the ratio is unchanged. Euclidean distance has no such property — stretching one vector moves it away from the other. The canonical demonstration uses raw term counts. Suppose a short document has counts `a = (1, 2, 1)` over three words, and a longer document with exactly the same word mix has `b = (3, 6, 3)`. Then: - `a . b = 3 + 12 + 3 = 18` - `||a|| = sqrt(1 + 4 + 1) = sqrt(6)`, and `||b|| = sqrt(9 + 36 + 9) = sqrt(54) = 3 * sqrt(6)` - `cos(theta) = 18 / (sqrt(6) * 3 * sqrt(6)) = 18 / 18 = 1.0` Cosine calls them identical, which is what you want if the question is *are these about the same thing*. Euclidean distance calls them quite different: `b - a = (2, 4, 2)`, so `d = sqrt(4 + 16 + 4) = sqrt(24)`, roughly 4.9. Here the distance is mostly measuring that one document is three times as long. ## Sign and range in practice For vectors with only non-negative entries — raw counts, frequencies, non-negative features — every product `a_i * b_i` is at least 0, so the dot product is non-negative and cosine similarity lands in `[0, 1]`. Negative cosine requires components of opposite sign, which appears once values have been centred or can be negative for other reasons. Candidates who assert *cosine is always between 0 and 1* have quietly assumed non-negative data. People often speak of **cosine distance**, usually defined as `1 - cos(theta)`. That reverses the ordering so that smaller means more similar, which is convenient when a system expects a dissimilarity. It is a transformation of the similarity, not a new measurement. ## The unit-length bridge If both vectors are rescaled to length 1 — replace `a` with `a / ||a||` — then `||a|| = ||b|| = 1` and the cosine formula collapses to the plain dot product: `cos(theta) = a . b`. That is why normalisation is such a common preprocessing step: after it, one cheap dot product gives you the angle. Normalisation also reconciles the two measures. Expanding the squared distance between unit vectors: `||a - b||^2 = ||a||^2 + ||b||^2 - 2 * (a . b) = 1 + 1 - 2 * cos(theta) = 2 - 2 * cos(theta)` So `||a - b|| = sqrt(2 - 2 * cos(theta))`, which strictly decreases as cosine increases. On unit-length vectors, ranking by Euclidean distance and ranking by cosine similarity produce the *same order*. The two only diverge when lengths are allowed to vary. ## Choosing between them Ask what the magnitude of your vector means. - If length is an artefact — document length, how many events a user generated, the units a feature happens to be recorded in — cosine similarity strips it out and compares composition. - If length is signal — a physical measurement, a spend amount, a count you genuinely care about — Euclidean distance keeps that information and cosine throws it away. A common failure is applying cosine to features on wildly different scales without any standardisation. Cosine removes the *overall* length of each vector, but it does nothing about one coordinate dominating because it is measured in a larger unit; that coordinate still drives the dot product. Scale invariance per vector is not the same as per-feature comparability.
- What happens to cosine similarity if you multiply one of the two vectors by 10?Nothing. Both `a . b` and `||a||` scale by 10, so the ratio is unchanged — cosine similarity is invariant to multiplication by any positive scalar. A negative scalar is different: it reverses the direction, so multiplying by -10 flips the sign of the similarity. Euclidean distance, by contrast, changes a great deal under either.
- If both vectors are rescaled to unit length, does Euclidean distance rank pairs the same way as cosine similarity?Yes. For unit vectors, `||a - b||^2 = 2 - 2 * (a . b)`, and `a . b` is exactly the cosine. Distance is therefore a strictly decreasing function of cosine similarity, so the two produce identical orderings — only the scores differ. This equivalence is why normalisation is often applied before a distance-based comparison.
- Can cosine similarity be negative between two vectors of raw word counts?No. Counts are non-negative, so every term `a_i * b_i` in the dot product is at least 0 and the similarity lies in `[0, 1]`; the worst case is 0, meaning the two documents share no words. Negative values need coordinates of opposite sign, which only arises once the data can take negative values.
Cosine similarity is a compass bearing: two hikers heading due north match perfectly whether one walked one kilometre or ten. Euclidean distance is the gap between where they ended up, so the ten-kilometre hiker looks far away.
saying these in an interview costs you the question
- Claims cosine similarity changes when one vector is scaled up
- Says cosine similarity is always between 0 and 1 for any real vectors
- Reports the raw dot product as similarity without dividing by the norms
- Treats a higher cosine value as meaning greater distance
- Thinks cosine similarity fixes features being on different units