skip to content

In k-NN over session page-view counts, why might cosine distance beat Euclidean?

level: middleimportance: should knowfreq 45%

answer

  1. direction kept, length divided out
  2. same mix, five times the volume
  3. nuisance dimension removed for free
  4. you also lose volume as a signal

basics

~20 s

Cosine distance compares the direction of two count vectors and discards their length, so two sessions with the same browsing mix are treated as close even when one user viewed five times as many pages.

solid answer

~50 s

Cosine distance is `1 - (x . y) / (||x|| * ||y||)`: it divides out each vector's length, so only the relative profile of the counts survives. On session page-view counts that is usually what you want. A visitor who viewed 2 product pages, 4 blog pages and 6 help pages and one who viewed 4, 8 and 12 have identical interests, but Euclidean puts them far apart because the second simply browsed more; cosine puts them at distance 0. Total activity is a *nuisance* dimension here, and cosine removes it. The price is that you also throw away activity volume, which may itself be predictive — of churn, of intent to buy. If so, keep cosine on the profile and reintroduce total views as a separate scaled feature, rather than expecting one metric to carry both signals.

go deeper

for a junior

Know that cosine compares the shape of two vectors and ignores how big they are, and that a session with twice as many views of everything is a perfect cosine match to the original.

for a middle

Explain the normalisation by each vector's length, work a small count example both ways, and state plainly what information the switch throws away.

for a senior

Justify the metric from the prediction target: nuisance volume versus predictive volume, and how you would keep the discarded magnitude as an explicit feature rather than losing it.

for a principal

Own the position that the metric is part of the model's inductive bias — choosing cosine encodes a scale-invariance assumption about users, and that assumption should be stated, tested and revisited, not inherited from a tutorial.

## The quantity being computed For two non-zero vectors, cosine similarity is the dot product divided by the product of the lengths, and cosine distance is one minus that: ``` cos_sim(x, y) = (sum_j x_j * y_j) / ( sqrt(sum_j x_j^2) * sqrt(sum_j y_j^2) ) cos_dist(x, y) = 1 - cos_sim(x, y) ``` For non-negative data such as counts, the similarity lies in [0, 1] and the distance in [0, 1] too: 0 when one vector is a positive multiple of the other, 1 when they share no active feature at all. ## Why magnitude is a nuisance for session data Represent a browsing session as counts over page categories. Visitor A has (2, 4, 6): a little product browsing, more blog, most help pages. Visitor B has (4, 8, 12): the same shape, five-and-a-bit times the volume — a longer session, a bored user, a bot, or someone on a fast connection. Euclidean distance between them is `sqrt(4 + 16 + 36) = 7.5`, larger than the distance from A to plenty of visitors with genuinely different interests. Cosine distance is exactly 0, because B is `2 * A` and scaling a vector does not change its direction. Which is right depends on what you are predicting. If the label is *what kind of visitor is this* — support-seeker, comparison shopper, researcher — the total volume is noise from session length and device, and Euclidean spends most of its budget on that noise. Cosine deletes the nuisance dimension in one step, without you having to model session length. If the label is *will this visitor convert* and heavy browsing is a genuine signal, cosine has just thrown away the best feature you had. ## The honest tradeoff State it as a tradeoff, not a preference. Cosine buys invariance to overall scale of the row; it pays with total blindness to it. Two useful ways to keep both: 1. Use cosine on the count profile and add total page views back as an explicit, separately scaled feature. Then the model sees profile and volume as different things and can weight them. 2. Convert counts to shares — divide each row by its own total so the entries sum to 1 — and use a summed metric such as Euclidean or Manhattan on the shares. This is similar in spirit to cosine (both remove the row's overall size) though not identical, since the two normalise by different notions of length. ## Things people get wrong **Cosine does not remove the need to put features on a common scale.** It removes the length of the whole vector, not a mismatch of units between columns. If one column is in dollars and another counts clicks, the dollar column supplies most of both the dot product and each norm, so it still dominates the angle. Cosine solves "this user is five times more active"; it does not solve "this column is a thousand times bigger than that one". **Cosine distance is not a true metric.** `1 - cos_sim` can violate the triangle inequality, so anything that assumes a metric — including pruning arguments used by some search structures — is not automatically valid under it. The angle itself, `arccos(cos_sim)`, is a metric on the relevant domain if you need one. **The all-zero vector has no direction.** A session with no page views at all makes the denominator zero and the cosine undefined. Decide explicitly what happens to such rows: drop them, or route them to a default prediction. **Sign matters for non-count data.** On features that can be negative, cosine similarity ranges over [-1, 1] and cosine distance over [0, 2], and two vectors pointing in opposite directions score as maximally dissimilar rather than merely different. That is often reasonable, but it is worth knowing before applying cosine to standardised, centred columns, where every feature has both signs and the geometric interpretation is much less intuitive than on raw counts. ## How to answer it in an interview Name the invariance, name the price, and name the data property that decides. Something like: cosine measures profile and ignores volume, which is right when volume is an artefact of session length and wrong when volume is the signal; on page-view counts I would default to cosine for interest-type prediction and keep total views as a separate feature if conversion is the target. That answer shows you are choosing a metric because of a property of the data rather than reciting a rule that cosine is for text and Euclidean is for numbers.

  • What have you given up by switching those count features to cosine distance?
    Total activity. Any signal carried by how much the visitor did — long sessions correlating with intent, tiny sessions with bounces — is divided out before the neighbours are found. If that signal matters, keep the cosine comparison on the profile and add total page views back as its own scaled feature, so the model can weight profile and volume separately instead of one metric silently deciding for you.
  • How would you handle a session vector that is all zeros?
    Handle it explicitly, because cosine is undefined there: the vector has zero length, so the denominator is zero and there is no direction to compare. Options are to filter such rows out before the neighbour search, or to route them to a fixed fallback prediction. What you must not do is let a silent zero-division or a substituted default distance decide, since every empty session would then land in the same arbitrary neighbourhood.
  • Is cosine distance a true metric?
    No. One minus cosine similarity can violate the triangle inequality, so guarantees that depend on a metric do not transfer to it automatically. The angle itself, arccos of the cosine similarity, is a proper metric. In practice this rarely changes results for a brute-force neighbour scan, but it matters if you rely on triangle-inequality pruning to skip candidates.

Two shopping baskets with the same recipe in different portion sizes: cosine asks what the dish is, Euclidean asks how many people it feeds.

saying these in an interview costs you the question

  • Says cosine also removes the need to scale mismatched columns
  • Calls cosine 'the text metric' with no reason given
  • Ignores that discarding magnitude can discard real signal
  • Assumes cosine distance obeys the triangle inequality
  • Never mentions the undefined all-zero vector case

context