skip to content

Why is raw cosine similarity on 1-5 star rating vectors replaced by Pearson or adjusted cosine?

level: middleimportance: must knowfreq 58%

answer

  1. no zero point on a star scale
  2. everyone lands in the positive corner
  3. generous raters look like everyone
  4. subtract the rater's own average first
  5. in item-item, still the user's mean

basics

~20 s

Raw cosine treats every star rating as a positive quantity, so two raters who use the scale differently look similar even when they disagree. Subtracting each rater's own mean first turns ratings into signed deviations, so disagreement becomes negative similarity.

solid answer

~50 s

On a 1-5 scale there is no zero, so every rating vector points into the positive orthant and raw cosine between any two raters is high almost by construction. A cook who rates everything 4 or 5 looks similar to everybody, and worse, two cooks who rank the same recipes in *opposite* order can still score 0.95. The fix is mean-centring: subtract the rater's own average before comparing, so "below my usual" becomes negative and genuine disagreement pushes the similarity down toward -1. Pearson correlation is exactly cosine on user-mean-centred vectors, computed over the co-rated items, and is the standard user-based similarity. Adjusted cosine is the item-based analogue: when comparing two item columns it subtracts the **rating user's** mean, not the item's mean, because the scale bias being corrected belongs to people, not to items.

code

python · 20 lines
python
from math import sqrt

# the 4 recipes both cooks rated, 1-5 stars - opposite orderings
generous = [5, 4, 5, 4]
harsh = [2, 3, 2, 3]

def cosine(u, v):
    dot = sum(x * y for x, y in zip(u, v))
    norm = sqrt(sum(x * x for x in u)) * sqrt(sum(y * y for y in v))
    return dot / norm

def centred(v):
    mean = sum(v) / len(v)
    return [x - mean for x in v]

print(round(cosine(generous, harsh), 3))
# 0.953  -> "almost identical taste"

print(round(cosine(centred(generous), centred(harsh)), 3))
# -1.0   -> Pearson: perfectly opposite taste

go deeper

for a junior

Remember the one-line reason: star scales have no zero, so every rating vector points the same way and raw cosine calls almost everyone similar. Subtracting each rater's average is the standard fix.

for a middle

Write the Pearson formula from memory and say which sums run over the co-rated items. Be precise that adjusted cosine subtracts the rating user's mean even though it compares item columns.

for a senior

Discuss the implementation choices you actually make: all-ratings mean versus co-rated mean, whether to normalise spread as well as level, and how a negative similarity is used rather than thrown away in the prediction.

for a principal

Frame normalisation as a modelling decision about what "agreement" means for your product, and be ready to argue when removing rater bias hurts — for example when the absolute rating level is itself the business signal you are ranking on.

## The defect in raw cosine on ratings Cosine similarity measures the angle between two vectors. That is a sensible summary when the coordinates can be positive or negative and zero means "neutral". A 1-5 star scale has neither property: every observed value is positive, and there is no rating that means zero. Every user's rating vector therefore points into the same corner of the space, and the angle between any two of them is small. Similarities cluster in the 0.85-0.99 band, which destroys the ranking you actually need — you are not asked whether two people are similar, you are asked *which* other people are the most similar. A sharper version of the failure: cosine cannot express disagreement at all on this data. Two cooks in a recipe community rate the same four dishes. One rates them 5, 4, 5, 4; the other rates them 2, 3, 2, 3. They rank the dishes in exactly opposite order — one likes precisely what the other dislikes. Raw cosine on those vectors is about 0.95, which the recommender reads as near-identical taste, and it will happily recommend to each the dishes the other enjoyed. ## Mean-centring Subtract each rater's own average rating from each of their ratings. The generous cook's row becomes `[+0.5, -0.5, +0.5, -0.5]` and the harsh cook's becomes `[-0.5, +0.5, -0.5, +0.5]`. The vectors now point in opposite directions and the cosine between them is -1.0, which is the truth. Centring achieves three things at once: - It removes the **rater's baseline**. A cook whose ratings run 4-5 and one whose ratings run 1-3 can now be compared on what they liked *relative to their own habit*, which is what "agreement" actually means. - It restores a meaningful **sign**. Negative similarity now means anti-correlated taste, and the prediction formula can use it (a neighbour who reliably disagrees with you is informative — just with a flipped contribution). - It puts similarities back on a **spread-out scale**, so a top-k neighbourhood is a real selection rather than a coin flip among values that all round to 0.9. ## Pearson correlation (the user-based form) Pearson correlation between two users is precisely cosine applied to their mean-centred vectors, restricted to the items both rated: ``` sim(u,v) = sum_i (r_ui - mean_u)(r_vi - mean_v) / sqrt( sum_i (r_ui - mean_u)^2 * sum_i (r_vi - mean_v)^2 ) ``` where `i` ranges over the co-rated items. One genuine implementation choice: `mean_u` can be the user's mean over **all** their ratings or only over the **co-rated** items. The all-ratings mean is more stable and cheaper (compute it once per user); the co-rated mean makes the statistic a textbook correlation on that sample but is noisy when the overlap is small. Most systems use the all-ratings mean and say so. ## Adjusted cosine (the item-based form) When the neighbourhood runs over items, the naive move is to centre each item column by that item's own mean rating. That corrects for items being generally well or badly received, but it does **not** correct for the thing that actually distorts the comparison: individual users' scale habits. Adjusted cosine, the standard item-based similarity, therefore subtracts the *user's* mean from every rating, then takes cosine over the users who rated both items: ``` sim(i,j) = sum_u (r_ui - mean_u)(r_uj - mean_u) / sqrt( sum_u (r_ui - mean_u)^2 * sum_u (r_uj - mean_u)^2 ) ``` Notice the same `mean_u` appears in both factors of each product — it is one user contributing a pair of their own deviations. This is the point of the word "adjusted": it is cosine between item vectors, adjusted for who did the rating. Getting this backwards — centring items by item means in an item-item similarity — is a common and quietly damaging mistake, because a community of generous raters then still inflates every item pair together. ## What centring does not fix Centring equalises **level**, not **spread**. A cook who uses only 4 and 5 and one who uses the whole 1-5 range still contribute deviations of very different magnitude. Cosine's normalisation absorbs most of this, since it divides by each vector's length, but if you compare raw centred dot products or feed deviations into a prediction, dividing by each user's standard deviation as well (a z-score normalisation) is the further step. Centring also does nothing about the reliability of a similarity computed from very few co-rated items — that is a separate problem needing a separate fix. ## What to say in an interview Lead with the concrete failure: on a positive-only scale, cosine says "similar" to almost every pair and cannot represent disagreement. Then name the fix and be precise about which mean gets subtracted in which direction — Pearson centres users when comparing users, adjusted cosine centres users when comparing items.

  • In adjusted cosine between two items, whose mean is subtracted and why?
    The rating user's mean, not the item's. The distortion being corrected is that individuals use the star scale differently — one rates almost everything 4-5, another almost everything 1-3 — and that bias travels with the person, so it must be removed per person even when the vectors being compared are item columns. Centring by item means instead would correct for items being generally liked, which is not the source of the inflation.
  • Should a user's mean be computed over all their ratings or only the co-rated items?
    Both are used. The all-ratings mean is stable, computed once per user, and does not change as you compare against different neighbours — it is the common production choice. The co-rated mean makes the statistic a textbook correlation on that exact sample, but with only a handful of overlapping items it is itself a noisy estimate and can swing the similarity wildly. Pick one and state it.
  • What bias does mean-centring leave behind?
    Differences in spread. Centring equalises where each rater's scale sits, not how widely they use it: a cook who only ever gives 4 or 5 produces small deviations, one who uses the full range produces large ones. Cosine's length normalisation absorbs much of this, but if you compare uncentred dot products or aggregate raw deviations, dividing by each rater's standard deviation is the extra step.

saying these in an interview costs you the question

  • Says cosine already handles rating scale differences
  • Centres item columns by item means in adjusted cosine
  • Thinks negative similarity is impossible or must be discarded
  • Claims centring also equalises how widely raters spread scores
  • Computes similarity over all items instead of the co-rated ones

context