skip to content

In a dot-product recommender, popular titles dominate results — why, and is that a bug?

level: seniorimportance: should knowfreq 47%

answer

  1. score grows with vector length
  2. head items accumulate more gradient
  3. compare dot ranking against cosine ranking
  4. magnitude can be a deliberate prior

basics

~20 s

Dot-product scores grow with vector length, and items with abundant training interactions often end up with longer vectors, so popular titles outrank better-matching niche ones. Whether that is a defect depends on whether popularity is a signal you deliberately want in the score.

solid answer

~60 s

The dot product is ‖a‖‖b‖cos, so a candidate can win either by pointing closer to the query or by simply being longer. In recommender embeddings, items with a lot of interaction data frequently acquire larger norms, which lets a mediocre-angle blockbuster outscore a well-matched niche title. That is not automatically a bug: inner-product retrieval is often chosen precisely because the norm carries a learned prior about item quality or engagement, and some two-tower setups even fold an explicit bias term into the vector. It becomes a bug when the effect is unintended and invisible — new or long-tail items are buried, and offline metrics look fine because popular items are also the ones users historically clicked. Diagnose it by correlating norm with rank position and by re-running the same queries under cosine to see how much the top-k moves. Fix it either by normalizing, which makes ranking purely directional, or by keeping the dot product and making popularity an explicit, tunable term you can dial rather than an accident of training.

code

python · 11 lines
python
import numpy as np

def unit(v):
    return v / np.linalg.norm(v)

query = np.array([1.0, 0.0])
niche = np.array([1.0, 0.0])      # perfect direction match, norm 1
popular = np.array([1.8, 2.4])    # cosine 0.6 with query, norm 3

print(query @ niche, query @ popular)                    # 1.0  1.8 -> popular wins
print(unit(query) @ unit(niche), unit(query) @ unit(popular))  # 1.0  0.6 -> niche wins

go deeper

for a junior

Recall that the dot product is affected by how long each vector is, while cosine is not, and that this alone can change which item ranks first.

for a middle

Explain the mechanism as norm times cosine, and show with two numbers how a worse-matching but longer vector outscores a better-matching short one.

for a senior

Demonstrate the diagnosis you would actually run — norm-versus-rank correlation, a cosine-versus-dot top-k diff, recall sliced by item age — and name why offline metrics hide the regression.

for a principal

Own the framing: decide per system whether magnitude is an endorsed prior or an artefact, and insist that any popularity prior lives in a tunable term the business can adjust per surface rather than inside the norms.

## The mechanism Write the dot product in its geometric form: a·b = ‖a‖ · ‖b‖ · cos(θ). Ranking candidates by this quantity for a fixed query means ranking by the product of two things — how well the item's direction matches the query, and how long the item's vector happens to be. Cosine, by construction, keeps only the first factor. So a candidate whose direction is a mediocre match can beat a candidate whose direction is an excellent match, provided its norm is large enough. Concretely: a niche title pointing exactly along the query with norm 1 scores 1.0, while a blockbuster at cosine 0.6 with norm 3 scores 1.8. Under cosine the niche title wins outright; under raw dot product it is beaten by a wide margin. ## Why popularity ends up in the norm Embeddings learned from interaction data are updated once per observed interaction. An item that appears in millions of interactions receives orders of magnitude more gradient signal than one appearing a hundred times, and with common objectives that pushes its representation further from the origin. Nothing in the training explicitly said "popular items should have longer vectors"; it falls out of the data distribution. The result is that a raw inner-product retrieval carries an implicit popularity prior nobody wrote down. ## When it is a feature, not a defect Maximum inner product search is a deliberate design choice in a lot of retrieval systems, and the magnitude sensitivity is the reason. If the business goal is engagement, a mild popularity prior is a genuinely useful default: popular items are popular partly because they satisfy many users. Some two-tower architectures make this explicit by reserving a dimension of the item vector for a learned per-item bias, so that the inner product literally computes "match score plus item prior" in one operation, which the retrieval index can then serve without a second pass. The honest interview answer is therefore not "switch to cosine." It is: decide whether magnitude is signal or noise for *this* system, and make the decision explicit. ## When it really is a bug Three symptoms mark the pathological case. First, cold start: newly added items have short vectors by construction and are structurally unable to reach the top-k, no matter how well they match. Catalogue coverage collapses toward the head. Second, invisible evaluation. Offline metrics computed against historical interaction logs are themselves popularity-biased, so a ranker that over-serves popular items scores *well* on them. The regression shows up only in long-tail recall, catalogue coverage, or a live experiment on new-item discovery. Third, the effect is untunable. Because it lives in the norms rather than in a coefficient, there is no knob. You cannot turn the popularity prior down by 20% for a discovery surface and up for a homepage carousel; you get whatever the training data produced. ## How to diagnose it A few cheap checks settle it: - Correlate item vector norm with the rank position it achieves across a sample of queries. A strong correlation means norm, not match quality, is driving the ordering. - Re-run a fixed query set under cosine and diff the top-k against the dot-product results. If the sets barely overlap, magnitude is doing most of the work. - Plot the norm distribution against item interaction counts. A clear upward trend confirms popularity is what the norms encode. - Slice recall by item age or interaction volume. A cliff at the low end is the cold-start symptom. ## The fix menu Normalizing every vector removes magnitude entirely and makes the ranking purely directional; this is the right move when you concluded the magnitude was an artefact. It costs you nothing at serving time because the resulting ranking is the one cosine would give. Keeping the dot product and adding an explicit popularity or recency term to the final score is the right move when you want the prior but want it controllable. You then normalize the embedding contribution and add the prior as its own weighted feature, so product can tune it per surface. A third option is to leave retrieval as-is and correct downstream: retrieve generously with the dot product, then re-rank a larger candidate set with a scorer where popularity is one explicit feature among many. This is common when retraining the retrieval model is expensive. What does *not* work is normalizing and expecting popularity bias to disappear from the system. Normalization removes only the magnitude channel; if the training data itself over-represents popular items, their *directions* will also sit closer to the average query, and that bias survives untouched.

  • How would you confirm that magnitude, rather than relevance, is driving the ranking?
    Run a fixed query set twice — once under the raw dot product, once with every vector normalized — and diff the top-k. Then correlate item norm with achieved rank position and with interaction count. If the two result sets barely overlap and norm correlates strongly with rank, magnitude is the ranker. Slicing recall by item age exposes the cold-start half of the same effect.
  • When is the magnitude effect exactly what you want?
    When the norm encodes a prior you endorse and the objective is engagement rather than pure semantic match. Some two-tower retrieval models deliberately fold a per-item bias into the vector so a single inner product computes match plus prior. The distinction is intent and control: a prior you chose and can tune is a feature, an accident of gradient counts you cannot dial is a bug.
  • Your offline metrics look great but discovery of new titles has collapsed. What happened?
    The offline set is built from historical interactions, which are themselves popularity-skewed, so a ranker that over-serves popular items scores well on it by construction. New items have short vectors and cannot reach the top-k regardless of match quality. Measure catalogue coverage and long-tail recall as first-class metrics, and validate discovery surfaces with a live experiment rather than replayed logs.
  • Does normalizing the vectors remove popularity bias from the system?
    Only the magnitude channel. If the training data over-represents popular items, their directions also drift toward where typical queries land, and normalization does nothing about that. Treat normalization as removing one specific mechanism, then measure the residual bias separately rather than declaring the problem solved.

saying these in an interview costs you the question

  • Claims dot product and cosine give identical results, so magnitude cannot be the cause
  • Asserts magnitude is always noise and normalizing is always correct
  • Reads a longer vector as meaning a more relevant item by definition
  • Believes normalizing removes popularity bias present in the training data
  • Trusts offline metrics built from historical clicks to catch a popularity skew

context