skip to content

Recommenders and Ranking

How a sparse user-item matrix becomes a ranked list: neighborhood similarity, matrix factorization fit by ALS, and pairwise ranking objectives. Cold start and offline metrics are the usual probes.

on this pageshow

explore

questions

page 2 of 2

What does Bayesian Personalised Ranking optimise over its (user, positive, negative) triplets?

level: middleimportance: nice to knowfreq 35%

basics

~20 s

Bayesian Personalised Ranking maximises the probability that a user's observed item scores above an un-observed one. Per triplet it maximises the log of a sigmoid of the two scores' difference, plus regularisation, so only differences matter.

open as a page

In confidence-weighted ALS on purchase counts, why does summing over every user-item cell stay tractable?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

The weighted normal equations split into a term shared by all users plus a small correction from that user's own interactions. The shared item Gram matrix is built once per sweep, so cost tracks observed interactions, not the full grid.

open as a page

How would you keep an item-item similarity matrix fresh for a 2M-item catalog?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

Never materialise the full matrix: two million items give roughly two trillion pairs. Compute only pairs sharing a rater by walking each user's rated list, cap oversized profiles, keep a few hundred neighbours per item, and rebuild on a schedule.

open as a page

Why can scoring a held-out item against 100 sampled negatives flip which recommender wins?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Sampling 100 negatives replaces the real task, ranking against a whole catalog, with a much easier one. It compresses the tail toward the top, unequally across models, so the sampled winner need not be the full-catalog winner.

open as a page

How would you launch a recommender for a new marketplace with no interaction history at all?

level: principalimportance: nice to knowfreq 28%

basics

~10 s

Ship logging before ranking, launch a non-personalised baseline from catalogue metadata and business rules, add content-based personalisation once users have any history, and define in advance the data density that justifies a collaborative model.

open as a page

How do you decide how much relevance to trade for catalog coverage on a feed?

level: principalimportance: nice to knowfreq 31%

basics

~10 s

Do not blend the two into one score. Make relevance a guardrail with an explicit tolerance, make catalog coverage the metric you move, and set that tolerance with a long-horizon experiment.

open as a page

Your implicit-feedback recommender surfaces only chart-toppers — how do you decide whether to correct that popularity bias?

level: principalimportance: nice to knowfreq 37%

basics

~20 s

Measure first: compare the model against a non-personalised most-popular ranker. Popularity is partly real taste, so correct it only where the head demonstrably crowds out items a user would prefer, or where tail supply matters commercially.

open as a page

When is replaying a candidate ranker on last month's impression log a trustworthy estimate?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Only when the log recorded the probability with which each item was shown, the candidate ranker mostly promotes items the old ranker did sometimes show, and enough logged sessions survive importance weighting to give a usable interval.

open as a page

showing 31–38 of 38