Why can scoring a held-out item against 100 sampled negatives flip which recommender wins?
answer
- 100 candidates is not the real choice
- rank 500 of 40,000 becomes rank 2
- the tail is compressed toward the head
- distortion differs from model to model
- ordering of methods is not preserved
basics
~20 sSampling 100 negatives replaces the real task, ranking against a whole catalog, with a much easier one. It compresses the tail toward the top, unequally across models, so the sampled winner need not be the full-catalog winner.
solid answer
~50 sThe sampled protocol scores the held-out course against 100 uniformly drawn negatives and asks whether it reaches the top ten of those 101. With a 40,000-course catalog, a course truly ranked 500th is beaten by a random negative only about 1.2 percent of the time, so on average roughly one of the 100 negatives outranks it and the protocol records a near-certain hit for an item that never appears on screen in production. The distortion is a compression of the entire tail toward the top, and it is not identical across models: a model with many plausible items crowding the head loses no more under sampling than one with everything scattered. That is why sampled and full-catalog evaluations have been shown to disagree about which method is better. Score the full catalog when you can; if you sample, fix the sample, share it across models, and never quote the number as the real metric.
code
python · 18 linesimport random
random.seed(7)
CATALOG = 40000
TRUE_RANK = 500 # held-out course is 500th of 40,000 by model score
SAMPLE = 100 # negatives drawn uniformly for the sampled protocol
TRIALS = 20000
p = (TRUE_RANK - 1) / (CATALOG - 1) # chance one negative outranks it
hits = 0
for _ in range(TRIALS):
beaten_by = sum(1 for _ in range(SAMPLE) if random.random() < p)
if beaten_by < 10: # target inside top 10 of the 101
hits += 1
print("one negative outranks target:", round(p, 5))
print("sampled hit rate at 10: ", hits / TRIALS)
print("full-catalog hit at 10: ", TRUE_RANK <= 10)go deeper
Know that the candidate set is part of the evaluation, and that scoring against a hundred random items is a much easier test than choosing from a whole catalog.
Explain the arithmetic: with a catalog of tens of thousands, an item ranked in the hundreds is beaten by only about one of a hundred random negatives, so it looks like a hit that production would never show.
Demonstrate the operational habits - spot saturation, fix and share the negative sample, re-score finalists on the full catalog, and refuse cross-protocol comparisons with published numbers.
Own the principle that a cheap evaluation whose bias depends on the models being compared is not a proxy at all, and decide where the organisation spends compute to keep the comparison sound.
## What the sampled protocol does For each held-out interaction, draw `m` items the learner has not interacted with (commonly `m = 100`), score them plus the held-out item, and evaluate the cutoff metric over that small candidate set. It exists for one reason: scoring 40,000 courses for every held-out learner is expensive, and 101 is cheap. ## The arithmetic of the distortion Suppose the model places the held-out course at rank `r` out of a catalog of `N`. Under uniform sampling, each drawn negative outranks it independently with probability `(r - 1) / (N - 1)`, so the number of sampled negatives above it is roughly Binomial(`m`, `(r-1)/(N-1)`), and the sampled rank is one plus that count. Put `r = 500`, `N = 40000`, `m = 100`. The per-negative probability is about `0.0125`, so the expected count above the target is about `1.25`. The sampled protocol therefore reports the target at roughly second place out of 101 - a comfortable hit inside any reasonable cutoff - while on the full catalog it sits 490 places below a ten-slot shelf and is simply never shown. ## Why this is bias, not noise Two properties matter. **It is compressive.** The map from true rank to expected sampled rank squashes the whole tail into the head. Everything from rank 200 to rank 2,000 lands in the same narrow band of sampled ranks. The metric therefore stops measuring the part of the problem that is hard - separating a few hundred plausible courses from each other - and mostly measures the part every serious model already gets right, namely pushing the obviously irrelevant 39,000 downward. **It is not a constant offset.** Different models have different distributions of true ranks for held-out items. A model that gets many targets into the top 50 and misses badly on the rest can score identically, under sampling, to one that puts almost everything around rank 400. Since the transformation is nonlinear and model-specific, the ordering of models under the sampled metric is not guaranteed to match their ordering on the full catalog, and published comparisons have found that ordering reversing. Increasing `m` shrinks the distortion but only fully removes it as `m` approaches `N`. ## Diagnosing it in your own numbers Saturation is the tell. If your sampled hit rate at ten sits near 0.9 or above across every candidate model while the models clearly differ in other respects, the protocol has stopped discriminating. Re-score the top two candidates against the full catalog; if the gap changes character, the sampled numbers were an artefact. ## Practical positions, in order of preference 1. **Score the full catalog.** For a 40,000-item catalog this is usually affordable, especially if the model factorises into an embedding dot product where the scoring is one matrix product per learner batch. Reserve sampling for catalogs where it genuinely is not. 2. **If you must sample, harden it.** Draw negatives by popularity rather than uniformly, so the candidates are plausible items rather than dead catalog. This makes the task harder and better correlated with the real one, though it still estimates a different quantity. 3. **Make the comparison internally valid.** Fix one negative sample per held-out item, by seed, and reuse it for every model. Redrawing negatives per model injects variance that can easily exceed the difference under test. 4. **Report the protocol with the number.** Sample size, sampling distribution, cutoff, exclusion rule. A sampled hit rate quoted next to a paper's full-catalog figure is a meaningless comparison in both directions. 5. **Confirm the finalists.** Whatever screens the field, the model you propose to put in front of learners is scored under the strictest protocol you can run before anyone argues about shipping it. ## The general lesson A cheap evaluation is not a noisy version of the expensive one. It can be a systematically different measurement whose bias depends on the thing you are trying to compare, which is the worst kind: it does not merely blur the answer, it can invert it.
- Does drawing negatives by popularity instead of uniformly fix the problem?It helps and does not fix it. Popularity-drawn negatives are plausible competitors rather than dead catalog, so scores drop, models separate again, and the sampled task resembles the real one more closely. But it is still a small candidate set estimating a different quantity than full-catalog ranking, and it changes the estimand rather than correcting the bias.
- How many negatives would you need for a sampled comparison to be safe?There is no threshold that makes it equivalent; the distortion only disappears as the sample approaches the catalog. The practical rule is to compare the sample size against how many items a decent model plausibly places above the target. If that count is comparable to the sample size, the metric is saturating and a hit rate near one is a symptom rather than a result.
- What has to be identical across models for a sampled comparison to be even internally valid?The same held-out interactions, the same negative sample for each held-out item fixed by seed, the same cutoff, and the same rule for excluding items the learner already took. Redrawing negatives per model adds variance that can dwarf the effect you are measuring, and changing the exclusion rule quietly changes the denominator.
saying these in an interview costs you the question
- Reports a hit rate over 100 negatives as catalog performance
- Compares a sampled number with a published full-catalog number
- Redraws negatives independently for each model
- Assumes sampling adds only variance, never bias
- Reads a near-perfect sampled hit rate as a strong model