skip to content

How does inverse propensity weighting turn position-biased clicks into an unbiased relevance signal?

level: middleimportance: should knowfreq 48%

answer

  1. reciprocal of the exposure chance
  2. rare evidence counts for more
  3. expectation cancels the rank factor
  4. long-tailed weights, unstable sums
  5. floor the propensity to cap the weight

basics

~20 s

Divide each logged click by the probability that its slot was examined, so a rank-8 click counts far more than a rank-1 click. In expectation that recovers relevance, but rare deep clicks make the estimate noisy.

solid answer

~50 s

Under the examination hypothesis a click's expected value is the slot's examination propensity times the item's relevance, so dividing the click by that propensity gives something whose expectation is relevance alone. Concretely, each logged click at rank r contributes `1 / p_r` instead of 1, and the estimate is the sum of those weights over all impressions of the item. If slot 8 is examined about 15 percent as often as slot 1, a single click there counts for roughly seven ordinary clicks. The price is variance: the weights have a long tail, and one lucky deep click can dominate a whole item's score. The standard mitigations floor the propensities at, say, 0.05 so no weight exceeds twenty, or self-normalise by dividing by the sum of the weights instead of by the impression count. Both trade a little bias for much less variance.

code

python · 20 lines
python
# examination propensity by slot, normalised so rank 1 equals 1.0
propensity = {1: 1.00, 2: 0.62, 6: 0.21, 9: 0.14}

# (rank shown, clicked) for two listings of equal rater-graded quality
item_a = [(1, 1), (1, 0), (2, 1), (2, 0), (1, 0)]   # always shown high
item_b = [(6, 1), (9, 0), (9, 0), (6, 0), (9, 0)]   # always shown deep

def naive_ctr(impressions):
    return sum(c for _, c in impressions) / len(impressions)

def ipw_relevance(impressions):
    return sum(c / propensity[r] for r, c in impressions) / len(impressions)

for name, impressions in (("A high", item_a), ("B deep", item_b)):
    print(name,
          "naive CTR", round(naive_ctr(impressions), 2),
          "| IPW relevance", round(ipw_relevance(impressions), 2))

# A high naive CTR 0.4 | IPW relevance 0.52
# B deep naive CTR 0.2 | IPW relevance 0.95

go deeper

for a junior

Know the direction of the correction: divide by the chance the slot was seen, so clicks from deep positions count for more, not less. Getting that direction backwards is the classic slip.

for a middle

Be able to derive it in two lines from the click factorisation and show that the rank factor cancels in expectation. Then state the catch — heavy-tailed weights — and name flooring or self-normalisation as the fix.

for a senior

Talk about operating it: what you floor the propensity at and why, how you monitor the weight distribution, how you sanity-check with a flatter and a steeper propensity curve, and when the deep ranks simply do not support a conclusion.

for a principal

Frame it as a bias-variance budget across the whole measurement stack, and be ready to argue when a slightly biased but stable estimate serves decision-making better than an unbiased one nobody can reproduce week to week.

## The estimator Start from the examination hypothesis: a click at rank `r` on item `d` happens only if the slot was examined and the item was relevant, so ``` E[click] = p_r * relevance(d) ``` where `p_r` is the examination propensity of rank `r`. Rearranged, that is a recipe. Divide the observed click by the propensity of the slot it happened in and the expectation is the quantity you wanted: ``` E[click / p_r] = (p_r * relevance(d)) / p_r = relevance(d) ``` So the **inverse propensity weighted** estimate of an item's relevance over N logged impressions is ``` rel_hat(d) = (1/N) * sum over impressions of ( click_i / p_{r_i} ) ``` This is the Horvitz-Thompson idea from survey sampling: if you sampled one group one time in twenty and another one time in two, you weight each respondent by the reciprocal of their sampling rate to reconstruct the population. Here the sampling rate is the chance the user's attention reached that slot. The same weights are what you attach to the training examples when the debiased clicks are fed to a ranking model: a click harvested from a rarely-examined slot is rare evidence, so it should move the model more. ## Reading the weights A typical estimated propensity curve on a results page might run 1.00, 0.62, 0.44, 0.33, 0.28 down to roughly 0.15 at rank 8. Notice two things. - Propensities are usually reported **relative to rank 1**, normalised so `p_1 = 1`. A common scale factor multiplies every weight equally, so it cancels whenever you compare items or minimise a weighted loss. - The reciprocal of a small number is a big number. At `p_8 = 0.15` a click is worth `1/0.15`, about 6.7 ordinary clicks. At `p = 0.02` it is worth fifty. ## What happens to non-clicks An impression with no click contributes zero to the numerator but still counts in `N`. It is not a negative label and it does not get an inverse weight; the correction acts on the clicks you observed, and the denominator keeps the estimate an average rather than a total. Deleting non-clicked impressions before weighting is a live production bug, because it removes exactly the evidence that an item was given its chance and refused. ## The variance problem Inverse propensity weighting buys unbiasedness and pays in variance, and the payment can be brutal. The variance of the estimator grows roughly with the reciprocal of the propensities, so the deepest slots — the ones you most wanted to rescue — carry the heaviest and noisiest weights. One click at rank 20 with an estimated propensity of 0.03 injects a term of 33 into a sum that is otherwise made of ones and zeros. Item scores start looking like a lottery over which deep click happened to land. Three standard responses: 1. **Clipping, also called propensity flooring.** Replace `p_r` with `max(p_r, tau)` for some floor such as 0.05. The maximum weight is then capped at `1/tau`. This is deliberately biased — buried slots are now under-credited — but the mean squared error usually improves a lot, because variance was dominating. `tau` is a tuning knob with a clear interpretation: the largest correction you are willing to trust. 2. **Self-normalisation.** Divide by the sum of the weights instead of by `N`. The self-normalised estimator is slightly biased in finite samples but consistent, and it is far more stable when one weight is huge, because that weight inflates the denominator too. 3. **Collecting better propensities where it matters.** If the weights at ranks 15 and beyond are untrustworthy, that region of the log may simply not support conclusions, and the honest move is to restrict the analysis rather than extrapolate. ## What the correction assumes Unbiasedness is conditional on three things, and a candidate who states them is answering at the right depth. - **The propensity model is right.** If your estimated `p_r` are wrong, you have swapped one bias for another. Sensitivity analysis — re-running with a flatter and a steeper curve — is cheap insurance. - **Every propensity is strictly positive.** A slot that is never examined has an undefined weight and contributes no identifiable information; no amount of data recovers it. This is the support condition. - **Examination really does depend only on rank** in the way the model says. Trust bias and presentation effects violate it, and no amount of reweighting fixes a mis-specified generative model. ## The one-line summary Inverse propensity weighting does not create information. It re-weights the information you have so that evidence collected under unequal exposure is combined fairly, and it charges you variance for every slot that was rarely looked at.

  • What do you do with impressions that got no click?
    Keep them. They contribute zero to the weighted numerator but they stay in the denominator, which is what makes the result a rate rather than a total. They are not negative labels either — a zero at a low-propensity slot mostly means the slot was never examined. Dropping non-clicks before weighting reintroduces exactly the bias you were removing, and it inflates every item that was ever shown at the top.
  • What breaks if an estimated examination propensity is zero?
    The weight is undefined and the estimate is not identified for that slot: no quantity of data tells you what would have happened at a position nobody ever looks at. In practice you floor the propensity, exclude that stratum and say so, or fix the data collection so every slot you care about has a positive chance of being examined.
  • Why does clipping propensities usually improve the estimate even though it adds bias?
    Because mean squared error is bias squared plus variance, and in click logs the variance term dominates. Capping the weight at, say, twenty removes the few enormous terms that were swinging item scores between refreshes, at the cost of systematically under-crediting the deepest slots. It is a knob you tune and report, not a silent default.

It is survey weighting. If you polled one neighbourhood one household in two and another one in twenty, you multiply each reply by the reciprocal of its sampling rate before combining them.

saying these in an interview costs you the question

  • Multiplies clicks by the propensity instead of dividing by it
  • Claims inverse propensity weighting reduces variance
  • Drops non-clicked impressions before applying the weights
  • Calls the estimate unbiased without questioning the propensity model
  • Ignores that a zero propensity makes the quantity unidentifiable

context