skip to content

How do you estimate per-rank examination propensities on a live results page?

level: seniorimportance: should knowfreq 40%

answer

  1. the curve you can plot is confounded
  2. intervene, do not observe
  3. hold the item fixed, move the slot
  4. only ratios are identifiable
  5. a small randomised slice, priced explicitly

basics

~20 s

Perturb the order for a small traffic slice so the same items land in different slots, then compare their click rates across slots. Holding the items fixed makes the ratio an examination ratio rather than a quality difference.

solid answer

~50 s

You cannot read propensities off click-through rate by rank, because the ranker deliberately puts better items at the top, so that curve mixes examination with quality. You need an intervention. The cheapest useful design is swap randomisation: on about 2 percent of sessions, swap the item at rank 1 with the item at some rank k, and compare the click rate the same items get in each slot. Because the item is held fixed, the ratio of click rates estimates `p_k / p_1`, and normalising `p_1 = 1` gives the curve. Run it across a range of k to fill in the ranks you care about. The cost is real: a swap degrades that session, so keep the slice small, prefer tail queries, and stop once the deep-rank estimates stabilise. Cheaper alternatives are intervention harvesting from natural rank variation and fitting a click model on the logs, both free but assumption-heavy.

go deeper

for a junior

Know that the propensity curve has to come from an experiment where the order is deliberately perturbed, not from grouping the existing log by slot.

for a middle

Explain why the observational curve is confounded — the ranker puts better items on top — and how swapping the same item between two slots removes that confound and yields a ratio of examination probabilities.

for a senior

Show you would run it: which slots you swap, what share of traffic, on which query segments, how you price the degraded sessions, and how you know when the deep-rank estimates are precise enough to stop.

for a principal

Own the tradeoff between measurement quality and short-term revenue as an ongoing budget, and decide whether the organisation invests in permanent randomised exposure or accepts model-based propensities and their assumption risk.

## Why you cannot just plot click rate by rank The tempting shortcut is to take the log, group by slot, and call the resulting click-rate curve the propensity curve. It is wrong for one reason: the ranker is not assigning items to slots at random. It is deliberately placing the items it believes are best at the top. The observed curve is therefore the product of two decaying things — examination decays with rank, and so does the quality of what was placed there. You cannot separate them from observational data alone, which is why estimating propensities is an **intervention** problem, not an analysis problem. ## Design 1: swap randomisation The workhorse. On a small slice of sessions — 2 percent is a common starting point — pick a pair of slots and swap the items in them before rendering. The classic pairing swaps rank 1 with rank k. Because the *same items* now appear in both slots across the randomised sessions, relevance is held fixed by construction. Under the examination hypothesis the ratio of the click rates the item collects in slot k versus slot 1 estimates `p_k / p_1`. Normalise `p_1 = 1` and read off the curve. Repeat over a set of k values, or randomise which pair is swapped per session, to cover the ranks you care about. Two practical notes. First, only relative propensities are identifiable, and that is enough: a common scale factor multiplies every inverse weight equally and cancels in any comparison between items or in the minimiser of a weighted loss. Second, precision at deep slots is the binding constraint, because those slots yield very few clicks; the standard error falls with the square root of randomised impressions, so the tail of the curve is what determines how long you run. ## Design 2: full result randomisation Uniformly shuffle the top-n results for a tiny fraction of sessions. This is the cleanest estimator — every item has an equal chance of every slot, so the click rate by slot *is* the propensity curve up to scale — and the most expensive, because a shuffled page is a visibly worse page. Reserve it for a very small slice, or for query classes where the top results are near-interchangeable. ## Design 3: intervention harvesting Free data hiding in the log. Over weeks, the same query-item pair lands at different ranks anyway: rankers get retrained, features drift, scores tie, experiments run. Collect the pairs that appeared at two or more ranks and compare their click rates within pair. The estimate costs nothing in revenue, but it leans on the assumption that whatever moved the item between ranks was unrelated to its relevance for that query — often violated, since the thing that moved it was usually a model that changed its mind about relevance. ## Design 4: click models fitted on the log Write down a generative model with a rank term and a relevance term and fit both by expectation-maximisation over the whole log. The position-based model is exactly the examination hypothesis with a free parameter per rank. The cascade model instead assumes users scan strictly top-down and stop at the first click, which turns skipped-over items into informative negatives. These need no intervention at all, but identification now rests entirely on the model's assumptions being true, and the different click models disagree with each other on real logs. ## Managing the cost of randomisation A senior answer treats the randomised slice as a budget rather than a constant. - **Bound the harm per session.** Swapping neighbouring or near-neighbouring slots hurts far less than shuffling the whole page. Swapping rank 1 with rank 2 is nearly free; rank 1 with rank 10 is not. - **Choose where to spend it.** Tail queries, low-monetisation surfaces and logged-out traffic usually cost less per degraded session than head commercial queries — with the caveat that propensity curves differ across those segments, so you cannot always estimate on cheap traffic and apply the numbers to expensive traffic. - **Size it against the precision you need.** Deep-rank propensities are the noisy ones and also the ones that produce the largest inverse weights. If ranks past 10 will be floored anyway, you do not need to measure them precisely. - **Estimate the revenue drag explicitly.** Roughly, the traffic share times the per-session metric drop measured on the randomised arm. Quote that number to whoever owns the surface before you turn it on, and re-quote it if you extend the run. ## Keeping the curve honest Propensities are not a constant of nature. Segment them by device and layout, because a mobile single column, a desktop grid and a page with promoted slots above the results decay differently. Re-estimate after any redesign that changes where the eye lands. Validate by holding out part of the randomised traffic and checking that the fitted curve predicts the held-out swap outcomes. And carry the uncertainty forward: a propensity with a wide interval produces an inverse weight with a much wider one.

  • Can you estimate propensities without randomising any traffic?
    Yes, two ways, both assumption-heavy. Intervention harvesting mines natural rank variation for query-item pairs that appeared at several slots and compares their click rates within pair; it is free but assumes whatever moved the item was unrelated to its relevance. Alternatively fit a click model such as the position-based model on the whole log by expectation-maximisation. Neither is as trustworthy as a real intervention, and different click models disagree on the same log.
  • How would you decide what share of traffic to randomise?
    Price both sides. The benefit is precision on the deep-rank propensities, and the standard error falls only with the square root of randomised impressions, so doubling the slice buys about forty percent tighter intervals. The cost is the traffic share times the per-session metric drop measured on the randomised arm. Start small, prefer near-neighbour swaps and low-stakes queries, and stop when the ranks you actually weight have stabilised.
  • Should one propensity curve serve the whole product?
    No. Examination depends on layout, so a mobile single-column list, a desktop grid and a page with promoted slots above the results each need their own curve. Query type matters too — a navigational query is answered at slot 1 and the page is abandoned. Segment where you have the traffic to support it, and re-estimate after any redesign that moves where the eye lands.

To measure the eye-level shelf effect, you rotate the same products between shelves for a while. Comparing two different products on two shelves would tell you nothing.

saying these in an interview costs you the question

  • Reads propensities straight off click-through rate by rank
  • Randomises the whole page with no cost estimate
  • Compares different items across slots instead of the same item
  • Assumes one propensity curve holds across devices and layouts
  • Never re-estimates after a page redesign

context