How does the examination hypothesis explain position bias in a ranked results page's click log?
answer
- clicks confound looking with liking
- two events must both happen
- one factor depends only on rank
- the other factor is the relevance you want
- slot 1 versus slot 5, equal quality
basics
~20 sThe examination hypothesis says a logged click happens only when the user both examined that slot and found the item relevant. Since examination falls steeply with rank, a top slot's higher click rate reflects position, not quality.
solid answer
~50 sPosition bias is the fact that an item shown higher in a list collects more clicks than an identical item shown lower, purely because of where it sat. The examination hypothesis models a click as the conjunction of two events: the user looked at the slot, and the item was relevant to the query. That gives the factorisation `P(click | query, item, rank) = P(examine | rank) * P(relevant | query, item)`. On a classifieds marketplace results page you routinely see slot 1 pull roughly six times the clicks of slot 5 for listings a human rater grades as equally good; under the factorisation that ratio is an examination ratio, not a relevance ratio. The practical consequence is that raw click-through rate by rank is not a quality signal, and a model trained on raw clicks mostly learns to reproduce the ordering that produced the log.
go deeper
Be ready to say in one sentence that items higher on the page get clicked more regardless of quality, so click counts by slot are not a fair comparison between items.
You are expected to write the factorisation out loud: a click needs examination of the slot and relevance of the item, so click probability is the examination probability of the rank times the relevance probability. Say why that makes deep non-clicks uninformative.
Show you know the assumption is an approximation. Name attractiveness effects, trust bias and device layout as places where examination stops depending on rank alone, and say how you would segment propensity curves in production.
Own the framing that the log is a missing-not-at-random sample created by your own ranker, and argue for what the organisation must log — slot, surface, and the propensity in force — before anyone can debias anything later.
## What a click log actually contains Every row of a ranked-results click log is an impression: a query or a user context, an item, the slot it was shown in, and whether it was clicked. What you want is relevance — would this user have wanted this item. What you have is a click, and a click is the outcome of a process with a step you never observe: whether the user's eyes ever reached that slot. **Position bias** is the systematic gap between the two. Hold quality constant and move an item down the page, and its click rate falls. That is not preference; it is geometry and attention. ## The examination hypothesis The examination hypothesis is the standard generative story that makes the problem tractable. It says a click occurs if and only if two independent events both occur: 1. the user **examined** the slot the item was shown in, and 2. the item was **relevant** to that user for that query. Written out for one impression: ``` P(click | query, item, rank) = P(examine | rank) * P(relevant | query, item) ``` The first factor is often called the **examination propensity** of the rank, written `p_r`. The second is the thing you actually want to learn. The strong part of the assumption is that examination depends on the **rank alone** — not on the item, not on the query. A model built on exactly this factorisation is usually called a position-based click model. ## Why factorising is the whole point Once clicks factor into a rank term times a relevance term, position bias stops being a vague complaint about logs and becomes a nuisance parameter you can measure and divide out. Two consequences follow immediately. **First, click rates are only comparable within a slot.** Two listings on a classifieds marketplace results page, both graded equally good by a human, will show wildly different click rates if one habitually lands at slot 1 and the other at slot 5 — a factor of roughly six is unremarkable on real pages. Any leaderboard built on raw click-through rate is really a leaderboard of where the current ranker happened to put things. **Second, the bias does not average out with more data.** This is the single most common misunderstanding. More impressions shrink the variance of the click rate; they do not touch the multiplicative `p_r` factor sitting in front of relevance. A biased estimator converges to the wrong number, and it converges to it more confidently. **Third, a non-click is not a negative label.** Under the hypothesis, a non-click at rank 9 has two explanations that the log cannot distinguish: the user saw it and did not want it, or the user never got that far. At a slot with low examination probability, the second explanation dominates, so treating deep non-clicks as hard negatives systematically slanders whatever the ranker buried. ## Where the rank-only assumption breaks The hypothesis is an approximation, and knowing its cracks is what separates a middle answer from a senior one. - **Attractiveness or presentation effects.** A large photo, a bold price badge, or a familiar brand raises the chance a slot is examined at all. Then examination is not a function of rank alone. - **Trust bias.** Users partly infer quality from position: an item at the top is judged relevant more readily than the same item lower down. That contaminates the *relevance* factor, not just the examination factor, and plain rank-only correction does not remove it. - **Layout and device.** A single-column mobile list, a grid, an infinite scroll, and a page with promoted slots above the results all decay differently. One propensity curve for all surfaces is wrong. - **Context within the session.** A user who has already found what they wanted stops examining anything; a user on their fifth query scans further. ## What the hypothesis buys you It turns an unidentifiable problem into an estimation problem in two steps. Estimate `p_r` for each rank, then correct the observed clicks for it — the correction being to weight each click by the reciprocal of its slot's examination propensity, so evidence gathered from rarely-examined slots counts for more. Both steps have their own difficulties, but neither is available at all until you have committed to a model of how a click is generated. That commitment is what the examination hypothesis is. If you take one thing away: a click log is a **missing-not-at-random** sample of user preferences, and the mechanism that decided what is missing is the ranker you are trying to improve.
- Where does the rank-only part of the examination hypothesis break down?Whenever something other than position changes whether a slot is looked at or how it is judged. A large photo or a trusted brand raises examination of its own slot; trust bias means users read position itself as a quality cue, which contaminates the relevance factor rather than the examination factor; and mobile lists, grids and pages with promoted slots above the results each have their own decay curve. In practice you estimate separate propensity curves per surface.
- If clicks are position-biased, why not train on purchases or long dwell instead?Because they are gated by exactly the same unobserved step: nobody buys what they never saw. Deeper engagement signals inherit the identical examination factor, so they need the same correction. They are also far sparser and arrive later, so the corrected estimates are noisier and staler. They are better targets, not unbiased ones.
- Does collecting a year of logs instead of a week fix position bias?No. Extra data reduces variance, not bias. The examination propensity multiplies relevance in every single impression, so the click rate converges to relevance times that factor no matter how many rows you add. Worse, a longer log usually reflects a ranker that has been reinforcing its own ordering for longer, so the exposure is even less even.
Products at eye level in a supermarket outsell identical products on the bottom shelf. Sales by shelf tell you about shelves, not about products.
saying these in an interview costs you the question
- Claims a higher click rate at rank 1 proves the item is more relevant
- Says enough logged data eventually averages position bias away
- Treats a non-click at a deep slot as a confirmed negative label
- Cannot state the two events a click requires under the hypothesis
- Confuses missing-not-at-random exposure with random label noise