Offline NDCG improved but online click-through fell after a ranking change. How would you diagnose that?
answer
- Two different populations, two different questions
- Fewer clicks is not always fewer answers
- Check what else shipped in the same window
- Ask whether the judged queries look like real traffic
- Slice the loss before explaining it
basics
~20 sCheck first whether the click drop is real harm or good abandonment, then look for the usual gaps: an unrepresentative judgment query sample, judges guessing intent differently from users, position and presentation bias in clicks, and non-relevance regressions such as latency.
solid answer
~50 sTreat it as two questions: is the offline win real, and is the online loss real? On the offline side, check whether the judged query set represents live traffic — head-heavy samples routinely hide tail regressions — whether the win came from a handful of queries, whether it is statistically significant on paired tests, and whether pooling bias flattered the new ranker. On the online side, remember click-through rate is a **proxy, not a truth**: it falls when users are satisfied without clicking, when snippets got better, and it rises when results are ambiguous enough to force exploratory clicking. Segment the drop by query class, position, device, and new-versus-returning users. Then check the non-relevance explanations: added latency, a UI change shipped alongside, index freshness, or an error rate on a query subset. Finally look at richer online signals — abandonment, query reformulation, time to first click, session success — before concluding either number is wrong.
go deeper
Know that offline relevance metrics and live click metrics measure different things and can legitimately disagree.
Explain position bias and good abandonment concretely, and name which additional online signals distinguish a satisfied non-click from a failed search.
Walk a structured diagnosis: validate the experiment, segment the loss, inspect real losing queries, then check latency, errors, and confounded deploys before blaming relevance.
Own the policy that online experiments decide launches while the offline suite is a cheap filter, and treat persistent disagreement as a defect in the suite to be fixed by reweighting and re-judging.
## The two worlds Offline evaluation scores rankings against human judgments on a fixed query set. Online evaluation observes real users on live traffic. They measure related but genuinely different things, and a disagreement between them is information, not necessarily a bug. The senior skill is a structured diagnosis rather than a reflex to trust one side. ## Is the click drop actually harm? Click-through rate is the most-used and least-trustworthy online metric. **Good abandonment.** If the change improved snippets, entity panels, or direct answers, users get what they need on the results page and do not click. CTR falls while satisfaction rises. Look for a matching fall in query reformulations and in sessions that end with a refined query — if reformulation also fell, the drop is likely benign. **Position bias.** Users click high results because they are high. Any metric built on raw clicks conflates presentation with relevance; comparing raw CTR across two different rankings is comparing two different position distributions. **Novelty and trust bias.** Returning users have learned where things sit. A reordering can depress clicks for a week purely because familiar items moved, then recover. Segment by new versus returning users and watch the trend, not the first day. **Metric definition.** CTR per query, per session, or per impression can move in opposite directions if the change altered the number of searches per session. Nail down the denominator before interpreting the delta. ## Is the offline win actually a win? **Query sample.** Judgment sets are usually built on a sample skewed toward head queries, because those are what someone thought to collect. If the change helps head queries and hurts the tail, offline says win and traffic says loss. Stratify the judged set across head, torso, and tail, and weight metrics by traffic if the headline is meant to predict production. **Judge-user intent mismatch.** Annotators infer intent from a query string with guidelines; users have context, history, and a task. For ambiguous queries, judges pick the interpretation the guidelines favour, which may be the minority intent in real traffic. **Pooling bias.** If the new ranker's results were freshly judged and the baseline's were not, the comparison is skewed. Fairness requires judging both systems' unjudged results. **Significance.** A mean NDCG that moved by a fraction of a percent on 50 queries is noise. Use a paired test across queries and report the win/loss/tie counts, which often reveal that a "win" is one query moving a lot rather than broad improvement. **Metric ceiling effects.** NDCG normalises against the judged set. Queries whose judgments contain two documents can score 1.0 while the ranking is poor, muting real differences. ## Non-relevance explanations A ranking change ships as code. Check whether it also changed: - **Latency.** Extra re-ranking or a second retrieval pass adds milliseconds; measurable engagement loss from added latency is a well-documented effect and is invisible to every offline relevance metric. - **Coverage and errors.** A subset of queries timing out, falling back, or returning zero results will crush CTR while the judged sample, which contains no such failures, shows nothing. - **Freshness.** If the new path uses a differently-refreshed index or a cached model, recent documents may be missing. - **Confounded deploys.** Another change shipped in the same window. This is the single most common cause and the reason experiments should isolate one change. ## The diagnostic sequence 1. Confirm the experiment: correct randomisation unit, no sample-ratio mismatch, sufficient exposure, one change under test. 2. Segment the online loss by query class, intent, device, locale, position, and user tenure. Losses are rarely uniform; find where it concentrates. 3. Pull the specific queries that lost the most clicks and inspect their result pages side by side. Human inspection of twenty real losers explains more than another week of aggregate dashboards. 4. Re-score the offline metric restricted to those losing queries. If offline also shows a loss there, the judgment set simply under-weighted them and the metrics agree after reweighting. 5. If offline still says win on queries that lost online, the judgments disagree with users — investigate the guidelines, and consider re-judging a sample with better intent context. 6. Check the non-relevance factors: latency percentiles, error rates, zero-result rate, index lag. 7. Look at deeper online metrics: abandonment, reformulation rate, time to first click, clicks below the fold, session-level success. If those improved while CTR fell, the change is probably good. ## The durable lesson Offline metrics exist to filter candidate changes cheaply and to give a reproducible regression signal; online experiments exist to decide launches. Neither replaces the other. A healthy team treats a persistent offline-online disagreement as a defect in the offline suite and fixes the suite — by reweighting queries to traffic, by refreshing judgments, by adding a segment that was missing — so that the next change's offline number is worth something. ## What an interviewer wants A structured diagnosis, explicit awareness that CTR can fall for good reasons, at least one offline-suite defect and one non-relevance cause, and the closing judgment that the online result governs the launch decision while the offline suite gets repaired.
- What is good abandonment, and how do you distinguish it from a real relevance regression?Good abandonment is a session ending with no click because the user got the answer from the result page itself — a snippet, a direct answer, an entity panel. Distinguish it by looking at what happened next: real regressions come with more query reformulations, more pagination, more searches per session, and shorter dwell on any click that does happen. If those all fell alongside CTR, the users were satisfied.
- Why is comparing raw click-through rate between two different rankings a biased comparison?Clicks depend heavily on position, so a document's click rate reflects where it was shown as much as whether it was relevant. Two rankings present a different position distribution, so their raw click rates are not measuring the same thing. Correct for it with a position-based propensity model, use interleaving so both rankings are exposed within one impression, or compare at fixed positions.
- Your offline suite keeps disagreeing with online results. What do you change about the suite?Rebuild the query sample from production logs stratified by traffic volume and intent class, and weight metrics by traffic so the headline predicts production. Refresh judgments and extend the pool to cover the systems under test. Add segments the suite was blind to — tail, locale, zero-result queries. Then backtest: score past launched and rolled-back changes offline and check the suite would have called them correctly.
saying these in an interview costs you the question
- Assuming lower click-through always means worse relevance
- Trusting offline metrics over a well-run online experiment
- Ignoring that judged query sets over-represent head traffic
- Not checking whether latency or another change shipped alongside
- Comparing raw clicks between rankings without correcting position bias