skip to content

Why does pooling to build a relevance judgment set bias evaluation against a newly built retrieval system?

level: seniorimportance: should knowfreq 42%

answer

  1. Judgments cover only what somebody retrieved
  2. Missing label is not treated as neutral
  3. New system finds documents nobody judged
  4. The penalty scales with how different the system is
  5. Fix costs annotator hours, not code

basics

~10 s

Pooling judges only the documents that the contributing systems retrieved, and everything unjudged is scored as non-relevant. A new system that surfaces relevant documents no contributor ever returned is therefore penalised for finding them.

solid answer

~50 s

Nobody can judge a whole corpus, so judgment sets are built by **pooling**: run several retrieval systems over a query set, take each system's top-n, judge the union, and treat that union as the relevant set. Every metric then applies the closed-world assumption that an unjudged document is irrelevant. That is fine for the systems in the pool, whose good documents were all judged. It is unfair to a system built later — a new embedding retriever, say — that surfaces genuinely relevant documents nobody pooled: those score zero gain, they push judged documents down, and both its precision and its NDCG look worse than reality. Recall is hit hardest, since the denominator itself is a pool artefact. The mitigations are to judge the new system's top-n before comparing, to use pool-tolerant measures such as bpref or inferred AP, and to treat absolute numbers on an old collection as unusable for a genuinely novel retrieval method.

go deeper

for a junior

Know that relevance judgments cover only a sampled subset of documents and that anything unlabelled is scored as not relevant.

for a middle

Explain how pooling is constructed from multiple systems' top-n and why pool depth changes the recall denominator and the measured relevant set.

for a senior

Show you would budget judging effort to extend the pool before believing a comparison, and name a pool-tolerant measure such as bpref or a reported residual.

for a principal

Own judgment-set provenance and refresh policy as an asset, and set the standard that absolute offline numbers from an inherited collection are never the basis for a launch decision.

## Why pooling exists An evaluation collection needs to know which documents are relevant to each query. Judging every document for every query is impossible past a few thousand documents — a million-document corpus with 50 queries would need 50 million judgments. Pooling is the standard workaround, introduced by the classic TREC evaluations and still how most enterprise judgment sets are built: take a set of retrieval systems, run them all over the same queries, take the top n results from each (the *pool depth*), form the union, judge only that union, and declare everything else non-relevant. The union is far smaller than the corpus and heavily enriched in relevant documents, because independent systems tend to agree on the obvious answers and diverge on the marginal ones. For the pooled systems, this is a good approximation. ## The closed-world assumption Every metric built on the resulting judgments — precision@k, recall@k, MAP, NDCG — needs a label for every retrieved document. Unjudged documents get label 0. This is not a bug in any one implementation; it is the only thing the metric can do with a missing label, and it is applied silently. ## Where the bias bites Now evaluate a system that did not exist when the pool was built. Suppose it retrieves documents by dense vector similarity while every pooled system was lexical. It will surface documents that share no query terms with the query — paraphrases, synonyms, differently-worded answers. Some are excellent. None were pooled. Each one: 1. contributes zero gain to DCG and zero to the precision numerator; 2. occupies a high rank, pushing judged relevant documents down and shrinking their discounted contribution; 3. never enters the relevant-set denominator, so recall's denominator understates the truth. The new system is therefore punished twice for doing something right. The size of the effect scales with how *different* the system is from the pool: a small tweak to a pooled ranker is measured almost correctly, while a genuinely new retrieval paradigm can be understated badly. The symmetric failure exists too. If the pool was built mostly from systems similar to yours, your good documents are all judged and your competitor's are not, so you look better than you are. Pool composition is part of the experimental design, not an implementation detail. ## Other pooling artefacts **Pool depth.** Judging the top 10 of each system rather than the top 100 finds fewer relevant documents, so |R| is smaller, recall is inflated, and deep-ranking differences become invisible. Deeper pools cost linearly more judging. **Query set composition.** Pools are usually built on a query sample. If that sample over-represents head queries, the judgments describe head behaviour and any tail regression is unmeasured. **Judgment staleness.** Corpora change. Documents get added, edited, and deleted; a judgment made two years ago may describe a document that no longer says what it said. Judgment sets need maintenance, not just creation. **Annotator disagreement.** Two judges on a 4-point scale routinely disagree by a grade. Guidelines, calibration rounds, and measured agreement matter before you believe a small metric delta. ## What to do about it **Extend the pool.** The direct fix: before comparing a new system, judge its unjudged top-n on the evaluation queries. This restores fairness at a known judging cost and is what mature relevance teams budget for on every significant retrieval change. **Use pool-tolerant measures.** *bpref* scores only judged documents, ignoring the unjudged ones instead of counting them as non-relevant, and is far more stable when judgments are incomplete. *Inferred AP* estimates average precision from a sampled judgment pool. *Rank-biased precision* has an explicit residual that quantifies how much of the score is uncertain because of missing judgments. Reporting a residual is honest in a way a single number never is. **Compare relatively, not absolutely.** On an incomplete collection, "system A beats system B on this judgment set" is a defensible claim; "our NDCG@10 is 0.71" is close to meaningless outside that collection. **Sample from production traffic.** Judgment queries drawn from real logs, stratified across head, torso, and tail, describe your users. Queries inherited from a public benchmark describe someone else's. **Cross-check online.** Because offline judgments are always incomplete, a ranking change that wins offline must still be validated with an online experiment before it is believed. Pooling bias is one of the concrete reasons offline and online results disagree. ## What an interviewer wants That you know judgment sets are pooled rather than exhaustive, that unjudged means non-relevant to the metric, that the penalty falls hardest on systems unlike the pool, and that you would budget judging effort for the new system's own results rather than trusting an inherited qrels file.

  • How does bpref differ from average precision on an incomplete judgment set?
    Average precision counts every unjudged document as non-relevant, so it punishes a system for retrieving them. bpref ignores unjudged documents entirely and scores a system by how often judged relevant documents are ranked above judged non-relevant ones. It degrades far more gracefully as judgment coverage falls, which makes it the standard choice when the pool is known to be shallow.
  • How would you decide the pool depth when commissioning a new judgment set?
    Work backwards from the cutoffs you will report and the depth downstream stages consume. If NDCG@10 is the headline, a depth of 20 to 50 per system covers rank churn; if a first-stage retriever feeding a re-ranker must be measured at recall@200, the pool has to go at least that deep. Then check the marginal yield — when extra depth stops finding new relevant documents, stop paying for it.
  • Your team inherits a two-year-old judgment set for the same corpus. What do you check before using it?
    Whether the judged documents still exist and still contain the judged content, whether the query sample still reflects current traffic, what systems built the pool and how similar they are to yours, the pool depth, and the annotation guidelines and agreement rate. Then re-judge a random sample to estimate drift before trusting any absolute number.

It is like grading an exam with an answer key written from three students' papers: a fourth student who solves the problem a new way gets marked wrong for an answer the key never anticipated.

saying these in an interview costs you the question

  • Assuming an unjudged document is scored as neutral
  • Believing pooled judgments enumerate every relevant document
  • Treating absolute NDCG on an old collection as comparable anywhere
  • Comparing a new retrieval paradigm without extending the pool
  • Building the query set only from head traffic

context