skip to content

How would you bootstrap and maintain a labelled query-document set for retrieval evaluation?

level: principalimportance: should knowfreq 42%

answer

  1. synthetic first, audited second
  2. one query per chunk is a bootstrap
  3. the query mirrors the chunk's wording
  4. single-chunk questions only
  5. chunk ids churn on re-chunking

basics

~20 s

Bootstrap with an LLM generating one query per chunk, then hand-audit a sample and expect to discard a large share. Blend in real logged queries for realistic distribution, and re-audit whenever the corpus or chunker changes.

solid answer

~50 s

I bootstrap and then correct the bootstrap's biases. Generating one synthetic query per chunk gives an instant labelled pair — the source chunk is the positive — and on a bank's compliance-procedure corpus that yields thousands of pairs in an afternoon. Then I hand-audit around 100 pairs and typically throw out roughly 30%: queries that are vague, unanswerable from the chunk, or duplicates. Two biases survive the audit and must be attacked deliberately. First, **lexical leakage** — the generated query reuses the chunk's wording, so retrieval looks better than it will in production; I paraphrase into user vocabulary and seed the set with real logged queries. Second, **single-chunk bias** — every synthetic query is answerable from one chunk, so multi-hop questions never appear and recall looks flattering. On maintenance: labels are pinned to chunk ids that churn on every re-chunk, so I anchor to document plus span and re-audit after corpus or chunker changes.

go deeper

for a junior

Know that retrieval metrics need query-document relevance labels, and that generating candidate queries from chunks with a model is the usual way to start.

for a middle

Explain why synthetic pairs flatter the retriever — leaked wording, single-chunk questions, uniform distribution — and what a hand audit of a sample actually corrects.

for a senior

Show the maintenance plan: labels anchored to stable spans rather than chunk ids, pooled judgments to limit incomplete-label bias, and re-auditing after corpus or chunker changes.

for a principal

Own the tradeoff between annotation spend and decision quality — set size versus realism, frozen versus refreshed slices, and how much metric movement you will treat as signal given the set's known defect rate.

## Why you need one at all Classical retrieval metrics — recall@k, precision@k, nDCG — all require knowing which documents are relevant to which queries. That ground-truth set is the expensive part of retrieval evaluation, and the quality of every number you report is bounded by it. Building it badly does not produce noisy metrics; it produces confidently wrong ones. ## Step 1: bootstrap synthetically The standard cold-start move is to have a model read each chunk and write a question that the chunk answers. The pair is labelled for free: the source chunk is a known positive. On a bank's internal compliance-procedure corpus this turns a few thousand chunks into a few thousand query-document pairs in an afternoon, which is enough to start comparing retrievers. It is a bootstrap, not a ground truth. Three biases come baked in. **Lexical leakage.** The generated question tends to reuse the chunk's distinctive vocabulary, because the model is looking straight at it. That hands both lexical and dense retrievers an unrealistically easy match and inflates every metric. Real users ask "do I need a second sign-off to release a payment?" while the document says "dual authorisation thresholds for outbound settlement". **Single-chunk bias.** Each query is answerable from exactly one chunk by construction, so genuinely multi-hop questions — the ones that need a policy clause plus its exception schedule — never enter the set. Recall looks flattering precisely on the queries that matter most. **Distribution mismatch.** Chunk-derived queries are uniform over the corpus. Real traffic is heavily skewed: a small set of topics carries most of the volume, and rare documents are almost never asked about. ## Step 2: audit a sample by hand Audit a sample — a hundred pairs is a workable first pass — and expect to discard a meaningful share; roughly 30% is a realistic yield loss. Typical discards: questions too vague to have a single right answer, questions the source chunk does not actually answer, near-duplicates of each other, and questions that are really about a heading rather than content. The audit does two jobs. It cleans the sample, and it *estimates the defect rate of the whole set*, which tells you how much to discount the metrics computed on the unaudited remainder. If 30% of pairs are junk, a 4-point difference between two retrievers is not a result. While auditing, also **paraphrase away the leaked vocabulary**: rewrite the query in the language a user would use. This is the single highest-value edit, because it converts an easy lookup into a realistic one. ## Step 3: mix in real queries Synthetic pairs fix coverage; real queries fix distribution. Take actual questions users asked, label their relevant chunks by hand or by pooling the results of several retrievers and judging the union. Real queries also surface things generation never invents: typos, acronyms, half-sentences, and questions the corpus cannot answer at all — the last category is valuable, since "correctly retrieving nothing" is a behaviour worth measuring. ## Step 4: handle incomplete judgments A hand-built set is almost never exhaustive. Chunks that are genuinely relevant but never judged get counted as false positives, so **precision is systematically understated** and a retriever that finds good material your annotators missed is punished for it. Mitigations: pool the top results from every retriever you compare and judge the union rather than one system's output; and be explicit that a chunk is *unjudged* rather than silently treating it as irrelevant. Recall carries a different bias — it is measured only against the relevant set you declared, so it over-reports whenever that set is incomplete. ## Step 5: plan for the set going stale This is the part teams skip, and it is why old evaluation sets quietly stop meaning anything. - **Chunk ids churn.** Change the chunk size, the splitter or the overlap and every label pinned to a chunk id is orphaned. Anchor labels to a stable document identifier plus a character span or a quoted passage, so relevance can be re-mapped onto new chunks mechanically. - **The corpus moves.** Documents are revised, superseded and deleted. A label pointing at a clause that no longer exists silently penalises the retriever for doing the right thing. - **The query mix moves.** New products and new user behaviour shift what people ask; a set frozen a year ago measures last year's system. A workable structure is two slices: a small **frozen regression slice** that never changes, so results stay comparable over time, and a **refreshed slice** re-sampled from recent traffic each cycle, so you see current reality. Report them separately — averaging them hides exactly the drift you built the second slice to catch. ## What to say about scale Size is a budget question, not a virtue. A few hundred well-audited, realistically-phrased queries will discriminate between retrievers more reliably than ten thousand leaky synthetic pairs, because the leaky set compresses every system toward the top of the scale. Spend the annotation budget on realism and on covering your hard slices — multi-hop questions, rare-but-critical topics, unanswerable queries — rather than on volume.

  • What is lexical leakage in a synthetic evaluation set, and how do you reduce it?
    When a model writes a question while looking at a chunk, it reuses that chunk's distinctive vocabulary, so both lexical and dense retrievers match far too easily and every metric is inflated. Reduce it by paraphrasing audited queries into the vocabulary users actually use, and by seeding the set with real logged questions, which carry the genuine vocabulary gap between how people ask and how documents are written.
  • Why do multi-hop questions rarely appear in a chunk-generated set, and why does that matter?
    Because the generation procedure shows the model one chunk and asks for a question it answers, so every pair is single-hop by construction. That matters because multi-hop queries — a clause plus its exception schedule — are where retrieval actually fails, and their absence makes recall look better than production behaviour. Author those cases deliberately, from real user questions, and track them as a separate slice.
  • Public retrieval benchmarks exist. Why not just use one instead of building your own set?
    Because retriever performance transfers poorly across domains: zero-shot comparisons across heterogeneous collections routinely show a model that leads on one domain falling well behind on another. A public benchmark tells you which retrievers are plausible candidates; it cannot tell you which one wins on your corpus, your vocabulary and your query mix. Use benchmarks to shortlist, your own set to decide.
  • How do you keep an evaluation set comparable over time while still tracking current traffic?
    Split it. Keep a small frozen regression slice that never changes, so a number from six months ago still means something, and a refreshed slice re-sampled from recent queries each cycle, so you see the system users actually face. Report them separately rather than averaged — the divergence between the two slices is the drift signal, and averaging destroys it.

saying these in an interview costs you the question

  • Treats LLM-generated query-chunk pairs as verified ground truth
  • Ignores that generated queries copy the chunk's wording
  • Pins relevance labels to chunk ids that churn
  • Assumes unjudged documents are irrelevant
  • Prefers a huge unaudited set to a small audited one

context