skip to content

In RAG retrieval evaluation, how do recall@k and precision@k differ, and which one caps answer quality?

level: middleimportance: must knowfreq 75%

answer

  1. one is a ceiling, one is noise
  2. the denominators differ
  3. what is missed cannot be cited
  4. distractors cost tokens and attention
  5. always report the k

basics

~20 s

Recall@k is the share of the relevant chunks that appear in the top k results; precision@k is the share of those k results that are relevant. Recall sets the ceiling — what retrieval misses, the generator can never cite.

solid answer

~40 s

Both are computed over the same top-k list but with different denominators. **Recall@k** divides the relevant chunks found by the total number of relevant chunks that exist for that query; **precision@k** divides them by k. In a RAG pipeline recall is the hard ceiling: if the chunk holding the deductible rule for an insurance claim never enters the prompt, no amount of generator quality recovers it, and the model either says nothing or invents something. Precision is a softer cost — irrelevant chunks burn tokens, dilute attention and give a plausible-sounding wrong passage for the model to anchor on. So I optimise recall first, then buy precision back with reranking or a smaller final k. Neither number means anything without the k it was measured at, so always report `recall@5`, never "recall 0.61".

code

python · 15 lines
python
def recall_at_k(retrieved, relevant, k):
    hits = len(set(retrieved[:k]) & set(relevant))
    return hits / len(relevant)


def precision_at_k(retrieved, relevant, k):
    hits = len(set(retrieved[:k]) & set(relevant))
    return hits / k


retrieved = ["c9", "c3", "c7", "c1", "c4"]
relevant = ["c3", "c1", "c8"]

print(round(recall_at_k(retrieved, relevant, 5), 3))     # 0.667
print(round(precision_at_k(retrieved, relevant, 5), 3))  # 0.4

go deeper

for a junior

Be able to state both definitions cleanly: recall@k is the fraction of relevant chunks you found, precision@k is the fraction of returned chunks that were relevant, and always name the k.

for a middle

Explain why the denominators differ, why recall rises and precision falls as k grows, and why retrieval recall caps what any generator can produce downstream.

for a senior

Show the diagnostic move: read recall at two different k values to separate a retrieval-input problem from a ranking problem, and quantify what low precision costs in tokens and distraction.

for a principal

Own the framing that recall is a ceiling and precision is a budget, and defend how much labelling effort a trustworthy retrieval number is worth against the alternative of judging answers end to end.

## The two definitions, side by side Both metrics look at the same object: the ordered list of chunks a retriever returned for one query, truncated at position k. What differs is the denominator. - **recall@k** = (relevant chunks in the top k) / (all relevant chunks that exist for this query). It answers *how much of what I needed did I get?* - **precision@k** = (relevant chunks in the top k) / k. It answers *how much of what I got did I need?* Both are per-query numbers, averaged over an evaluation set. A single query with six relevant chunks, of which three appear in a top-10 list, scores recall@10 = 0.5 and precision@10 = 0.3. The two move against each other as k grows. Increasing k can only add relevant chunks, so recall@k is monotonically non-decreasing in k; precision@k usually falls, because the tail of the ranking is mostly noise. This is why quoting either number without its k is meaningless. ## Why recall is the ceiling in RAG RAG has a strict information bottleneck: the generator sees only the chunks retrieval placed in the prompt. Everything downstream — reranking, prompt design, a stronger model, self-critique — can only reorganise, ignore or summarise that set. It cannot conjure a fact that was never retrieved. Concretely, imagine an insurance claims-handling knowledge base. An adjuster asks whether a windscreen replacement is subject to the standard deductible. The rule lives in one clause of the policy-exceptions document. If that clause is not in the top k, the best possible outcomes are a refusal or a confident answer derived from the general deductible chunk — which is wrong. Retrieval recall has capped the system at *wrong or silent*, and no generation-side metric will explain why. That is precisely why RAG evaluation is split by stage: a low retrieval recall localises the bug before anyone starts blaming the model. ## Why precision still matters It is tempting to conclude that you should just make k huge and stop worrying. Three costs push back. 1. **Token and latency budget.** Context is paid for on every request, and long prompts raise time-to-first-token. 2. **Distraction.** Models are demonstrably pulled off course by passages that are topically close but factually wrong — a superseded policy version, a chunk about a different product line. A plausible distractor is more dangerous than an obviously irrelevant one. 3. **Attention dilution.** As the retrieved set grows, the signal-to-noise ratio of the prompt falls, and the relevant sentence competes with more surrounding text. So the practical target is: get recall high at a generous k, then raise precision *at the k you actually ship* by reordering, not by retrieving less. ## Multi-chunk questions and the hit-rate trap **Hit rate** (sometimes called success@k) scores 1 if *any* relevant chunk is in the top k, and 0 otherwise. It is the crudest of the family and it is fine when every question is answerable from exactly one chunk. Most real corpora are not like that. A question such as "is a windscreen claim excluded, and what is the notification window?" needs two clauses. Hit rate awards a perfect 1.0 for retrieving one of them, hiding a systematic failure. Per-query recall@k, with the full required set as the denominator, exposes it: that query scores 0.5 and the average drops. Whenever a corpus supports multi-hop or composite questions, prefer recall@k over hit rate and make sure the ground-truth set lists *every* chunk a correct answer needs. ## Reading them together A useful diagnostic pattern: - **Recall@k low at every k** — the failure is upstream of ranking: chunking that splits the answer, an embedding model that does not understand the domain vocabulary, or queries whose phrasing never matches the documents. Fix retrieval inputs, not the ranker. - **Recall@k high at a large k but low at a small k** — the right chunk is in the candidate pool but ranked badly. This is a ranking problem, and a reranker is the cheap fix. - **Precision@k low while recall@k is fine** — you are paying tokens for noise; shrink the shipped k or rerank. One last caution: both metrics assume your relevance labels are complete. If a genuinely relevant chunk is unlabelled, it is counted as a false positive and precision is understated. Incomplete judgments make precision the less trustworthy of the two numbers on a hand-built set.

  • If the generator can simply ignore irrelevant chunks, why track precision@k at all?
    Because ignoring is not free and not reliable. Every extra chunk is billed tokens and adds latency, and models do get anchored by passages that are topically close but factually wrong — a superseded policy clause reads exactly like the current one. Low precision also dilutes attention across a longer prompt, so the relevant sentence competes with more text. Precision is the metric that tells you what you are paying for that noise.
  • A question needs two chunks to answer. Why is hit rate the wrong metric there?
    Hit rate scores 1 if any relevant chunk appears in the top k, so retrieving one of the two required clauses looks like a perfect result while the answer is still unanswerable. Per-query recall@k, with both required chunks in the denominator, scores that query 0.5 and surfaces the gap. Use hit rate only when every question in the set is genuinely single-chunk.
  • How does incomplete labelling bias these two metrics differently?
    Unlabelled-but-relevant chunks are counted as misses in the numerator of precision, so precision@k is systematically understated on a hand-built set — you punish the retriever for finding something your annotators never judged. Recall is biased differently: it is measured against the relevant set you declared, so it silently over-reports if that set is incomplete. Treat precision as the noisier number and audit judgments before acting on it.

saying these in an interview costs you the question

  • Quotes a recall number without saying at what k
  • Claims high precision@k proves retrieval is healthy
  • Thinks a better generator can recover chunks retrieval missed
  • Uses hit rate on questions that need several chunks
  • Assumes every query has exactly one relevant chunk

context