When would you use nDCG instead of MRR to score a RAG retriever's ranking?
answer
- one stops at the first hit
- the other grades and discounts
- binary versus graded relevance
- normalised by the ideal ordering
- graded labels cost more to produce
basics
~20 sMRR records only where the first relevant result landed, so it fits lookup queries with one right answer. nDCG scores the whole top-k list with graded relevance, so it fits queries where several results matter by different amounts.
solid answer
~40 sMRR is the mean of 1/rank of the *first* relevant hit. It is binary and it stops looking after that hit — everything below is invisible. nDCG grades each result (say 3 = fully answers, 1 = partially useful, 0 = irrelevant), discounts each gain by its position, and normalises by the ideal ordering so queries with different numbers of relevant documents are comparable. On an e-commerce query set that difference is the whole point: for "wireless earbuds", the in-stock current model, a discontinued variant and a charging case are not equally good, and only nDCG can reward putting the in-stock one first. The price is labelling: nDCG needs graded judgments, and with purely binary labels it tells you little that a rank-aware precision measure does not.
code
python · 15 linesimport math
def dcg(gains):
return sum(g / math.log2(i + 2) for i, g in enumerate(gains))
def ndcg_at_k(gains, k):
ideal = sorted(gains, reverse=True)
return dcg(gains[:k]) / dcg(ideal[:k])
# graded relevance of the results, in the order the retriever returned them
ranked_gains = [1, 3, 0, 2]
print(round(ndcg_at_k(ranked_gains, 4), 3)) # 0.788go deeper
Know that MRR looks at the rank of the first relevant result and nDCG scores the whole ranked list, and that both are reported at a stated k.
Explain the three parts of nDCG — graded gains, positional discount, normalisation by the ideal ranking — and why binary relevance makes it redundant.
Argue the choice from the query mix and the labelling budget you actually have, and say what each metric would hide about a retriever you are debugging.
Own the measurement strategy: which metrics get reported to whom, what grading rubric keeps labels consistent, and when ordering quality is worth the extra annotation spend at all.
## What each metric actually computes **MRR (mean reciprocal rank).** For each query, find the position of the first relevant result and score 1/position: rank 1 scores 1.0, rank 2 scores 0.5, rank 5 scores 0.2, nothing relevant in the list scores 0. Average across queries. Two properties follow directly: relevance is binary (a document either counts or does not), and the metric ignores everything after the first hit. **nDCG (normalised discounted cumulative gain).** Assign each retrieved document a *gain* from a graded scale — commonly 0-3. Discount each gain by its rank, conventionally dividing by log2(rank + 1), and sum to get DCG. Then divide by the DCG of the ideal ordering of the same judged documents (the IDCG), which yields a number in [0, 1] where 1 means "perfectly ranked". The three moving parts are therefore: graded gains, a positional discount, and normalisation. ## When MRR is the right tool MRR fits **known-item retrieval**: there is one correct document and the question is how fast the user or the generator reaches it. An internal support bot that maps "how do I reset my MFA device?" to exactly one runbook is a clean MRR case. It is cheap to label — you only need to mark the one right answer — it is easy to explain to stakeholders, and it is directly meaningful when the shipped k is 1 or 2. MRR breaks down the moment several documents are legitimately relevant with different value. It cannot distinguish a ranking that puts the best answer first and three good ones behind it from one that puts a barely-relevant document first and buries the best answer at rank 9 — both score by the position of *a* relevant hit only. ## When nDCG is the right tool nDCG fits queries where relevance is a spectrum and the whole visible list matters. Take an e-commerce catalogue and the query "wireless earbuds". A reasonable grading is: 3 for the current in-stock model that directly satisfies the intent, 2 for a sibling model in the same line, 1 for a compatible charging case, 0 for a discontinued variant nobody can buy. Under binary labels the in-stock model and the discontinued one look identical — both "relevant". nDCG's graded gains encode the difference, and the positional discount means that promoting the 3 above the 0 raises the score while burying it lowers it. The normalisation step matters more than people expect. Raw DCG grows with the number of relevant documents, so a query with eight good matches would dominate the average over a query with one. Dividing by the ideal DCG makes each query's score a fraction of the best achievable ranking *for that query*, which is what makes averaging across a heterogeneous query set defensible. ## How this maps onto RAG specifically The consumer of the ranking in RAG is the generator, not a human, and that shifts the emphasis slightly. - The prompt window is a hard cutoff, so metrics should be reported **at the shipped k** (nDCG@5, MRR@10) rather than over an unbounded list. - The positional discount is a reasonable proxy for the real effect that models weight earlier context differently and that a distractor at rank 1 does more damage than one at rank 5. - Graded relevance maps naturally onto "fully answers the question" versus "gives useful background" versus "same topic, wrong specifics" — a distinction that binary labels destroy and that determines whether an answer is right or subtly wrong. ## The cost side nDCG is not free. It requires graded judgments for every query-document pair you intend to count, which is several times the labelling effort of binary marks, and it requires the grading rubric to be consistent — if one annotator's 2 is another's 3, the metric moves for reasons that have nothing to do with the retriever. Graded scales also invite false precision: a 0-3 scale with a written rubric and worked examples is usually more reliable than a 0-10 scale. If your labels are binary, nDCG degenerates into a rank-aware version of precision and adds interpretive burden without adding information. In that situation MRR (for single-answer queries) or recall and precision at k (for multi-answer ones) communicate more clearly. ## A practical combination Most teams end up reporting more than one number, because each answers a different question: recall at a generous k for "can the retriever find it at all", nDCG at the shipped k for "is the ordering good", and MRR only where the corpus genuinely has one right answer per query. Reporting a single headline metric hides which of those is failing.
- What is the point of the normalisation step in nDCG?Raw DCG grows with the number of relevant documents a query has, so a query with eight good matches would dominate the average over a query with one. Dividing by the ideal DCG — the score of the best possible ordering of the judged documents for that query — turns each query's score into a fraction of what was achievable, in [0, 1]. That is what makes averaging across a heterogeneous query set meaningful.
- Why does nDCG discount by the logarithm of the rank rather than linearly?A log discount falls quickly across the first few positions and then flattens, which matches how attention actually works: the difference between rank 1 and rank 2 matters far more than the difference between rank 18 and 19. A linear discount would treat all those gaps as equal and over-reward deep-list improvements that nobody, human or generator, ever sees.
- Your labels are binary. Is nDCG still worth computing?Mostly not. With gains of 0 and 1 nDCG collapses into a rank-aware precision measure, so it adds interpretive overhead without adding information, and stakeholders read the number as more sophisticated than it is. Report MRR if queries have one right answer, or recall and precision at the shipped k if they have several, and invest the saved effort in graded labels if ordering quality genuinely drives your product.
saying these in an interview costs you the question
- Says nDCG and MRR are interchangeable ranking metrics
- Computes nDCG from binary labels and calls it graded
- Reports nDCG without stating the k
- Forgets that MRR ignores everything after the first hit
- Skips normalisation and compares raw DCG across queries