skip to content

In search evaluation, how do MAP and MRR differ, and when is each the right metric?

level: middleimportance: should knowfreq 56%

answer

  1. One metric stops looking after the first hit
  2. The other scores at every relevant position
  3. Reciprocal of a rank, averaged
  4. Its denominator is the count of relevant documents
  5. Known-item search versus exploratory search

basics

~20 s

Mean reciprocal rank averages 1 divided by the position of the first relevant result, so it only sees one document per query. Mean average precision averages precision measured at every relevant position, rewarding a ranking that surfaces all relevant documents high.

solid answer

~50 s

**MRR** takes the rank of the first relevant result, inverts it, and averages over queries: first hit at rank 1 scores 1.0, rank 2 scores 0.5, rank 4 scores 0.25. Everything after the first relevant document is invisible to it, which makes it the right metric for **known-item and navigational** queries where exactly one answer is correct — a login page, a specific product, a question with one canonical document. **MAP** is built from average precision: walk the ranking, record precision@i at each position holding a relevant document, sum those, and divide by the number of relevant documents; then average across queries. That rewards packing *all* the relevant documents near the top, so it fits **recall-oriented, multi-answer** queries. Both assume binary judgments — if your labels are graded, NDCG is the better fit. MRR is also brutally noisy per query, since a single position change can halve the score.

code

python · 5 lines
python
ranks = [1, 3, 6]      # positions of relevant results
R = 5                  # total relevant documents for this query

rr = 1 / ranks[0]                                       # 0.333
ap = sum((i + 1) / r for i, r in enumerate(ranks)) / R  # (1/1 + 2/3 + 3/6) / 5 = 0.433

go deeper

for a junior

Recall that MRR uses only the position of the first relevant result and that a hit at rank 1 scores 1.0, rank 2 scores 0.5.

for a middle

Be able to compute average precision on a short worked list, get the denominator right, and justify picking one metric over the other from the query workload.

for a senior

Demonstrate that you know both flatten graded judgments, that MRR is noisy on small query sets, and that excluded zero-score queries silently flatter a report.

for a principal

Decide which metric the organisation headlines for which surface, and resist a single company-wide number that hides that navigational and exploratory search need different measures.

## Two different questions Rank-aware metrics differ mainly in what they consider a success. MRR asks "how fast did the user reach *an* answer?". MAP asks "how well did the system rank *all* the answers?". Choosing between them is really choosing which of those user tasks your search serves. ## Mean reciprocal rank For one query, the reciprocal rank is `1 / rank_of_first_relevant_result`, or 0 if no relevant result appears within the cutoff. MRR is the mean of that over a query set. The score drops steeply: 1, 0.5, 0.33, 0.25, 0.2. Moving the answer from rank 3 to rank 1 gains 0.67; moving it from rank 9 to rank 7 gains about 0.03. That shape encodes the assumption that the user stops at the first useful result, which is exactly right for navigational intent and largely wrong for exploratory intent. Because each query contributes one of a small set of discrete values, MRR is noisy. A query set of 50 known-item queries can swing meaningfully when three of them shift one position. Report it over hundreds of queries, and prefer a paired test over comparing two means. ## Average precision and MAP Average precision uses the whole list. Walk down the ranking; each time you hit a relevant document at position i, record precision@i. Sum those precisions and divide by the total number of relevant documents R: `AP = (Σ over relevant positions i of precision@i) / R` Example: R = 5, relevant documents at ranks 1, 3, and 6. Precision@1 = 1/1 = 1.0, precision@3 = 2/3 ≈ 0.667, precision@6 = 3/6 = 0.5. Sum ≈ 2.167, divided by 5 gives AP ≈ 0.43. The two relevant documents never retrieved contribute nothing, which is how AP folds recall into the score — a system that finds only three of five relevant documents cannot score above 0.6 no matter how it orders them. MAP is the mean of AP across queries. It is the workhorse of classic IR benchmarking because it is a single number that is sensitive to both ordering and coverage. ## Choosing between them - **Navigational / known-item**: MRR. There is one right document; how far down it sits is the whole story. Question answering over a corpus with one gold passage is the same shape. - **Informational, several relevant documents**: MAP. A researcher, a shopper comparing options, or a support agent scanning candidate articles benefits from every relevant document being high. - **Graded relevance available**: NDCG instead of either. MAP and MRR both flatten grades to relevant/not, discarding the distinction between a perfect and a marginal match. - **Set-oriented downstream consumption**: recall at the candidate depth, since ordering inside the candidate set barely matters when a re-ranker will redo it. ## Common mistakes **Dividing AP by the number of retrieved relevant documents instead of R.** That inflates every score and destroys the recall sensitivity that makes MAP worth using. With relevant hits at ranks 1, 3, 6 the correct denominator in the example is 5, not 3. **Truncation.** In practice AP is computed at a cutoff (AP@k). Decide whether the denominator is R or min(R, k) and be consistent, because the two conventions give different numbers for queries with many relevant documents. **Reading MRR as "average rank".** It is the average of a reciprocal, not the reciprocal of an average, and the two differ substantially. A system whose answers sit at ranks 1 and 9 has MRR ≈ 0.56, not 1/5 = 0.2. **Using MRR when the query has many valid answers.** It will declare victory as soon as any relevant document reaches rank 1, even if the other nine relevant documents are on page four. That is a genuinely misleading result for an exploratory workload. **Ignoring the zero case.** Queries where no relevant result appears within the cutoff contribute 0 and drag the mean down; excluding them silently inflates the metric and hides your worst failures. Keep them and report the coverage rate separately. ## Relationship to the rest of the metric family MRR is a special case of the family that scores only the first success; MAP generalises precision@k across all cutoffs where a relevant document appears; NDCG generalises further by admitting grades and a tunable discount. All three are order-sensitive, unlike raw precision@k, and all three inherit whatever bias sits in the judgment set. None of them tells you anything about latency, snippet quality, or whether the user was actually satisfied — that is what online metrics are for. ## What an interviewer wants Correct formulas, an honest statement that MRR looks only at the first hit, the ability to compute AP on a small worked example without dividing by the wrong denominator, and a workload-driven justification for choosing one.

  • Why does average precision divide by the total number of relevant documents rather than by the number of relevant documents retrieved?
    Dividing by the total makes the metric recall-sensitive: relevant documents never retrieved contribute zero to the numerator but still inflate the denominator, capping the score. Dividing by the retrieved count instead would let a system that finds one relevant document at rank 1 and misses nine others score a perfect 1.0, which is exactly the failure the metric exists to catch.
  • You have graded judgments on a five-point scale. Why would you still report MAP alongside NDCG?
    MAP is easier to interpret and compare against published baselines, and it stresses coverage in a way NDCG does not, since NDCG normalises against only the judged set. Collapsing grades to binary at a chosen threshold and reporting MAP gives a second, differently-biased view; if the two metrics disagree about a change, that disagreement is itself informative.
  • Why is MRR considered a noisy metric on small query sets?
    Each query contributes one value from the discrete set 1, 0.5, 0.33, 0.25..., and the gaps between them are large at the top. A handful of queries whose answer moves from rank 2 to rank 1 can shift the mean noticeably, so differences on a few dozen queries are usually noise. Use hundreds of queries and a paired significance test.

saying these in an interview costs you the question

  • Saying MRR accounts for all relevant results in the list
  • Computing average precision by dividing by the number of hits found
  • Treating MRR as the reciprocal of the average rank
  • Using MRR for exploratory queries with many valid answers
  • Dropping zero-scoring queries from the mean

context