skip to content

Why does reciprocal rank fusion combine ranks rather than raw BM25 and cosine scores?

level: middleimportance: must knowfreq 62%

answer

  1. two scores, no shared axis
  2. normalization is estimated per query
  3. best of a bad list becomes 1.0
  4. positions survive any monotone rescaling
  5. 1/(k+rank), summed across lists

basics

~20 s

BM25 scores are unbounded and corpus-dependent while cosine similarity sits in a fixed range, so they cannot share a scale. Reciprocal rank fusion sums 1/(k + rank) over the lists, commonly with k=60, using only positions — no calibration needed.

solid answer

~50 s

Score fusion requires putting two incomparable quantities on one axis. BM25 produces an unbounded score whose magnitude depends on term rarity and document length in *this* corpus; cosine similarity is bounded and its typical values depend entirely on the embedding model. The usual fix, min-max normalizing each list per query, quietly breaks: it maps the best hit of every query to 1.0, so a query where nothing matched lexically gets its garbage top hit promoted to a perfect lexical score. Reciprocal rank fusion sidesteps calibration entirely — each document scores `sum over lists of 1/(k + rank)`, with k conventionally 60, and the fused list is that sum re-sorted. It is scale-free, needs no per-query statistics, and rewards agreement between the legs: with k=60 a document ranked 20 in both lists (2/80 = 0.025) beats one ranked 1 in only one list (1/61 ≈ 0.016). The cost is real — RRF throws away magnitude, so it cannot tell a runaway top match from a marginal one.

code

python · 10 lines
python
def rrf(rankings, k=60):
    scores = {}
    for ranking in rankings:
        for rank, doc_id in enumerate(ranking, start=1):
            scores[doc_id] = scores.get(doc_id, 0.0) + 1.0 / (k + rank)
    return sorted(scores, key=scores.get, reverse=True)

dense = ["d1", "d2", "d3"]
bm25 = ["d3", "d1", "d9"]
print(rrf([dense, bm25]))  # ['d1', 'd3', 'd2', 'd9']

go deeper

for a junior

Know the formula shape — each document scores 1/(k + rank) summed over the lists it appears in — and be able to say that BM25 and cosine scores are not on the same scale.

for a middle

Explain concretely why per-query min-max normalization misleads: it forces the top hit of a hopeless query to 1.0. Then show that rank fusion needs no calibration and rewards cross-leg agreement.

for a senior

Demonstrate the tradeoff you accepted: RRF discards magnitude, so name what you lost — thresholding, abstention, user-facing confidence — and how you preserved those signals alongside the fused order.

for a principal

Own fusion as a platform decision: one scale-free default that survives embedding-model swaps and corpus changes without retuning, versus per-team score calibration that quietly rots.

## The problem: two scores that do not live on the same axis A hybrid retriever returns two ranked lists for the same query. To produce one answer, you must merge them. The naive move is to add the scores — but the two numbers mean different things. - **BM25** returns an unbounded, corpus-relative score. Its magnitude depends on the inverse document frequency of the query terms, the term frequencies in the document, and the document's length relative to the corpus average. A score of 31 is not "good" in any absolute sense; it is only meaningful next to the other scores for the same query on the same index. - **Cosine similarity** between embeddings is bounded (typically reported in [0, 1] for normalized vectors on models trained that way), but its distribution is a property of the embedding model. Some models cluster nearly everything above 0.7; others spread scores widely. There is no calibration guaranteeing that 0.8 means "relevant". Adding these directly means an arbitrary implicit weighting that changes with the corpus and the model. ## Why per-query normalization breaks The common patch is min-max normalization: for each list, map its top score to 1 and its bottom score to 0. This is done *per query*, over the returned window, and that is precisely the flaw. Normalization is computed from the results themselves, so it destroys the information about whether the query matched anything at all. Consider a manufacturing support index. Query A contains a rare fault code and BM25 returns scores 31.2, 12.0, 9.4 — a decisive top match. Query B is a vague prose question with no distinctive terms and BM25 returns 2.1, 1.9, 1.8 — noise. After min-max, both top hits are 1.0 and both bottom hits are 0.0. The fusion step now believes the lexical leg is equally confident on both queries, and the noise gets weighted as heavily as the signal. Z-score normalization softens this but shares the disease: it is estimated from a small, truncated, query-specific sample. Other score-fusion pitfalls: the statistics shift when you change per-leg top-k (a different window changes the min and max), and they shift again whenever you swap embedding models or reindex. ## What RRF does instead Reciprocal rank fusion discards magnitudes and keeps only positions: `RRFscore(d) = sum over each list L of 1 / (k + rank_L(d))` where `rank_L(d)` is the document's 1-based position in list L, and documents absent from a list simply contribute nothing. The constant k is conventionally 60, from the original formulation of the method. Documents are then sorted by that sum. Three properties follow directly: 1. **Scale invariance.** Any monotone transformation of either retriever's scores leaves the fused output unchanged. No calibration, no per-query statistics, no retuning after a model swap. 2. **Agreement is rewarded.** A document found by both legs accumulates two terms. With k=60, rank 20 in both lists gives 1/80 + 1/80 = 0.025, which beats rank 1 in a single list (1/61 ≈ 0.0164). Cross-retriever consensus is treated as evidence. 3. **Top-heaviness is damped by k.** The gap between rank 1 and rank 2 is 1/61 − 1/62 ≈ 0.00026, small relative to the value of appearing in a second list. Lower k sharpens the advantage of top positions; higher k flattens the curve and lets deeper agreement matter more. k is the one knob, and 60 is a reasonable default rather than a magic number. ## What RRF gives up Rank fusion is deliberately blind to confidence. If the dense leg returns one document at similarity 0.94 and everything else at 0.31, RRF sees only "rank 1, rank 2, rank 3" and cannot express that the first is in a different class. It also cannot express "this query matched nothing anywhere" — every list has a rank 1, so the fused list always looks populated even when it is empty of real matches. If you need an abstention signal or a relevance threshold, you must keep the raw scores alongside the fused order and consult them separately. RRF also treats the two legs as equally authoritative unless you add per-list weights (`w_L / (k + rank)`), which reintroduces a tuning parameter — but a rank-space one, which is far more stable across corpora than a score-space one. ## How to argue it in an interview The strong framing is not "RRF is better". It is: score fusion demands a calibration you cannot reliably obtain per query, and its failure mode — promoting the best of a bad list to a perfect score — is silent. Rank fusion trades away magnitude, which you often do not need for ordering, in exchange for removing the calibration problem entirely. That is why it is the sensible default and why teams that keep score fusion usually do so because they specifically need the confidence signal.

  • What does raising k from 10 to 60 actually change in the fused ordering?
    It flattens the reciprocal curve. With small k, rank 1 dominates: 1/11 versus 1/12 is a large relative gap, so a single list's top hit is hard to beat. With k=60 the gaps near the top shrink to fractions of a percent, so appearing in both lists at moderate ranks outweighs being first in one. In effect, larger k shifts weight from within-list position toward cross-list agreement.
  • When would you deliberately keep score fusion instead of RRF?
    When you need the magnitude, not just the order. Thresholding to decide 'no good answer, abstain', calibrating a confidence shown to users, or feeding a downstream component that consumes a score all need the raw values, which RRF destroys. A common compromise is to order by RRF and carry the original per-leg scores forward as metadata for those decisions.
  • How would you weight the lexical leg more heavily while still using rank fusion?
    Use weighted RRF: give each list a multiplier, so a document scores `sum of w_L / (k + rank_L)`. Doubling the lexical weight makes a lexical rank count roughly like two independent list appearances. This keeps the scale-invariance property — you are weighting sources, not comparing incomparable score magnitudes — and the weight tends to transfer across corpora better than a score-space alpha.

RRF is ranked-choice voting: each retriever submits a ballot ordering the candidates, and only the positions on the ballot count — not how strongly each voter claims to feel.

saying these in an interview costs you the question

  • Adds BM25 and cosine scores directly with no normalization
  • Believes min-max normalization makes scores comparable across queries
  • Treats k=60 as a mathematically derived constant
  • Thinks RRF preserves how confident each retriever was
  • Claims a document must appear in both lists to be fused

context