skip to content

Why add BM25 keyword retrieval alongside dense vectors, and how does reciprocal rank fusion combine them?

level: middleimportance: must knowfreq 56%

answer

  1. opposite failure modes
  2. rare literals versus paraphrases
  3. fuse by rank, not by score
  4. one over sixty plus rank
  5. agreement across lists wins

basics

~20 s

Dense embeddings match meaning but smear rare exact tokens such as an identifier, a part number or a docket number; BM25 matches those literally. Reciprocal rank fusion merges the two result lists by rank position, so the systems' incomparable raw scores never have to be reconciled.

solid answer

~50 s

Dense and lexical retrieval fail in opposite directions. An embedding model maps text into a semantic space, so it finds a paraphrase that shares no words with the query — but a rare literal string, like a case number or an internal ticket id, gets absorbed into a general-purpose representation and the exact document ranks nowhere. BM25 is the opposite: it nails the exact token and misses the paraphrase. Running both and fusing the results recovers most of each side's recall. **Reciprocal rank fusion** does the merge without touching scores: each document gets `sum over lists of 1 / (k + rank)`, with k conventionally 60, and the fused list is sorted by that sum. Because only rank positions matter, you never have to normalize a BM25 score against a cosine similarity — two quantities on unrelated scales. In 2026 practice hybrid retrieval plus a reranking stage is the default retrieval shape, not an optimization you add later.

code

python · 10 lines
python
def rrf(rankings, k=60):
    scores = {}
    for ranking in rankings:
        for rank, doc_id in enumerate(ranking, start=1):
            scores[doc_id] = scores.get(doc_id, 0.0) + 1.0 / (k + rank)
    return sorted(scores, key=scores.get, reverse=True)

bm25 = ["doc-9", "doc-3", "doc-1"]
dense = ["doc-3", "doc-7", "doc-9"]
print(rrf([bm25, dense]))   # ['doc-3', 'doc-9', 'doc-7', 'doc-1']

go deeper

for a junior

Be able to say that keyword search matches exact words and vector search matches meaning, and that running both and merging catches queries either one alone would miss.

for a middle

Explain BM25's ingredients — term frequency saturation, inverse document frequency, length normalization — and write the reciprocal rank fusion formula, including why fusing on rank avoids normalizing incomparable scores.

for a senior

Show how you would prove hybrid is worth its cost: labelled query set, recall@100 per leg and fused, sliced by query class, plus the operational invariant that both indexes hold the same document set.

for a principal

Own the retrieval architecture as a whole — two indexes to build and keep in sync, when learned sparse retrieval or tuned weighted fusion replaces plain RRF, and where fusion ends and reranking begins.

## Two retrievers, two failure modes Consider a legal e-discovery corpus: several million emails and attachments in a single matter, searched by paralegals under a production deadline. A paralegal searching *"who approved the shipment delay in Q3"* needs semantic matching — the responsive email may say "signed off on pushing the freight window". Dense retrieval handles that and BM25 does not, because there is no lexical overlap. The same paralegal searching for docket number `1:21-cv-04447` needs the opposite. That string is a rare token the embedding model has effectively never seen; it is tokenized into fragments and its contribution to a 1024-dimensional vector is diluted by the surrounding text. Dense retrieval returns documents that are *about* litigation. BM25 returns the exact document, because a rare term carries enormous inverse-document-frequency weight. Neither retriever is a superset of the other, which is why hybrid retrieval exists. ## What BM25 does BM25 is a lexical ranking function over an inverted index. It scores a document by summing, over the query terms it contains: term frequency with saturation (the tenth occurrence of a word adds far less than the second), inverse document frequency (rare terms count for much more than common ones), and a length normalization so long documents are not automatically favoured. It has no notion of meaning at all — synonyms, translations and paraphrases are invisible to it — but it is exact, cheap, interpretable and needs no model. ## What dense retrieval does An embedding model encodes query and chunk independently into vectors, and an approximate index returns the nearest chunks under a similarity metric. It generalizes across wording, handles morphology and often crosses languages. Its weaknesses are the complement of BM25's: rare literals, negation, and precise numeric or symbolic constraints. ## Reciprocal rank fusion The naive merge is to normalize both score sets and take a weighted sum. It is brittle. BM25 scores are unbounded and corpus-dependent; cosine similarities cluster in a narrow band whose location shifts by model and by query. Min-max normalization over a page of results makes the fused ranking sensitive to whichever outlier happened to land in that page, and the weights need retuning whenever the model, the analyzer or the corpus changes. RRF sidesteps this by discarding the scores. Each retriever contributes `1 / (k + rank)` for each document it returned, and contributions sum across retrievers. The constant k, conventionally 60, damps the influence of the very top ranks so that a single retriever's confident first place cannot dominate a document that both retrievers ranked respectably. Worked example with `k = 60`. BM25 returns `[doc-9, doc-3, doc-1]`; dense returns `[doc-3, doc-7, doc-9]`. - doc-3: `1/62 + 1/61 = 0.03252` - doc-9: `1/61 + 1/63 = 0.03227` - doc-7: `1/62 = 0.01613` - doc-1: `1/63 = 0.01587` Fused order: doc-3, doc-9, doc-7, doc-1. Documents both retrievers agreed on rise; single-list finds still survive, which is exactly the behaviour you want when either retriever alone is blind to a whole class of queries. ## Tradeoffs and alternatives RRF's virtues are that it is parameter-light, robust and needs no training data. Its cost is that it throws away genuine confidence information — a BM25 hit on an exact rare identifier is *much* stronger evidence than its rank alone conveys. Alternatives worth knowing: weighted score fusion after careful per-retriever normalization, which can beat RRF when you have labelled data to tune on; a weighted RRF variant that scales each retriever's contribution; and learned sparse retrieval (SPLADE-style), which produces a sparse term-weighted representation from a model and so gets some semantic expansion while remaining searchable on an inverted index. RRF also does not fix ordering *quality* — it fixes coverage. The usual pipeline is: retrieve maybe 50 to 100 candidates from each retriever, fuse, then send the fused head to a reranker that does the fine-grained ordering. ## Operating it Measure each leg separately. Recall@100 for BM25 alone, for dense alone, and for the fusion, computed over a labelled query set drawn from real usage, tells you whether hybrid is earning its keep and which leg is carrying which query class. If dense adds nothing on your corpus, you may have a chunking or embedding-model problem rather than a fusion problem. Slice by query type — natural-language questions versus identifier lookups — because the aggregate number hides exactly the split that motivates hybrid in the first place. The operational costs are real: two indexes to build, keep in sync and reindex, and two subsystems that can drift apart when only one ingestion path is updated. Treat "same document set in both indexes" as an invariant you actually assert, not one you assume.

  • Why not just normalize both score sets and take a weighted sum?
    Because the scales are not commensurable and not stable. BM25 scores are unbounded and depend on corpus statistics; cosine similarities sit in a narrow, query-dependent band. Min-max normalizing over a result page makes the fusion hostage to outliers, and any tuned weight has to be retuned when the model, the analyzer or the corpus changes. Weighted fusion can beat RRF when you have labelled data and are willing to maintain the tuning; RRF is the robust default when you are not.
  • What does the constant k in reciprocal rank fusion actually do?
    It damps the advantage of the very top ranks. With k around 60, the gap between rank 1 and rank 2 is small, so a document both retrievers placed in their top ten can outrank a document only one retriever put first. A small k makes the fusion behave more like a winner-takes-all over first places; a large k flattens the ranks towards uniform voting. Sixty is a convention worth sanity-checking on your own labelled set, not a law.
  • How would you tell whether the lexical leg is still earning its cost?
    Ablate it. Run your labelled query set through dense-only and through the fusion, and compare recall@100 sliced by query type rather than in aggregate. Identifier and code-like lookups are where BM25 earns its keep; if that slice barely moves, the lexical index may be misconfigured — a tokenizer splitting your identifiers, or documents missing from one of the two indexes.

saying these in an interview costs you the question

  • Dense retrieval handles exact identifiers fine, so BM25 is obsolete
  • Reciprocal rank fusion needs the two score scales normalized first
  • Hybrid retrieval always doubles query latency
  • Fusion improves the ordering quality, so no reranker is needed
  • BM25 understands synonyms because it uses inverse document frequency

context