skip to content

How does Elasticsearch's rrf retriever combine a BM25 query with a kNN search?

level: seniorimportance: should knowfreq 60%

answer

  1. It fuses positions, not scores
  2. Avoids normalizing incomparable score scales
  3. One over a constant plus the rank
  4. The constant defaults to 60
  5. A document outside the window cannot be rescued

basics

~20 s

The rrf retriever runs each sub-retriever separately, then fuses their ranked lists by summing 1 divided by (rank_constant plus each list's rank). It uses positions, not scores, so incomparable BM25 and vector scores never have to be normalized.

solid answer

~50 s

You wrap two sub-retrievers — typically a `standard` retriever holding a BM25 query and a `knn` retriever — inside an `rrf` retriever. Each runs independently and produces its own ranked list of up to `rank_window_size` documents. Reciprocal rank fusion then scores every document as the sum over lists of `1 / (rank_constant + rank)`, with `rank_constant` defaulting to 60, and the fused ranking is returned. The point is that it never compares a BM25 score to a vector similarity: those live on different, query-dependent scales, and normalizing them is fragile. RRF only needs an ordering. The cost is that magnitude information is thrown away — a document that BM25 loved overwhelmingly gets exactly the same credit as one that barely edged into first place — and a document must appear in a list's window to contribute at all.

code

json · 27 lines
json
POST /articles/_search
{
  "retriever": {
    "rrf": {
      "retrievers": [
        {
          "standard": {
            "query": {
              "match": { "body": "how do i reset my password" }
            }
          }
        },
        {
          "knn": {
            "field": "body_embedding",
            "query_vector": [0.05, -0.31, 0.62],
            "k": 50,
            "num_candidates": 200
          }
        }
      ],
      "rank_constant": 60,
      "rank_window_size": 100
    }
  },
  "size": 10
}

go deeper

for a junior

Know the shape: two child retrievers nested inside an rrf retriever in one _search call, fused by rank position rather than by score.

for a middle

Reproduce the formula — one over rank_constant plus rank, summed across lists — and explain why rank fusion avoids normalizing incomparable BM25 and vector scores.

for a senior

Show the tuning judgment: size rank_window_size against the kNN leg's k, evaluate the hybrid against each leg on a judged query set, and account for the doubled per-shard cost.

for a principal

Own when hybrid is the right architecture at all, what the relevance evaluation programme measures, and when the answer is a reranking stage or a better embedding model rather than more fusion tuning.

## The problem RRF solves A BM25 score and a vector similarity are not on the same scale, and worse, BM25 scores are not even comparable across queries: they depend on term statistics, field lengths and the number of clauses. So `bm25_score + knn_score` is arithmetic on incompatible units. Min-max normalizing each list first is possible but brittle — the normalization depends on the outliers in each list, so one anomalous top hit rescales everything beneath it. Reciprocal rank fusion sidesteps the problem by discarding scores entirely and using only positions. ## The formula For a document `d` appearing in a set of ranked result lists: ``` score(d) = Σ over lists 1 / (rank_constant + rank_in_that_list(d)) ``` With the default `rank_constant` of 60, a document ranked first in one list contributes 1/61 ≈ 0.0164; ranked second, 1/62 ≈ 0.0161. A document absent from a list contributes nothing from it. The differences between adjacent ranks are deliberately tiny, which is the whole design: appearing in *both* lists matters much more than winning either one. A document ranked 5th by BM25 and 5th by kNN beats a document ranked 1st by BM25 and absent from the vector list. `rank_constant` controls how flat that curve is. A larger constant flattens the differences further, so consensus across lists dominates even more; a smaller constant sharpens the advantage of top positions. ## The Elasticsearch shape Retrievers are a search-request abstraction: a `standard` retriever wraps an ordinary query, a `knn` retriever wraps an approximate vector search, and compound retrievers like `rrf` and `text_similarity_reranker` take other retrievers as children. So a hybrid request nests naturally, and the whole thing is still a single `_search` call — one round trip, one set of shard requests. Key parameters on the `rrf` retriever: - **retrievers** — the child retrievers to fuse. Two is typical; more is legal. - **rank_constant** — the constant in the denominator, default 60. - **rank_window_size** — how many documents each child contributes to the fusion. It must be at least as large as the request's `size`, and it is the real quality dial: a document that never entered a child's window cannot be rescued by fusion. If your lexical leg only returns 10 documents and the right answer is lexically 40th, RRF will never see it. Each child's own limits still apply beneath that: the kNN retriever's `k` and `num_candidates` govern how many and how well vector neighbours are found before fusion sees them. ## What you give up RRF's indifference to magnitude is both its strength and its limitation: - **No confidence signal.** A near-exact vector match and a mediocre one both occupy rank 1. If the vector leg is confidently right, RRF still only gives it a rank-1 share. - **Limited weighting.** You cannot easily say "trust lexical twice as much as semantic". Tuning `rank_constant` shifts the shape of the curve but does not weight the legs against each other. - **Score opacity.** The final `_score` is a fusion artefact in a tiny numeric range; it means nothing to a downstream threshold, and "why is this document 3rd" is harder to explain than a BM25 explanation. When you need magnitude or asymmetric trust, the usual answer is not to hand-roll score normalization but to add a reranking stage — fuse cheaply with RRF to build a candidate set, then reorder the top slice with a cross-encoder style reranker that scores query and document together. ## Practical guidance - Set `rank_window_size` generously — larger than `size`, large enough that a merely-decent match on either leg still enters the window. It costs a bit of merge work, not a full extra search. - Keep the kNN leg's `k` in line with `rank_window_size`; a `k` of 10 under a window of 100 wastes the window. - Evaluate the hybrid, not the legs. Build a judged query set and compare hybrid nDCG against lexical-only and vector-only. Hybrid is usually better, but not universally — on corpora of exact identifiers, part numbers and code, a strong lexical leg diluted by a mediocre semantic leg can lose ground. - Watch latency: the request now pays for both a lexical search and a graph traversal on every shard. ## Version note Retrievers arrived in Elasticsearch 8.14 and the `rrf` retriever matured over the later 8.x releases, superseding the earlier top-level `rank: { rrf: ... }` form. On a 9.x cluster the retriever syntax is the one to write.

  • Why not just normalize the BM25 and vector scores and add them together?
    Because BM25 scores are unbounded and query-dependent — they shift with term statistics, field length and clause count — so any normalization is computed from the current result list and moves with its outliers. One anomalous top hit rescales everything below it, and a weighting tuned on one query set drifts on another. RRF needs only an ordering, which is stable, and that is why it is the default hybrid strategy.
  • What is the practical effect of raising rank_window_size in an rrf retriever?
    Each child retriever contributes more documents to the fusion, so a document that ranks moderately on one leg and well on the other can still surface. It is the main lever against "the right answer was never in the pool". The cost is more documents to merge and slightly more work per request, not a second search — so it is usually cheap relative to the recall it buys.
  • When would you add a reranking stage on top of RRF instead of tuning rank_constant?
    When ordering within the top slice matters more than assembling the candidate set. RRF is good at pooling; it is deliberately blind to how strong a match is. A `text_similarity_reranker` scores query and document together over the top few dozen fused results, restoring the magnitude information RRF discarded. You pay real latency for it, so it goes on the top 50, not the top 500.

Two judges rank the same finalists on different scoring systems. Rather than converting their marks, you add up each contestant's placings — whoever both judges put near the top wins, even if neither judge put them first.

saying these in an interview costs you the question

  • Thinks RRF averages the BM25 and vector scores
  • Assumes the fused _score is comparable across queries
  • Leaves rank_window_size at the request size
  • Believes hybrid always beats lexical-only retrieval
  • Says rank_constant weights one retriever against the other

context