How does LangChain's EnsembleRetriever combine BM25 and vector search results?
answer
- Two retrievers, incomparable score scales
- Fuse on position, not on magnitude
- A constant flattens the top ranks
- Consensus beats one list's favourite
- Exact identifiers versus paraphrase
basics
~20 sBy reciprocal rank fusion, not by score arithmetic. EnsembleRetriever runs each wrapped retriever, converts every result to its rank position, sums weight divided by (c + rank) across retrievers, and returns the deduplicated candidates ordered by that fused score.
solid answer
~50 s`EnsembleRetriever(retrievers=[bm25, dense], weights=[0.4, 0.6])` runs each wrapped retriever on the query and fuses their result lists with reciprocal rank fusion: each document scores `weight / (c + rank)` from every list it appears in, summed, with `c` defaulting to 60. Documents surfaced by both retrievers rise; documents unique to one still get in. The key design point is that fusion happens over **ranks**, never raw scores — a BM25 score and a cosine similarity live on incomparable scales, so averaging them is meaningless, while rank position is a common currency. `weights` biases the mix toward lexical or semantic evidence. The practical payoff is covering each method's blind spot: BM25 nails exact identifiers, error codes and rare product names that embeddings smear together, and dense retrieval handles paraphrase. Note `BM25Retriever` is an in-process index over a document list, so at real scale you want a store with server-side lexical search.
code
python · 9 linesfrom langchain_community.retrievers import BM25Retriever
bm25 = BM25Retriever.from_documents(chunks)
bm25.k = 10
dense = vectorstore.as_retriever(search_kwargs={"k": 10})
hybrid = EnsembleRetriever(retrievers=[bm25, dense], weights=[0.4, 0.6])
docs = hybrid.invoke("error E4032 on refund submission")go deeper
Know that hybrid retrieval combines keyword and vector search, and that LangChain does the combining with a rank-based fusion rather than by adding scores.
Explain the reciprocal rank fusion formula, the role of the weights, and why rank is used instead of raw scores from incomparable scales.
Bring the operational side: doubled query cost, in-process BM25 limits at scale, capping the fused list, and measuring recall on identifier-heavy versus paraphrase-heavy queries.
Decide whether hybrid belongs in the application layer at all versus in a store with native hybrid search, and own the evaluation set that justifies the weights.
## Why hybrid at all Dense and lexical retrieval fail in opposite directions, and the failures are predictable enough to engineer around. Embeddings generalise. That is their strength — "how do I get my money back" finds a page titled "Refund eligibility" — and their weakness: they smear rare, arbitrary tokens together. Error code `E4032`, part number `AX-119-B`, an internal service name, a person's surname — these carry almost no semantic signal, so the nearest neighbours of a query containing them are documents that are *topically* similar and contain the wrong identifier entirely. BM25 is the inverse. It is exact-term matching with term-frequency and document-length weighting, so it finds `E4032` unerringly and fails completely when the user's vocabulary does not overlap the document's. A corpus of technical documentation contains both kinds of query, often in the same sentence. Hybrid retrieval is how you stop choosing. ## The fusion mechanism The obvious approach — average the two scores — does not work. BM25 produces unbounded scores whose scale depends on corpus statistics; cosine similarity sits in a fixed narrow band; several vector stores return distances where lower is better. There is no principled normalisation between them, and any constant you pick to make them comparable silently stops being right when the corpus grows. Reciprocal rank fusion sidesteps the problem by throwing the scores away and keeping only the ordering. For each retriever, a document at rank `r` contributes `weight / (c + r)`. Contributions are summed across retrievers, and the fused list is sorted by that sum. The constant `c` (default 60) flattens the top of the curve: without it, rank 1 would dominate rank 2 overwhelmingly and a single retriever's top hit could never be outvoted. With `c=60`, ranks 1 and 2 differ by only a few percent, so consensus between retrievers matters more than any one list's top pick. That is the property you want — a document both retrievers liked outranks a document one retriever loved. `weights` scales each retriever's contribution, letting you bias toward lexical or semantic evidence based on your query mix. Equal weights are the default and a fine starting point. ## Wiring it in LangChain `EnsembleRetriever` takes any list of retrievers, not just two and not just these kinds — a metadata-filtered vectorstore retriever, an MMR retriever and a lexical retriever can all be fused together. Each sub-retriever keeps its own configuration, so their `k` values independently control how many candidates each contributes. The output is the deduplicated union ordered by fused score, so it can be larger than any single retriever's `k`. Cap it deliberately — through each sub-retriever's `k`, or by following the ensemble with a reranking or compression stage that cuts to a budget. An `id_key` option lets deduplication key on a metadata identifier instead of comparing content, which matters when the same source document arrives from two retrievers with slightly different text. ## The BM25Retriever caveat `BM25Retriever` in the community distribution builds an in-memory index over a Python list of Documents (via `from_documents`, with `k` settable on the instance) and depends on a pure-Python BM25 implementation. That is excellent for a demo or a few thousand chunks and wrong for production at scale: the whole corpus sits in every worker process, the index rebuilds on every start, and it is not shared, incremental or persistent. When your corpus outgrows that, the answer is a document store or search engine that does lexical scoring server-side — many vector databases now expose hybrid search natively — and you point the ensemble's lexical slot at its retriever instead. Its tokenisation is also whitespace-oriented by default, which degrades for languages that do not delimit words with spaces; that is a reason to prefer a real search engine's analyzer chain for multilingual corpora. ## Costs and evaluation Hybrid retrieval means two searches per query, so latency is the slower of the two (they can run concurrently) plus fusion, and infrastructure cost roughly doubles. The fused list is wider, which pushes work onto whatever stage follows. And because RRF ignores magnitudes, it also discards a genuinely useful signal: a document that BM25 scored overwhelmingly highly contributes exactly the same as a document it scored barely above the next one. In return you get a fusion that needs no calibration and does not drift as the corpus changes — usually the right trade, but worth naming as a trade rather than a free win. Measure it. Build a labelled query set that deliberately includes identifier-style queries and paraphrase-style queries, and compare recall at k for lexical, dense and fused. Hybrid usually wins overall while losing on some individual queries, and the weights are only tunable against numbers.
- Why not just normalise and average the two retrievers' scores?Because there is no stable mapping between them. BM25 scores are unbounded and depend on corpus statistics that change as you ingest, cosine similarity sits in a fixed band, and some stores return distances where lower is better. Any normalisation constant you calibrate today drifts as the corpus grows. Rank fusion needs no calibration because position is already comparable across systems.
- What does the constant c=60 in reciprocal rank fusion actually do?It damps the difference between top ranks. The contribution is weight divided by (c + rank), so with c=60 the gap between rank 1 and rank 2 is about one and a half percent rather than the fifty percent you would get with c=0. That stops any single retriever's top hit from dominating and makes agreement between retrievers the strongest signal in the fused ordering.
- How would you tune the weights argument?Only against a labelled query set that reflects your real traffic, split by query shape. Score recall at k for the lexical retriever alone, the dense one alone, and several weightings of the fused pair. Corpora full of identifiers, codes and rare proper nouns tilt lexical; conversational corpora tilt dense. Picking weights by intuition and shipping them is how hybrid retrieval ends up worse than either half.
saying these in an interview costs you the question
- Says it averages the two retrievers' similarity scores
- Thinks hybrid means running whichever retriever wins
- Ships BM25Retriever over a million-chunk corpus
- Assumes the fused output is capped at k
- Tunes the weights without a labelled query set