How do you build hybrid BM25-plus-vector retrieval in LlamaIndex?
answer
- two legs, one merged list
- the store can do it, or the framework can
- merge by rank, not by raw score
- one leg lives in your process
- new documents may reach only one leg
basics
~10 sEither use a vector store that supports native hybrid search through as_retriever(vector_store_query_mode="hybrid"), or run a BM25Retriever and a vector retriever side by side and merge them with QueryFusionRetriever using mode="reciprocal_rerank".
solid answer
~50 sThere are two routes and they have different operational profiles. If the backing store supports hybrid search, `index.as_retriever(vector_store_query_mode="hybrid")` pushes both the keyword and vector legs into the store, which does the blending server-side; the weighting parameter (often `alpha`) and its exact semantics belong to that integration, so check the store's docs rather than assuming. The framework-side route is `QueryFusionRetriever`: give it a vector retriever plus a `BM25Retriever` from the `llama-index-retrievers-bm25` package, set `mode="reciprocal_rerank"` for reciprocal-rank fusion, `similarity_top_k` for how many fused results survive, and `num_queries=1` to disable its LLM query generation unless you want it. The catch is that `BM25Retriever` builds an in-memory index from the nodes or docstore you hand it at construction, so it does not track vector-store writes: you must rebuild or persist and reload it when the corpus changes, and it holds the whole corpus in the app process. Hybrid earns its keep on exact tokens — part numbers, error codes, names — that embeddings blur.
code
python · 16 linesfrom llama_index.core.retrievers import QueryFusionRetriever
from llama_index.retrievers.bm25 import BM25Retriever
vector_retriever = index.as_retriever(similarity_top_k=10)
bm25_retriever = BM25Retriever.from_defaults(
docstore=index.docstore, similarity_top_k=10
)
fusion = QueryFusionRetriever(
[vector_retriever, bm25_retriever],
similarity_top_k=10,
num_queries=1,
mode="reciprocal_rerank",
use_async=True,
)
nodes = fusion.retrieve("error code E-4471 retry policy")go deeper
Know that hybrid means combining keyword search with embedding search, and that LlamaIndex offers both a store-native mode and a fusion retriever that merges two retrievers' results.
Explain how the fusion retriever is wired — child retrievers, a fusion mode, a top-k for the merged list — and why merging by rank avoids comparing incomparable score scales.
Bring the operational failure: the lexical leg is an in-memory structure that goes stale against a live vector store, and say concretely how you would rebuild, persist or push it into the store instead.
Decide whether hybrid earns its complexity at all — measure recall per leg on a labelled set, identify the query class that benefits, and weigh a store with native hybrid against carrying a second index in your own process.
## Why hybrid at all Dense vector retrieval matches meaning and is weak precisely where users are most literal: SKU codes, error identifiers, function names, rare proper nouns, version strings. Those tokens carry little semantic signal, so their embeddings sit near everything and nothing. Lexical scoring — BM25 — is the opposite: it nails exact and rare terms and is blind to paraphrase. Hybrid retrieval runs both and merges, so the recall floor is the union rather than either leg alone. In LlamaIndex there are two ways to assemble it. ## Route 1: native hybrid in the vector store Several vector-store integrations implement hybrid search themselves, keeping a sparse or keyword index alongside the dense one. In that case you ask the retriever for it: `index.as_retriever(vector_store_query_mode="hybrid", similarity_top_k=10)` The store executes both legs and returns one merged, ranked list. Some integrations expose a weighting parameter — commonly `alpha` — but the meaning and even the direction of that parameter is defined by the store, not by LlamaIndex, so verify it against the specific integration rather than porting a value from another product. Advantages: one network round trip, one system of record, and the keyword index stays consistent with writes because the store owns both. Disadvantage: you inherit whatever fusion the store implements, and you can only use it if your store supports it. ## Route 2: framework-side fusion When the store has no hybrid mode, or when you want control over the merge, build the two retrievers separately and fuse them. `BM25Retriever` lives in the `llama-index-retrievers-bm25` package and is constructed with `BM25Retriever.from_defaults(nodes=nodes, similarity_top_k=10)` or from a docstore, `from_defaults(docstore=index.docstore, similarity_top_k=10)`. It tokenizes the corpus and builds a BM25 index **in memory, at construction time**. `QueryFusionRetriever` takes the list of retrievers and merges their outputs. The parameters that matter: - `mode` — `"reciprocal_rerank"` for reciprocal rank fusion, which combines by *rank position* rather than raw score. This is the safe default precisely because BM25 scores and cosine similarities are on incomparable scales; RRF never has to compare them. Score-based modes exist (relative-score and distance-based normalization) for when you want magnitude to matter. - `similarity_top_k` — how many fused results the retriever finally emits. Each child retriever keeps its own `similarity_top_k` for the candidates it contributes; setting only the outer one starves the fusion. - `num_queries` — by default the fusion retriever asks an LLM to generate several query variants and retrieves for each. That is query expansion bundled into the retriever, and it adds an LLM call plus more store round trips to every retrieval. Set `num_queries=1` to turn it off when you only want fusion. - `use_async` — issue the child retrievals concurrently, which matters once you have two legs times several generated queries. ## The operational trap in BM25Retriever This is the part interviewers probe. The vector store is a live service: you write new nodes and the next query sees them. `BM25Retriever` is a process-local structure built from the nodes you handed it. Consequences: - **Staleness.** Documents added after construction are invisible to the lexical leg. Your hybrid pipeline silently degrades to vector-only for new content, and nobody gets an error. - **Memory and startup cost.** The corpus text lives in the application process, and building the index costs time on every boot. That is fine for tens of thousands of nodes and untenable for tens of millions. - **Replica skew.** Each service replica builds its own copy, so two replicas can disagree about lexical results if they started at different times. Mitigations: rebuild on a schedule tied to ingestion, persist and reload the built retriever so boots are fast, or move the lexical leg into a store that supports native hybrid so consistency is someone else's problem. At real scale the last option is usually the answer. ## Evaluating whether it helped Hybrid is not automatically better; it costs a second retrieval path and, with query generation on, an LLM call. Measure it: take a labelled question set, compute recall@k for the vector leg alone, the BM25 leg alone, and the fusion. If the union recall is not meaningfully above the vector leg, the corpus is not lexical enough to justify the moving parts. Where hybrid does win, the win is usually concentrated in a query class — identifiers, codes, names — which is worth knowing because it tells you whether a cheaper fix (better chunk metadata, a metadata filter, an exact-match prefilter) would do. Finally, hybrid and reranking are complements, not alternatives: fusion widens the candidate pool across two retrieval philosophies, and a cross-encoder reranker then cuts it down to the few nodes the LLM should actually read.
- Why does reciprocal rank fusion suit merging BM25 with vector results?Because it combines rank positions rather than raw scores. BM25 scores are unbounded and corpus-dependent while cosine similarities sit in a small bounded range, so any direct comparison needs normalization that is fragile across corpora. RRF sidesteps that entirely — a document ranked highly by either leg surfaces, and documents ranked well by both rise to the top.
- You add 50k new documents to the vector store and hybrid quality drops for them. What is the likely cause?The BM25 leg is stale. It was built in memory from the nodes available at construction, so it never saw the new documents and the lexical half of every query over them returns nothing useful. Rebuild or reload it as part of the ingestion pipeline, or move the lexical leg into a store that supports native hybrid search.
- What does QueryFusionRetriever do by default that you might not want, and how do you stop it?By default it generates several rewritten queries with an LLM and retrieves for each before fusing. That adds an LLM call and multiplies store round trips on every retrieval — real latency and cost you may not have budgeted. Set `num_queries=1` to retrieve only the original query and use the retriever purely as a fusion layer.
- Does hybrid retrieval remove the need for a reranker?No — they solve different halves. Fusion improves recall by widening the candidate pool across two retrieval philosophies; a cross-encoder reranker improves precision by scoring query-document pairs and cutting the pool to the few nodes worth sending to the LLM. The usual production shape is hybrid retrieval into a reranker, not one instead of the other.
saying these in an interview costs you the question
- Assuming BM25Retriever queries the vector store live
- Merging BM25 scores and cosine similarities by simple addition
- Leaving LLM query generation on without budgeting for the extra calls
- Setting only the fusion retriever's top_k and starving its child retrievers
- Adding hybrid retrieval without measuring recall against vector-only