skip to content

In LangChain, what does MultiQueryRetriever buy you, and what does it cost?

level: seniorimportance: should knowfreq 50%

answer

  1. One question sampled several ways
  2. An extra model call before searching
  3. Union of results, not a fused ranking
  4. Recall lever, not a ranking lever
  5. Cannot find what was never indexed

basics

~20 s

MultiQueryRetriever asks an LLM to rewrite the question into several phrasings, runs the base retriever on each, and returns the deduplicated union. It buys recall against vocabulary mismatch; it costs one extra LLM call of latency and a larger, unranked document set.

solid answer

~50 s

`MultiQueryRetriever.from_llm(retriever=base, llm=llm)` wraps any retriever. On each query it prompts the model to produce several alternative phrasings — the default prompt asks for three — runs the wrapped retriever once per generated query, and returns the union with duplicates removed; `include_original=True` keeps the user's own wording in the set too. What it fixes is vocabulary mismatch: the user says "can I get my money back", the manual says "refund eligibility", and one embedding of one phrasing missed it. The costs are real. You pay a full LLM round trip *before* retrieval starts, so it is on the critical path of every request. You then run N searches instead of one. And the union is not re-ranked or truncated to k, so you can hand the prompt three times the documents you budgeted for. Pair it with a reranker or compressor, and cache aggressively.

go deeper

for a junior

Know that it uses an LLM to rephrase the question several ways and pools the results, and that this is about improving recall.

for a middle

Explain the flow concretely — generate variants, search per variant, merge and deduplicate — and state that the output is not re-ranked or capped at k.

for a senior

Own the tradeoff in production terms: extra latency on the critical path, N-fold search amplification, context bloat, and non-determinism in tests.

for a principal

Decide when recall is worth a pre-retrieval model call at all, and set the standard that expansion stages must always be followed by a budgeted ranking stage.

## The problem it targets Dense retrieval embeds one query into one vector and looks for neighbours. That is a single sample from a distribution of ways the question could have been asked, and a single embedding can simply land in the wrong neighbourhood: colloquial phrasing against formal documentation, a symptom described where the docs describe a cause, an acronym the corpus spells out. `MultiQueryRetriever` attacks this by sampling the distribution more than once. An LLM generates several alternative phrasings of the user's question, each is embedded and searched independently, and the results are pooled. A document that only one of the phrasings would have found still makes it into the pool. ## Mechanics Construction is `MultiQueryRetriever.from_llm(retriever=base_retriever, llm=llm)`. The default prompt instructs the model to generate a small number of alternative versions of the question — three — one per line, and the retriever parses the lines into queries. You can supply your own prompt when the default's phrasing does not suit the domain, for example to force the rewrites to stay inside a controlled vocabulary. Each generated query is passed to the wrapped retriever. Because it wraps *any* retriever, the base can itself be a filtered vectorstore retriever, an MMR retriever or a fused hybrid — the expansion composes on top. The results are then merged with duplicates removed, and this is where the interesting property lives: the output is a **union**, not a fused ranking. Nothing re-scores the pooled set, and nothing truncates it back to `k`. With `k=4` and three generated queries you can receive up to twelve documents (plus the original query's results if `include_original=True`). For debugging, the retriever logs the generated query variants when you raise its module logger to INFO. Do this the first time you wire it up: seeing the rewrites is the fastest way to discover that the model is producing three near-identical restatements, which buys you nothing and costs you three searches. ## What it costs **Latency, on the critical path.** The rewrite call happens before any retrieval, so its full round trip is added to every request that misses your cache. On a chat interface this is the difference between a snappy and a sluggish first token, and it is paid even for questions that the base retriever would have answered perfectly. **Search amplification.** N queries means N searches. Against a hosted vector database this is N times the query cost and N times the load; against an in-process index it is mostly CPU. **Context bloat.** The unranked union is the most-missed cost. Teams add multi-query for recall, feed the result straight into the prompt, and quietly triple their token spend per request while pushing genuinely relevant chunks into the middle of a long context. The fix is to treat multi-query as a *candidate generation* stage and always follow it with a ranking or compression stage that cuts back to a budgeted set. **Non-determinism.** The rewrites come from a model, so the same question can retrieve differently on two runs. That complicates regression testing and makes production incidents harder to reproduce. Setting temperature to zero helps but does not fully remove it. ## What it does not fix This is the discriminating half of the answer in an interview. - **Content that is not in the index.** No amount of rephrasing finds a document you never ingested. Multi-query turns a *ranking* miss into a hit; it cannot turn a *coverage* miss into one. - **Bad chunking.** If the answer is split across two chunks such that neither is self-contained, every phrasing retrieves a fragment. - **Ranking.** It adds candidates; it never reorders them by relevance. If your complaint is "the right document is in the top twenty but not the top four", you want a reranker, not query expansion. - **Embedding mismatch.** If the index was built with one embedding model and queries are embedded with another, all N searches are equally broken. ## When to reach for it Good fits: short, ambiguous, colloquial user queries over a formal corpus; domains with heavy synonymy; systems where recall matters much more than latency, such as offline research or batch enrichment. Poor fits: latency-sensitive interactive chat, corpora with a controlled vocabulary the users already speak, and any pipeline that cannot afford a reranking stage after it. In an interactive product, a cross-encoder rerank over a wider single retrieval is often the better spend of the same latency budget.

  • How would you keep multi-query expansion from blowing the context budget?
    Treat it as candidate generation only. Follow it with a ranking or compression stage — a cross-encoder reranker or an embedding-similarity filter — that cuts the pooled union back to a fixed number of documents before the prompt is assembled. Without that stage the union of N searches at k each goes straight into the prompt, multiplying tokens per request.
  • Your rewrites all come back nearly identical. What is going wrong?
    Usually the prompt or the model. The default instruction asks for alternative phrasings, and a small or heavily aligned model often produces three cosmetic restatements that embed to almost the same vector, so you pay three searches for one result set. Supply a domain-specific prompt that forces genuinely different angles — a synonym-swapped version, a formal-register version, a symptom-versus-cause version — and inspect the logged variants.
  • When is a reranker the better spend than query expansion?
    When the right document is already being retrieved but ranked too low — a ranking problem, not a coverage problem. Expansion adds candidates and adds an LLM call before retrieval; a cross-encoder rerank runs after a single wider retrieval and directly fixes ordering. If you can only afford one extra stage in an interactive product, retrieve wider and rerank.

saying these in an interview costs you the question

  • Thinks it reranks or scores the merged results
  • Believes it can find documents missing from the index
  • Ignores that the rewrite call precedes every retrieval
  • Feeds the whole union into the prompt unbudgeted
  • Assumes results are reproducible across runs

context