In RAG, when do you fan out query paraphrases instead of decomposing into sub-questions?
answer
- two shapes, not two names for one
- same answer worded differently vs different answers
- parallel paraphrases, possibly dependent sub-questions
- check whether the parts live in different documents
- every extra shape costs another call
basics
~20 sFan out paraphrases when the question has one intent but is worded unlike your documents; decompose when it genuinely contains several answerable parts. Paraphrases are independent searches merged into one list; sub-questions are retrieved and answered separately, then joined.
solid answer
~50 sThey are two different shapes. **Multi-query fan-out** takes a single-intent question, asks a model for three or four alternative phrasings, embeds each, retrieves for each, and merges the result lists. It is a recall fix: one embedding is one point in vector space, and a terse or oddly-worded question may sit far from the passage that answers it. **Decomposition** applies when the question is genuinely compound — "how did our EMEA gross margin change after the 2023 acquisition?" needs the pre-acquisition figure, the post-acquisition figure, and possibly the acquisition date. Those facts live in different filings, so you retrieve per sub-question and join the sub-answers. The tell: if the parts have different answers in different documents, decompose; if they are the same answer said differently, fan out. Sub-questions can also be *dependent*, which forces sequential execution; paraphrases are always parallel.
code
python · 16 linesplan = {
"fanout": [
"EMEA gross margin trend",
"European segment gross margin 2023",
"gross margin by geographic region",
],
"decompose": [
{"id": 1, "q": "What was EMEA gross margin before the 2023 acquisition?", "needs": []},
{"id": 2, "q": "What was EMEA gross margin after the 2023 acquisition?", "needs": []},
{"id": 3, "q": "How large is the difference between them?", "needs": [1, 2]},
],
}
parallel = [s["q"] for s in plan["decompose"] if not s["needs"]]
print("run concurrently:", parallel)
print("paraphrases are always parallel:", len(plan["fanout"]))go deeper
Know that a RAG system can rewrite the user's question before searching, and that generating several phrasings is one way to find documents worded differently from the question.
Be ready to explain the two shapes precisely: independent paraphrases merged into one result set, versus sub-questions retrieved and answered separately then joined, and to say which a given question needs.
Show you have debugged this in production — log the generated paraphrases and sub-question plans, recognize rewriter drift and bad splits from the traces, and know that both failures look like a retrieval problem until you read the plan.
Own the policy question: transformation multiplies calls, so decide per query class which shape is applied, what it buys in measured recall, and where a fixed pipeline or a better index would deliver the same win for less.
## The problem both techniques attack A vector retriever searches with whatever the user typed. That text is embedded into a single vector, and the index returns the chunks whose embeddings are nearest to it. Two things routinely go wrong. First, the user's phrasing may simply not resemble the phrasing of the passage that answers it, so the right chunk is not in the top-k. Second, the user may have asked several things at once, and no single passage in the corpus answers all of them — the top-k fills with documents that match the loudest term in the question while the other parts go unretrieved. Multi-query fan-out addresses the first problem. Decomposition addresses the second. They are frequently confused because both start by calling an LLM on the query and both end with more than one retrieval, but the shapes differ in a way that determines your execution plan. ## Multi-query fan-out You prompt a model: "generate four alternative phrasings of this question." Each phrasing is embedded and searched independently, and the result lists are merged into one candidate set for the generator. Nothing about paraphrase two depends on paraphrase one, so all searches run in parallel and the wall-clock cost is one LLM call plus one round of concurrent retrievals. The intuition is coverage. An embedding is a point; the region of the corpus you actually want may be an area. Several paraphrases scattered around the original give you several nearby points, and their union covers more of that area. It is most valuable when the input is short and idiosyncratic — a two-word bug-tracker title, a search box query, an internal abbreviation — because short text embeds noisily and small wording changes move the vector a long way. Note that the *answer* is one answer; you are not splitting the information need, only the way you ask for it. ## Sub-question decomposition Here you prompt a model to split the information need itself: "list the sub-questions that must be answered to answer this." The output is a small plan. For "how did our EMEA gross margin change after the 2023 acquisition?" a reasonable plan is (1) what was EMEA gross margin in the periods before the acquisition, (2) what was it after, (3) what else changed in the segment definition. Each is retrieved for separately — likely hitting different 10-K filings and different sections — and each gets its own short answer. A final call joins those sub-answers into the response. Decomposition changes the *unit of retrieval*. Instead of hoping one top-k contains evidence for every part, you give every part its own budget of chunks. That is why it rescues compound questions that fan-out cannot: paraphrasing a two-part question just produces four two-part questions, all of which retrieve the same lopsided results. ## Independent versus dependent sub-questions A decomposition plan is a small dependency graph, not always a list. Some sub-questions are independent and can be retrieved concurrently. Others are dependent: you cannot ask "what is that carrier's claims SLA?" until you know which carrier serves the lane. Dependent chains serialize, so latency grows with depth rather than breadth, and each hop compounds the risk that a wrong intermediate answer sends the next retrieval somewhere useless. Where the sub-questions are independent, keep them independent — it is cheaper and more robust. ## Choosing between them Useful heuristics: - Do the parts of the question have *different* answers found in *different* documents? Decompose. - Is the question one intent whose vocabulary probably differs from the corpus? Fan out. - Is the question a single lookup already phrased in corpus vocabulary? Do neither — plain retrieval is fine, and transformation only adds latency. - Is the question compound *and* oddly worded? Decompose first, then optionally fan out each sub-question. Do not do both by default; the call count multiplies. ## What it costs and how it fails Both techniques trade calls for recall. Fan-out adds one planning call plus N retrievals. Decomposition adds a planning call, N retrievals, usually N answering calls, and a join call — an easy 5x on both tokens and latency for a three-part question. The characteristic failure of fan-out is *drift*: the model rewrites the question into paraphrases that quietly change its meaning, pulling in confidently-retrieved but irrelevant chunks that then outvote the correct ones. Constrain the rewriter to preserve entities, numbers and dates. The characteristic failure of decomposition is a *bad split*: the plan omits a part, invents a part the user never asked about, or splits along a seam the corpus does not follow, so each sub-retrieval is individually plausible and the joined answer is subtly wrong. Log the generated sub-questions — when a decomposition-based system gives a strange answer, the plan is usually where it went wrong, and it is the cheapest artifact to inspect.
- How would you constrain the rewriter so paraphrases don't drift away from the original intent?Pin the invariants explicitly: instruct the model to preserve named entities, dates, figures and negations, and to vary only phrasing and level of specificity. Give it few-shot examples of acceptable and unacceptable rewrites. Then verify cheaply — if a paraphrase drops an entity that appeared in the original, discard it rather than searching with it. Logging paraphrases alongside their retrieved chunks makes drift obvious in review.
- What happens if you decompose a question the corpus answers in a single passage?You pay the full cost for no gain, and you can lose quality. Each sub-question retrieves a slice of the same passage plus unrelated neighbours, the sub-answers are partial, and the join step has to reassemble something that was already coherent. That reassembly is where fabricated connective tissue creeps in. Single-passage questions should bypass the transform entirely.
- Can you combine fan-out and decomposition, and when is that justified?Yes — decompose first, then fan out each sub-question — but reserve it for high-value, low-volume queries such as a research or analyst workflow where a slow, thorough answer is acceptable. The call count is the product of the two fan-outs, so a three-part question with four paraphrases each means twelve retrievals plus the planning and joining calls. For interactive traffic that is rarely worth it.
saying these in an interview costs you the question
- Calling multi-query and decomposition the same technique
- Assuming paraphrases can answer a compound question
- Decomposing every query regardless of shape
- Ignoring that dependent sub-questions must run sequentially
- Believing more queries always improves answer quality