skip to content

Which RAG queries justify multi-query fan-out when the latency budget is one second?

level: principalimportance: should knowfreq 45%

answer

  1. count the sequential model calls, not the searches
  2. most traffic should skip the transform
  3. route by query shape and by surface
  4. measure lift per class, and the tail
  5. cheaper levers exist at ingest time

basics

~20 s

Only queries whose failure is expensive and whose shape needs it — compound, terse or vocabulary-mismatched questions. Gate the transform behind a cheap classifier, because the added sequential model calls, not the extra vector searches, are what break a one-second budget.

solid answer

~50 s

Start from where the time goes. Extra vector searches cost single-digit milliseconds and run in parallel; the added *model* calls are the bill. A plan-then-retrieve-then-answer-per-sub-question-then-join pipeline can turn a 900ms answer into six seconds for a live analyst tool, and those calls are sequential, so concurrency does not save you. That means the decision is per query class, not global. Route cheaply: a small classifier or a few rules — does the question contain multiple clauses, comparisons, or entities the index barely covers? — decides whether a query pays for the transform, and everything else goes straight to plain retrieval. Then measure the lift honestly on each class; many teams find fan-out helps a narrow slice and does nothing for the rest. Finally, pick a different lever where you can: better chunking, a hybrid index, or reserved context slots often buy the same recall for zero added latency.

go deeper

for a junior

Know that transforming a query means extra model calls before the answer, and that this makes responses slower, so it is not something to switch on for every question.

for a middle

Be able to break the added latency down — one planning call for fan-out, several sequential stages for decomposition — and explain why the extra vector searches are not the expensive part.

for a senior

Show how you would gate it: a cheap classifier on query shape, per-class evaluation of the lift, path logging in production, and percentile latency rather than averages.

for a principal

Own the spend decision. Define which query classes and which product surfaces justify the cost, insist the win is measured per class, and compare transformation against cheaper levers like chunking, metadata and hybrid retrieval before committing latency to it.

## Account for the latency before arguing about quality Query transformation is usually discussed as a quality technique and paid for as a latency technique. Do the arithmetic before deciding anything. A plain RAG turn is roughly: embed the query (tens of milliseconds), search (single-digit to low tens of milliseconds), generate (the dominant term, hundreds of milliseconds to seconds). Add multi-query fan-out and you insert one model call to produce the paraphrases — sequential, before any retrieval can start — plus N searches that run concurrently and therefore cost about one search's time. So fan-out adds roughly one model round trip. Add sub-question decomposition and the shape is worse: a planning call, then N retrievals, then N answering calls, then a join call. Even with the sub-answers generated concurrently, you have three sequential model stages instead of one. For a live analyst tool this is exactly how a 900ms answer becomes a six-second one — five model calls and five retrievals, most of them on the critical path. Users experience that as a different product. The first conclusion is structural: the extra vector searches are nearly free and the extra model calls are nearly the whole cost. Any optimization that attacks the searches is attacking the wrong term. ## Not every query deserves it The correct default is that most queries do not get transformed. In a typical support or internal-knowledge assistant, the majority of traffic is single-fact lookups already phrased in corpus vocabulary — plain retrieval answers them well, and a transform adds a second of latency for no measurable gain, sometimes for a loss when drifted paraphrases pull in noise. The queries that do justify it share recognizable shapes: - **Compound questions** with multiple clauses, comparisons, or a before/after structure. These need decomposition because no single passage contains all the evidence. - **Terse or jargon-heavy inputs** — a two-word ticket title, an internal abbreviation — where the embedding is noisy and paraphrases materially change what is retrieved. - **High-stakes questions** where a wrong answer costs far more than a slow one: compliance, financial reporting, clinical or legal review. Here seconds are cheap. - **Asynchronous or batch workloads**, where there is no interactive budget at all and you should be more aggressive than in a chat box. ## Gating, not global policy The operational answer is a router. Classify the incoming query cheaply — rules on clause count and question words, a small classifier, or a fast model with a tight output schema — and send only the shapes above through the transform. The classifier must be materially cheaper than what it gates, or you have moved the cost rather than removed it; a rule-based prefilter that catches the obvious single-lookups before any model call is often the highest-leverage piece. Routing also lets you differentiate by *product surface* rather than only by query text. The same engine may back an interactive chat box with a strict budget and a research workflow where a ten-second answer is fine. Those deserve different policies even for identical questions, and hard-coding one global transform pipeline forecloses that. ## Measure the lift per class, not in aggregate Aggregate evaluation hides everything. If fan-out lifts recall substantially on ten percent of queries and does nothing on the rest, the aggregate number looks like a small win and tempts you to apply it everywhere — paying the latency on ninety percent of traffic for nothing. Slice the eval set by the same classes the router uses, and require each class to justify its own transform independently. Log which path each request took so production behaviour can be attributed. Also measure the *distribution*, not the mean. Transformation lengthens the tail more than the median, because a slow planning call and a slow join call compound. A p50 that moves from 900ms to 1.2s may be acceptable while a p95 that moves to eight seconds is not, and only the percentile view reveals that. ## Consider the cheaper levers first Query transformation is one option among several for the same underlying problem, and it is the one that spends the most latency: - **Chunking and metadata.** Many transformation wins are really compensating for chunks that split a rule from its context, or for missing filters. Fixing ingestion costs nothing at query time. - **Hybrid retrieval.** Lexical matching alongside vector search recovers much of what paraphrasing was meant to recover, especially for exact identifiers and rare jargon, and adds negligible latency. - **Larger top-k with a downstream selection stage.** Sometimes retrieving more and choosing better beats searching more times. - **Caching.** Repeated or near-repeated queries can reuse a previously computed plan or result set, removing the planning call entirely on the hot path. A principal-level answer names transformation as a deliberate spend with a measured return on a defined query class — not as a pipeline upgrade applied everywhere because it improved a demo.

  • Why doesn't running the extra retrievals concurrently fix the latency problem?
    Because the searches were never the expensive part. Vector searches cost single-digit to low-tens of milliseconds and already parallelize; the added cost is model calls — planning, per-sub-question answering, joining — and those form a sequential chain where each stage needs the previous stage's output. Parallelism helps within a stage but cannot collapse the stages, so the critical path stays long.
  • How do you keep the routing classifier from becoming its own latency problem?
    Put rule-based checks first so obvious single-fact lookups never reach a model at all — clause count, question form, presence of comparison words and multiple entities catch a large share. Where a model is needed, use a small one with a tight, short output. And bound the classifier's budget explicitly: if it cannot decide fast, default to plain retrieval rather than blocking the request.
  • How would you present the tradeoff to a product owner who wants transformation on by default?
    With per-class numbers and a percentile view. Show which query classes gained measurable accuracy, what the p50 and p95 latency became for each, and what the token cost per thousand requests is. Then propose the gated policy as the way to keep the gain where it exists without charging every user for it, and name the cheaper alternatives — chunking, hybrid retrieval — that were tried first.
  • When would you deliberately transform every query despite the cost?
    When there is no interactive budget or the stakes dominate: offline and batch enrichment, overnight report generation, or review workflows in compliance, financial reporting and legal, where an extra few seconds is invisible and a missed document is expensive. Even then, log which queries the transform actually changed the retrieved set for — if it rarely does, the spend is not buying anything and the classifier is still worth adding.

saying these in an interview costs you the question

  • Blaming the extra vector searches for the added latency
  • Turning transformation on globally because a demo improved
  • Judging the lift only on aggregate evaluation numbers
  • Watching median latency and ignoring the tail
  • Reaching for query rewriting before fixing chunking or hybrid search

context