When does HyDE's per-query generation cost stop paying, and what would you reach for instead?
answer
- a generation call on the critical path
- serial latency before retrieval even starts
- pay per query, or pay once
- route it, do not blanket-apply it
- measure lift per segment, not in aggregate
basics
~20 sHyDE stops paying once the retrieval encoder already handles the question-to-passage gap, because you are then buying a generation call and its latency for lift you already have. The alternatives are a retrieval model built for that asymmetry, selective routing, and caching.
solid answer
~50 sHyDE puts an LLM call on the critical path of every query it touches: hundreds of extra tokens, extra latency before retrieval can even start, and one more component that can fail or drift. That is a real price on an interactive system, and it is worth paying only while the lift is real. The same asymmetry between question text and passage text can also be handled inside the retrieval model itself — encoders built for question-to-passage matching absorb much of the gap at no per-query cost — so the honest framing is two answers to one problem: pay per query at inference, or pay once in the retrieval stack. As of mid-2026, most teams get the bulk of the benefit from a stronger retrieval setup and reserve HyDE for the queries that need it: specialist vocabulary the encoder never saw, terse queries, or corpora where measured recall without it is poor. Routing HyDE selectively and caching drafts for repeated queries keeps the cost proportional to the benefit.
go deeper
Remember the basic cost: HyDE adds an LLM call before every search it touches, so queries get slower and more expensive even when the answer would have been found anyway.
Explain that the generation step is serial ahead of retrieval, so its latency is fully visible to the user, and that a smaller drafting model plus caching removes much of the cost without losing the benefit.
Show the routing design: cheap retrieval first, escalate to a drafted probe only on weak evidence, cache the head of the query distribution, and measure recall lift and p95 latency per segment before and after.
Own the framing as a build-versus-recur decision — a recurring per-query inference cost against a one-time investment in the retrieval stack — and be willing to say plainly that as of mid-2026 there is no universal answer, only measurements that decide it for your corpus and volume.
## What HyDE actually costs The technique's expense is easy to under-count because it looks like one extra step. Per query it adds: - **A generation call before retrieval.** Not after, not in parallel — retrieval cannot start until the draft exists, so its full latency is serial with everything downstream. - **Output tokens.** A paragraph-length draft, billed at generation rates, for every query. - **A failure surface.** The drafting model can be rate-limited, slow, or resolve an ambiguous term into the wrong subject, and the last case fails quietly by retrieving confidently irrelevant material. - **An operational dependency.** Retrieval now depends on a model endpoint, which changes your availability story and your capacity planning. On a batch pipeline none of this matters much. On an interactive assistant where users are watching a cursor blink, adding a serial generation step before the first byte of grounded output is a visible product regression that has to be earned back. ## The two answers to the same problem The asymmetry HyDE addresses — queries and passages being different genres of text — has two structurally different remedies. **Pay per query, at inference.** Generate a hypothetical passage each time, moving the query into passage space on the fly. Requires no training, no labelled data and no re-indexing; you can turn it on this afternoon. Cost recurs forever, per request. **Pay once, in the retrieval stack.** Use a retrieval model designed for question-to-passage matching, which handles the asymmetry internally. The per-query cost is then just an embedding lookup. The price is paid up front — selecting or adapting the model, re-embedding the corpus, validating the change — and again whenever the corpus or the model changes. Neither dominates. The first is the right call when you are exploring, when the corpus is small or volatile, or when you cannot re-index. The second is the right call once the system is load-bearing and query volume makes a recurring generation call the dominant cost line. A team that ships HyDE and never revisits it is often paying inference forever to avoid a one-time retrieval investment. ## Making the cost proportional: selective application The strongest production pattern is not on-or-off but **routing**. Apply the expensive path only to queries that plausibly need it: - **Terse or vague queries.** A three-word query gives the encoder little to work with; a long, well-phrased one usually does not need help. - **Vocabulary-gap signals.** Queries containing internal jargon, unfamiliar abbreviations, or laypersons' phrasing against a specialist corpus. - **Weak first-pass evidence.** Retrieve cheaply first; if the top results are weakly scored or thin, *then* spend the generation call and retrieve again. This inverts the pipeline so the expensive step only runs where the cheap one visibly failed, and it caps the added latency to the subset of queries that were going to fail anyway. - **Segment measurement.** Route by what the offline evaluation shows, not by intuition — measure recall with and without HyDE per query segment, because an aggregate lift often hides one slice improving a lot and another degrading. **Caching** is the other lever. In most real systems the query distribution has a heavy head: the same questions recur across users. Caching the draft (or the resulting probe vector) keyed on a normalised query removes the generation cost for repeats entirely. A smaller, faster model for drafting is a third lever — drafting needs the right register and terminology, not deep reasoning, so the largest available model is rarely the right choice for this step. ## What to measure before and after Any argument about whether HyDE pays should rest on four numbers, not on the technique's reputation: 1. **Recall (or hit rate) at your working k**, with HyDE on and off, per query segment. 2. **Added end-to-end latency at the tail**, not the mean — the p95 is what users feel. 3. **Marginal cost per query**, including the drafting model's output tokens. 4. **Drift rate**: how often the HyDE result set diverges sharply from the plain-query result set on queries the plain path already handled. If recall lift is concentrated in one segment, route to it. If it is broad but small and latency is tight, the retrieval-side investment is the better use of the same effort. If it is broad and large, HyDE is genuinely earning its cost and the question becomes how to make it cheaper rather than whether to keep it. ## The honest state of practice As of mid-2026 this remains a judgment call rather than a settled answer. Retrieval models have improved enough that the naive framing — always generate a hypothetical passage — is hard to defend as a default on high-volume interactive systems. But specialist corpora with vocabulary far outside any general encoder's training distribution still show real, measurable lift, and there the per-query cost is straightforwardly worth it. The defensible position in an interview is to describe the trade, name the measurements that decide it, and refuse to give a universal answer.
- How would you structure a pipeline so HyDE only runs where it is needed?Retrieve cheaply with the raw query first. If the top results are weakly scored or too few clear the threshold, spend the generation call and retrieve again with the drafted probe. This caps the added latency to queries that were already failing, and gives you a natural signal for how often the expensive path is actually being used.
- Why is caching unusually effective for this particular step?Query distributions are heavy-headed — a small number of questions account for a large share of traffic. Caching the draft or its embedding, keyed on a normalised query, removes the generation cost for every repeat. It works because the draft depends only on the query, not on the user or the current index state.
- Should the drafting step use your strongest model?Usually not. The draft needs the right topic and terminology, not correctness or reasoning depth, and its factual content barely reaches the embedding. A smaller, faster model produces a probe of nearly equal quality at a fraction of the latency and cost, which is exactly what a step on the critical path wants.
- What single number would most change your mind about keeping HyDE?Recall lift at your working k, measured per query segment against the plain-query baseline. If the lift is broad and large, keep it and optimise its cost; if it is confined to one segment, route rather than blanket-apply; if it is small and your p95 latency is tight, invest the same effort in the retrieval model instead.
saying these in an interview costs you the question
- Treating HyDE as a free accuracy improvement
- Ignoring that the generation call is serial before retrieval
- Applying it to every query regardless of query shape
- Judging it on aggregate lift while segments cancel out
- Using the largest model to draft a throwaway probe