What is HyDE, and why does embedding a hypothetical answer beat embedding the question?
answer
- queries and documents are different genres
- move the query into answer space
- search with a fabricated passage
- the draft is discarded after embedding
- topic survives compression, facts do not
basics
~20 sHyDE (Hypothetical Document Embeddings) asks a model to draft a plausible answer to the query, then embeds that draft instead of the question. Answers share vocabulary and phrasing with real documents, so the search vector lands nearer the right passages.
solid answer
~50 sDense retrieval puts queries and documents in one vector space, but they are different genres of text: a query is short, interrogative and often plain-language, while a passage is long, declarative and full of domain terminology. That mismatch is question-space vs answer-space asymmetry, and it costs recall. HyDE closes it by prompting an LLM to write the passage that *would* answer the query, embedding that fabricated passage, and searching with its vector. The draft is thrown away — it is never shown to the user and never used as an answer, only as a search probe. It works because the encoder is largely insensitive to the draft's specific facts; what survives compression is topic and terminology, which is exactly what needs to match. Gains are biggest in zero-shot settings where you have no labelled query-passage pairs to tune retrieval on.
go deeper
Be able to say plainly what HyDE stands for and the three steps: generate a fake answer, embed it, search with it. Remember that the fake answer is thrown away.
Explain the asymmetry that motivates it — short plain-language questions versus long jargon-heavy passages — and why an embedding tolerates a factually wrong draft. Name the zero-shot, no-labelled-data setting as its home ground.
Show judgment about when it does not pay: FAQ-shaped corpora, keyword lookups, encoders already tuned for question-to-passage matching. Be ready to describe how you would measure lift rather than assuming it.
Frame HyDE as one of several places to spend on the asymmetry — pipeline-side generation versus a retrieval model that handles it internally — and be ready to argue which one your team should own given latency budget and corpus churn.
## The problem HyDE attacks Dense retrieval works by pushing both the user's query and every document chunk through an embedding model, then finding the chunks whose vectors sit closest to the query vector. The implicit assumption is that a question and its answer land near each other. Often they do not. Imagine a rare-disease patient forum. A parent types: *"why does my daughter's skin blister whenever we change her nappy?"* The literature that actually answers this says something like: *"epidermolysis bullosa is a mechanobullous disorder characterised by dermal-epidermal separation following minor friction or trauma."* The two texts share almost no vocabulary, no register and no length. One is a short first-person interrogative; the other is a long third-person declarative packed with clinical terminology. This is **question-space vs answer-space asymmetry**: questions cluster with other questions, passages cluster with other passages, and the bridge between the two clusters is exactly what a general-purpose encoder is weakest at. ## What HyDE does HyDE — Hypothetical Document Embeddings, introduced in 2022 as *Precise Zero-Shot Dense Retrieval without Relevance Labels* — attacks the asymmetry by moving the query into answer space before searching: 1. Prompt an instruction-following model: "write a passage that answers this question." 2. Take the generated passage — the *hypothetical document*. 3. Embed that passage with the same encoder used for the corpus. 4. Search the index with that vector. 5. Discard the draft. Step 5 is the part candidates most often miss. The hypothetical document never reaches the user and is never used as grounding. It exists only to produce a better probe vector. The final answer is still generated from the real retrieved passages. A common refinement is to sample several drafts and average their embeddings, sometimes together with the original query's embedding, into a single search vector. That is still one retrieval — it is not a fan-out of several searches whose result lists get merged. ## Why a fabricated passage works The intuition that trips people up is: if the model does not know the answer, how can its invented answer help? Because the embedding is a lossy compression. What survives compression is the topical and terminological signal — the fact that the passage is about blistering skin, friction-induced dermal separation, paediatric dermatology. What mostly does not survive is the fine-grained factual detail: an invented statistic, a fabricated citation, a made-up dosage. The draft is judged by the *neighbourhood* it lands in, not by its truth value. Its job is to speak the corpus's language, and an LLM is good at producing the register of a domain even when shaky on its particulars. ## Where it pays, and where it does not HyDE helps most when: - You are **zero-shot** on a new corpus, with no labelled query-passage pairs to adapt retrieval with. - There is a genuine **vocabulary gap** between how users write and how the corpus is written — laypeople vs specialists, natural language vs internal jargon, one language vs another. - Queries are **short or underspecified**, giving the encoder very little signal to work with. It helps least when: - The corpus is already written as questions — FAQs, support tickets with question-shaped titles, forum threads. There the query already sits in the right space, and the drafted passage can move it away. - The retrieval encoder was already trained on question-to-passage matching, so the asymmetry is handled inside the model instead of in the pipeline. - Queries are essentially keyword lookups (an error code, a part number) where the literal string is the strongest signal and the draft dilutes it. ## What it does not fix HyDE is a **pre-retrieval** technique. It changes what you search with; it does not change what is in the index or how results are ordered afterwards. If the answer is not in the corpus, if chunks are split badly, or if the top results are poorly ordered, HyDE will not save you. It also does not make retrieval cheaper — it adds a generation call on the critical path of every query it is applied to, which is why teams tend to apply it selectively rather than universally. ## Practical shape The drafting prompt is usually short and domain-specific: telling the model to write in the style of the corpus (a clinical abstract, a policy document, a runbook) tends to matter more than telling it to be correct. Keeping the draft short — a paragraph, not an essay — keeps latency and cost down and avoids diluting the topic. Many implementations concatenate or average the original query so that a badly drifting draft cannot fully hijack the search.
- Is the hypothetical document ever shown to the user or used as grounding for the final answer?No. It is a search probe only. It is embedded, used to retrieve, and thrown away; the answer is generated from the real passages that came back. Treating the draft as evidence would defeat the point of retrieval entirely, since it is unverified model output.
- Some implementations generate several hypothetical documents. How is that different from issuing several searches?The drafts' embeddings are averaged into a single vector, so there is still exactly one retrieval and one result list. Averaging smooths out an individual draft that drifts off-topic. That is different from running separate searches per variant and merging the returned lists, which is a distinct pattern with different cost and merge semantics.
- On which corpora would you expect HyDE to give little or no lift?Corpora already written in question form — FAQ banks, support-ticket titles, Q&A forums — because the user's query already sits in the same space as the indexed text. Also keyword-style lookups such as error codes or SKUs, where the literal token is the strongest signal and a generated paragraph dilutes it.
It is like searching a medical library by first writing the paragraph you hope to find, then asking the librarian for shelves that look like that paragraph — rather than reading her your question.
saying these in an interview costs you the question
- Claiming HyDE's draft is shown to the user as the answer
- Assuming the draft must be factually correct to help
- Thinking HyDE changes the index or the documents
- Confusing it with reordering results after retrieval
- Believing it helps equally on every corpus