skip to content

A HyDE draft invents a false drug dosage — does that break retrieval, and when does invention actually hurt?

level: seniorimportance: should knowfreq 45%

answer

  1. two kinds of wrong, very different
  2. wrong number versus wrong subject
  3. the probe commits to one sense
  4. fails loudly? no — fails confidently
  5. anchor the probe with the raw query

basics

~20 s

Usually not. A wrong number with the right terminology still lands the search vector in the correct clinical neighbourhood. Invention hurts only when it moves the draft to a different topic — the wrong specialty, the wrong sense of an acronym — because then you retrieve confidently from the wrong region.

solid answer

~50 s

The draft's factual content barely reaches the embedding; its topic and vocabulary do. So a hypothetical passage that names a plausible-but-wrong dose for a paediatric dermatology drug still sits among real dermatology passages, and retrieval is fine — the draft is discarded and the answer comes from the retrieved documents. The dangerous failure is **domain drift**: the model misreads the query, drafts about cardiology instead of dermatology, and the search vector lands in a coherent but wrong region. Acronym collisions are the classic trigger, since the model must guess a sense before it has seen any corpus evidence. The symptom is nasty because retrieval still returns high-similarity results — it fails silently rather than returning nothing. Mitigations: blend the original query vector with the draft's, keep drafts short, and monitor overlap between HyDE and no-HyDE result sets so drift shows up as a metric rather than a complaint.

go deeper

for a junior

Remember that the draft is thrown away after embedding, so an invented number does not become part of the answer. Know that the search is only as good as the topic the draft landed on.

for a middle

Explain why embeddings tolerate factual errors but not topical ones, and name acronym or polysemy ambiguity as the usual cause of a draft wandering into the wrong subject.

for a senior

Demonstrate that you would measure this: overlap between HyDE and no-HyDE result sets, recall on a labelled slice of ambiguous queries, and human spot-checks. Then name concrete guards such as blending the raw query vector in.

for a principal

Own the risk framing: silent confident retrieval failure in a regulated or safety-relevant domain is a different class of incident from an empty result page, and that should shape whether HyDE ships unguarded, guarded, or not at all.

## Two very different kinds of wrong When people hear that HyDE searches with a model-generated passage, the immediate objection is: the model hallucinates, so the search will be wrong. The honest answer is that hallucination in a HyDE draft splits into two cases with completely different consequences. **Case one: wrong number, right neighbourhood.** Suppose a parent on a rare-disease forum asks about a topical treatment for their child's blistering skin, and the drafted passage confidently states a specific milligram-per-kilogram dose that is simply invented. The draft still reads like a paediatric dermatology passage: it uses the disease name, the anatomical vocabulary, the treatment class. When the encoder compresses it to a vector, the invented number contributes almost nothing to the direction of that vector, while the terminology contributes almost everything. The search lands among genuine dermatology passages. The draft is then discarded, and the real retrieved documents — which carry the real dose — are what the generation step sees. Retrieval was not harmed. **Case two: domain drift.** Now suppose the query is terse and ambiguous, and the model resolves it the wrong way — drafting a passage about a cardiac condition when the question was dermatological. The draft is internally coherent and fluent, but it is *about something else*. Its embedding sits in a different region of the space, and the search returns passages that are highly similar to the probe and entirely irrelevant to the user. This is the failure mode that matters. ## Why drift fails silently A lexical search that goes wrong tends to return nothing, which is a loud signal. A dense search with a drifted probe returns a full page of confidently-scored results. The similarity scores look healthy because the probe and the retrieved passages genuinely are similar — to each other. Nothing downstream can tell that the probe was about the wrong subject. If the generation step is instructed to answer from context, it will produce a fluent answer grounded in irrelevant documents, which is worse than an obvious empty result. ## The acronym-collision trigger Drift is most commonly triggered when the query contains a term whose sense depends on the corpus, and the model must guess that sense with no corpus evidence available. Acronyms are the worst offenders because the same three letters carry unrelated meanings across industries, and the drafting model picks whichever sense dominated its training data rather than whichever sense your index uses. Ordinary polysemous words behave the same way: a single word can pull an entire drafted paragraph into the wrong professional register. Once the draft commits to a sense, the whole vector commits with it — expansion amplifies the error instead of hedging it. ## Detecting it Because the symptom is confident wrongness, you have to look for it deliberately: - **Result-set overlap.** Run the same query with and without HyDE and measure how much the top-k sets intersect. Near-zero overlap on a query where the plain query already worked is a drift alarm, not a win. - **Offline recall on a labelled set.** Keep a held-out set of queries with known relevant documents and track recall with HyDE on and off. Report it per segment, because HyDE can lift recall overall while destroying it on the ambiguous-acronym slice. - **Score distribution is not a signal.** Do not use similarity scores as a drift detector; drifted probes score well by construction. ## Mitigating it Several cheap guards exist, and they compose: - **Blend with the original query.** Average the draft's embedding with the raw query's embedding, or concatenate the query text into the drafted passage before embedding. The user's own words act as an anchor that limits how far a drifting draft can pull the probe. - **Average several drafts.** Sampling a few hypothetical passages and averaging their vectors makes a single bad draft a minority vote rather than the whole probe. - **Keep drafts short.** A one-paragraph draft has less room to wander into a second subject than a long one. - **Constrain the prompt with domain context.** Telling the drafting model what kind of corpus it is writing for — a clinical archive, an HR policy set, an incident-response knowledge base — collapses most acronym ambiguity before it starts. - **Fall back.** If the HyDE-retrieved set has near-zero overlap with the plain-query set, retrieving with both and taking the union is a reasonable safety net for high-stakes domains. ## The framing an interviewer wants The strong answer is not "hallucination is fine" or "hallucination breaks it" — it is the distinction. Factual errors inside the right topic are tolerated by the geometry; topical errors are not, and they fail quietly. That distinction also explains why HyDE is safer than it sounds in specialist domains where the model knows the register but not the specifics, and riskier than it sounds in domains full of internal acronyms.

  • How would you detect domain drift in production without labelled relevance judgments?
    Compare the HyDE result set against the plain-query result set for the same request and alert on near-zero overlap, especially on queries the plain path already handled. Sample those cases for human review. Similarity scores are useless here — a drifted probe scores high by construction — so overlap and human spot-checks are the practical signals.
  • Why does averaging the draft's embedding with the original query's embedding help?
    It anchors the probe. The user's literal words are the one part of the pipeline that cannot have drifted, so blending them in bounds how far a bad draft can move the search vector. It costs nothing at inference and mainly trades a little of HyDE's lift for a lot of worst-case protection.
  • Does a factually wrong draft risk contaminating the final answer?
    Not directly, because the draft is discarded after embedding and never enters the generation context. The indirect risk is that a drifted draft retrieves irrelevant passages, and the model then grounds a fluent answer in them. Guarding the retrieval step, not the draft's accuracy, is what protects the answer.

saying these in an interview costs you the question

  • Assuming any hallucination in the draft ruins retrieval
  • Treating high similarity scores as proof the probe was right
  • Feeding the hypothetical draft into the answer as context
  • Ignoring acronym ambiguity in domain-specific corpora
  • Expecting drift to show up as empty result sets

context