Why does a RAG chat need history-aware query rewriting before retrieval?
answer
- the retriever has no conversation
- pronouns are not searchable
- two phenomena: reference and omission
- rewrite for the index, not the user
- measure recall on multi-turn traces
basics
~20 sRetrieval sees only the query string. A follow-up like "what about the later one?" carries no searchable content, because its referent sits in earlier turns, so a rewriter must turn it into a standalone question first.
solid answer
~40 sA retriever maps one string to a ranked list of passages; it has no memory of the thread. Users, though, write the shortest turn that is unambiguous to a human reading above it, so follow-ups are full of coreference ("the later one", "that fee") and ellipsis (the verb is simply missing). Embedded, such a turn lands nowhere useful; lexically it contains no rare term at all. History-aware rewriting fixes this: a small model sees the recent turns plus the current one and emits a self-contained question — "What is the change fee for the 19:40 Lisbon flight on booking QX7742?" — that names every referent. Crucially the rewrite is a retrieval artefact only: it goes to the index, while the generation prompt still receives the user's actual words and the real conversation.
go deeper
Be ready to say plainly that search only sees the text you send it, so "what about the later one?" has to be turned into a full question naming the flight and the booking before it can retrieve anything.
Explain the mechanics: coreference and ellipsis in the user's turn, a small model that reads recent history and emits a standalone question, and the fact that the rewrite goes to the index while the original still drives generation.
Show you have operated it — history-window tuning, topic-shift contamination, skipping the rewriter on first or self-contained turns, and evaluating recall@k on multi-turn traces rather than single questions.
Own the tradeoff: an extra model call on the critical path buys most of the quality gap between demo and production chat. Argue where it belongs (merged with routing, distilled into a small model) and what you budget for its latency.
## What the retriever actually sees A retriever — dense, sparse or hybrid — is a function from **one string** to a ranked list of passages. It has no access to the conversation. Everything the chat UI shows above the input box is invisible to it unless you put it into the query text. The moment a product supports follow-up turns, this creates a mismatch: people write the shortest thing that is unambiguous *to a human reading the thread*, and that is almost never a good search query. Take an airline support chat. The user asks about their Lisbon options on booking QX7742, the assistant lists a 12:05 and a 19:40 departure, and the user types "what about the later one?". As a search query that string is close to worthless. It contains no rare term for a keyword index to match, and its embedding sits in a region populated by every other vague follow-up in the corpus. Top-k comes back as noise or as generic policy boilerplate, and the generator then either refuses or invents a fee. ## Coreference and ellipsis Two linguistic phenomena cause this, and naming them well in an interview is worth real credit. **Coreference** is a pronoun or definite description pointing at an entity introduced earlier: "it", "the later one", "that fee", "the second option". Resolving it means replacing the reference with the entity itself. **Ellipsis** is material the user omits entirely. "What about the later one?" drops the whole predicate — the user never repeats "what is the change fee for". Reconstructing it means recovering intent from the previous turns, not just substituting a noun phrase. Humans do both effortlessly from the thread. A retriever does neither. History-aware rewriting — also called standalone-question rewriting or contextual query reformulation — is the repair step: a model receives the last N turns plus the current user message and emits one self-contained question. Good rewrites are explicit and specific: entity names, identifiers, dates and the missing predicate all restored. ## Two queries, two jobs A detail weaker candidates miss: the rewrite does **not** replace what the user said. It is an internal retrieval artefact. The rewritten query goes to the index; the generation prompt still carries the genuine conversation and the user's own wording. If you substitute the rewrite into the visible dialogue, two things go wrong — the assistant starts answering a question the user did not literally ask, and every rewriter error becomes user-visible rather than merely degrading recall. Keep the pair, log the pair, and let the generator see the original. ## When to rewrite, and what it costs Rewriting sits on the critical path and costs an extra model call before retrieval can even start. Common economies: skip it on the first turn of a session, where there is no history to resolve; use a small fast model, since the task is mechanical; cap the history window; and, where the pipeline also routes queries, fold rewriting and routing into a single call so you pay one round trip instead of two. Some systems gate the rewriter behind a cheap check — if the turn contains no pronoun, no definite "the X" and looks self-contained already, pass it through untouched. The history window itself is a tuning knob with a failure mode in both directions. Too short and you lose the entity the user is referring to. Too long and the rewriter drags in stale entities from a topic the user abandoned ten turns ago, producing a query about a booking the conversation has moved on from. Systems that handle long sessions usually add a topic-shift signal or summarize older turns rather than feeding raw history. ## How you know it is working Evaluate on **multi-turn conversations**, not single questions — a single-turn eval set cannot see this failure at all. Build a golden set of (history, current turn, expected standalone question) triples drawn from real traces, and score two things: rewrite quality against the reference, and, more importantly, retrieval recall@k measured on the rewritten query against the passages that actually answer the turn. A cheap online guardrail is an unresolved-reference check: if the rewritten query still contains a bare pronoun or a phrase like "the other one", the rewriter did not do its job and you can log or retry. The honest framing for an interview is that this is one of the cheapest, highest-yield fixes in conversational RAG. Teams typically discover it after shipping a single-turn prototype that scores well in evals and then falls apart in production the moment users start asking follow-ups.
- Which string do you show the user, and which one goes to the index?Only the index sees the rewrite. The user's own turn stays in the visible dialogue and in the generation prompt, so the assistant answers what was actually asked. Keeping them separate also means a bad rewrite degrades recall quietly instead of putting words in the user's mouth — and logging both halves of the pair is what later lets you audit the rewriter.
- How much conversation history should the rewriter see?Enough to resolve the current reference, rarely more. A short window (a few turns) covers most coreference; a long raw window invites the rewriter to import entities from an abandoned topic, producing a confidently wrong standalone query. For long sessions, summarize older turns or add a topic-shift signal rather than growing the window.
- Would you rewrite on the very first turn of a session?No — there is nothing to resolve, so it is pure added latency and risk. Skip it when there is no history, and optionally skip it when the turn already looks self-contained (no pronouns, no bare definite references). The rewriter earns its call only on turns that actually depend on prior context.
saying these in an interview costs you the question
- Claiming the retriever can see the chat history
- Replacing the user's message with the rewrite in the dialogue
- Passing the whole conversation to the embedding model as the query
- Evaluating a conversational RAG system on single-turn questions only
- Thinking a bigger top-k fixes an unresolved pronoun