How do you detect a query rewriter that adds constraints the user never stated?
answer
- the rewrite is itself a generation
- two directions, not one
- check every added term is grounded
- hedge by searching with both strings
- recall on traces, not answer scores
basics
~20 sAudit the rewrite, not just the answer. Log every original-and-rewritten pair, then check whether each added constraint — a date, a booking, a product name — actually appears in the conversation. Unsupported additions silently shrink recall.
solid answer
~50 sRewriters fail in two symmetrical ways. **Over-specification**: the model invents a constraint nobody stated — a date range, a cabin class, a version number — and the narrowed query no longer matches the passage that would have answered the question. **Under-specification**: it drops a constraint the user did state, so a question about the 19:40 Lisbon flight retrieves generic policy text. Both are silent; the pipeline returns confident answers to a subtly different question. Detection starts with logging the (history, original turn, rewrite) triple and auditing samples — a rule or judge checks whether every entity in the rewrite is grounded in the conversation. Prevention: instruct the rewriter to use only terms present in the history or the turn, keep rewrites short, and hedge by retrieving with both the raw turn and the rewrite and merging. Measure recall@k on multi-turn traces, since answer-level evals hide the cause.
go deeper
Know that the rewriting step is itself a model call and can get things wrong, and that the safe habit is to log both the user's original wording and the rewritten query.
Explain both directions — added constraints that narrow retrieval and dropped constraints that broaden it — and describe a concrete check that every entity in the rewrite appears somewhere in the conversation.
Demonstrate operational instinct: online grounding guardrails, judge sampling calibrated against human labels, recall@k on multi-turn traces, and hedging by retrieving with both the raw turn and the rewrite.
Frame it as risk placement. Decide where an unfaithful rewrite is tolerable and where it is not, insist that rewrite-derived filters carry less authority than user-supplied ones, and own the eval gate that any rewriter model change must pass.
## Why this failure is worth its own question Standalone-question rewriting is usually presented as a pure win: resolve the pronoun, retrieve better. In production it is a generative step sitting upstream of everything else, and generative steps hallucinate. Because the rewrite is invisible to the user and usually invisible in the answer, a broken rewriter produces the worst kind of bug — plausible answers to a question nobody asked, with no error anywhere in the trace. ## The two directions of failure **Over-eager rewriting (over-specification).** The model helpfully adds detail. In an airline support thread it turns "what about the later one?" into "What is the change fee for the 19:40 Lisbon flight on booking QX7742 in economy for a same-day change?" — where cabin class and "same-day" were never mentioned. Every invented constraint is a term the retriever now weights. In a dense index the embedding drifts toward same-day-change documents; in a hybrid or filtered setup it can be worse, because an invented value may be lifted into a metadata filter and eliminate the correct document outright. Recall drops and nothing logs an error. **Under-eager rewriting (under-specification).** The mirror image: the user did name the 19:40 flight, and the rewrite collapses to "What is the change fee?". Retrieval returns the generic fee policy, the generator answers confidently, and the user is told a number that does not apply to their booking. This variant is more dangerous, because a plausible generic answer looks correct to everyone including the evaluator. A third, subtler variant is **entity drift**: with a long history window the rewriter grabs the entity from an abandoned topic — the user switched from flights to baggage three turns ago, and the rewrite still says QX7742's Lisbon leg. ## Detecting it The key move is to make the rewrite a first-class logged artefact. Store the triple (history window, original turn, rewritten query) for every request, keyed to the retrieval result and the final answer. Then: - **Grounding check.** For each noun phrase, number, date or identifier in the rewrite, assert that it appears in the history or the current turn, allowing for lemmatization and simple normalization. Anything ungrounded is a candidate hallucination. This is cheap enough to run online as a guardrail, not just offline. - **Constraint-preservation check.** The inverse: identifiers and qualifiers that appeared in the user's turn should survive into the rewrite. A dropped booking reference is easy to catch mechanically. - **Judge sampling.** A model judge scores a sample of pairs on faithfulness — did the rewrite change the meaning of the turn given the history? Judges are noisy, so calibrate against a small human-labelled set before trusting the numbers. - **Retrieval-level metrics.** Score recall@k on multi-turn traces with known answer passages. If recall drops on turns where the rewrite added tokens, you have your signal. Answer-level metrics alone will not localise the cause. - **Length and delta heuristics.** Rewrites that are dramatically longer than the turn plus its referents are suspicious; simple token-delta alerting surfaces a drifting prompt or a model upgrade regression quickly. ## Preventing it Constrain the task rather than trusting the model's judgement. Instruct the rewriter to resolve references and restore omitted predicates **using only wording present in the conversation**, and to prefer copying the user's phrasing over paraphrase. Show it what to do when there is nothing to resolve — return the turn unchanged — because a rewriter with no null action will invent work. Cap the history window so old entities cannot leak in, and add a topic-shift signal for long sessions. Architecturally, hedge. Retrieving with both the raw turn and the rewrite and merging the two result lists means a bad rewrite costs latency instead of the right document; it is a cheap insurance policy for high-stakes domains. Where the pipeline applies metadata filters, treat filter values derived from a rewrite as lower-trust than values the user typed — an invented date range that becomes a hard predicate removes the answer with no recourse, whereas an invented term in the free-text query only reweights it. ## What good looks like in an interview Say that the rewriter is a model in the loop and must be evaluated like one; name both failure directions rather than only hallucinated additions; propose a grounding check that a non-ML engineer could implement; and be clear that the metric that exposes this is retrieval recall on conversational traces, not end-to-end answer quality. Then note the asymmetry: dropped constraints usually hurt users more than added ones, because they yield confident, generic, wrong answers.
- Which is worse in practice — an invented constraint or a dropped one?Usually the dropped one. An invented constraint narrows retrieval and often produces an obvious miss or a refusal. A dropped constraint retrieves the generic document, which reads perfectly well and yields a confident answer that simply does not apply to the user's case — a wrong answer nobody flags.
- Why is an invented value especially dangerous when it lands in a metadata filter?A free-text term only reweights ranking; the right passage can still surface. A filter is a hard predicate — an invented date range or category removes the answering document from the candidate set entirely, and no amount of top-k or reranking recovers it. Treat filter values derived from a rewrite as lower-trust than values the user typed.
- How would you catch a rewriter regression after a model upgrade?Keep a frozen conversational eval set with reference standalone questions and known answer passages, and run it on every rewriter change: recall@k, a grounding-violation rate, and a token-delta distribution. A silent regression shows up as more ungrounded terms and longer rewrites before it shows up in end-to-end answer scores.
saying these in an interview costs you the question
- Assuming a rewrite can only be too vague, never too specific
- Evaluating only the final answer and never the rewritten query
- Not logging the original turn alongside the rewrite
- Letting rewriter-invented values become hard metadata filters
- Giving the rewriter no way to leave a self-contained turn unchanged