How does pseudo-relevance feedback expand a query, and when does it drift off topic?
answer
- two passes, no human in the loop
- the top results are only assumed relevant
- expansion vocabulary comes from the corpus
- one bad first pass compounds
- keep the original query dominant
basics
~20 sPseudo-relevance feedback runs a first search, assumes the top few results are relevant, pulls terms or a vector centroid from them, and re-searches with the enriched query. It drifts when those top results are off topic, because the loop amplifies the mistake instead of correcting it.
solid answer
~50 sPseudo-relevance feedback (PRF) is a two-pass technique. Pass one retrieves with the user's query. Pass two assumes — without any human judgment — that the top-k results are relevant, and reformulates the query toward them: classically the Rocchio method moves the query vector toward the centroid of those documents and away from ones presumed non-relevant, and term-based variants harvest their most distinctive terms. Then it searches again. The appeal is that the expansion vocabulary comes from your actual corpus, so it speaks the index's language for free — no generation call, no external model. The risk is **query drift**: the assumption of relevance is unverified, so if two of the top five documents are off topic, the second query inherits their vocabulary and the results get steadily worse with high confidence. Control it with a small feedback set, a modest weight on the feedback term relative to the original query, and a score floor below which no feedback is taken at all.
go deeper
Know the shape: search once, treat the top few hits as if they were relevant, enrich the query from them, search again. Remember that nobody confirmed those hits were actually relevant.
Explain the Rocchio combination of original query and feedback centroid, name query drift as the failure mode, and give the standard controls — small k, conservative weight, a single iteration.
Show that you would gate it: a score floor so a failed first pass never feeds the second, metadata filtering of the feedback set, and offline measurement of recall with feedback on and off per query segment.
Position it against generation-based expansion on a cost and risk basis — a second index lookup versus a model call, corpus-sourced vocabulary versus parametric vocabulary — and decide which failure your domain can better absorb.
## The idea Pseudo-relevance feedback comes from classical information retrieval and long predates language models. True relevance feedback asks a human: you show results, the user marks which ones were useful, and you reformulate the query toward the marked ones. It works well and nobody does it, because users will not label results. Pseudo-relevance feedback keeps the mechanism and drops the human: it simply *assumes* the top-ranked results of the first search are relevant and feeds them back automatically. The loop is: 1. Search with the user's original query. 2. Take the top-k results (k is small — often 3 to 10). 3. Extract signal from them: distinctive terms, or a centroid vector. 4. Build a reformulated query combining the original with that signal. 5. Search again and return the second pass's results. ## Rocchio, concretely The canonical formulation is the Rocchio algorithm. It expresses the new query as a weighted combination of three things: the original query vector, the centroid of the documents presumed relevant, and (negatively) the centroid of documents presumed non-relevant. Three weights control the balance. In the pseudo variant, "presumed relevant" is just the top-k, and the non-relevant term is often dropped or given a very small weight because guessing non-relevance from a ranked list is even shakier than guessing relevance. The practical consequence of the weighting is that you keep the original query dominant. If the feedback centroid outweighs the user's own query, one bad first pass can replace the user's intent entirely. ## A worked scenario Consider an agricultural extension-service archive: decades of advisory bulletins for farmers. A grower searches for guidance on a soil condition using regional, informal phrasing. The first pass returns five bulletins, of which three are squarely on the soil problem and two are about unrelated machinery maintenance that happened to share a couple of words. Re-weighting on all five drags the query toward machinery vocabulary. The second pass now returns more machinery bulletins, which look highly relevant to the reformulated query and are useless to the grower. Nothing in the loop can detect this — the system has no ground truth, only its own first guess. Had the same loop run on a first pass where all five bulletins were on topic, the expansion would have added exactly the regional agronomy vocabulary the archive uses and the second pass would have been notably better. That is the whole bargain: PRF amplifies whatever the first pass gave you, in both directions. ## Controlling drift The standard levers: - **Small k.** Feeding back three documents rather than twenty limits how much off-topic mass can enter the centroid. - **Conservative weighting.** Keep the original query's weight high so feedback nudges rather than replaces. - **A score floor.** If the first pass's top results are all weakly scored, that is evidence the query failed; taking feedback from weak results is the worst case. Skip feedback entirely below a threshold. - **Filter the feedback set.** Drop documents that disagree with each other, or restrict feedback to documents sharing metadata (same section, same crop, same year range) with the majority of the top results. - **Single iteration.** Running the loop repeatedly compounds drift quickly; one round is the usual practice. ## How it compares with generating a hypothetical passage Both techniques attack the same underlying problem — the user's query does not speak the corpus's language — but from opposite directions. - PRF sources its expansion vocabulary from **inside the corpus**, after a first retrieval. It cannot invent a term the corpus does not contain, which is a real safety property. Its cost is a second retrieval round, which is cheap. - Generating a hypothetical answer sources vocabulary from the **model's parametric knowledge**, before any retrieval. It can supply terminology the first pass would never have surfaced, which is exactly what you want when the first pass returns nothing useful. Its cost is a generation call, which is much more expensive than a second index lookup. That suggests the natural split: PRF is a good default when the first pass is decent and you want to sharpen it; generation-based expansion earns its cost when the first pass is hopeless because the user and the corpus share no vocabulary at all. They can also be combined, though each added stage multiplies the ways the pipeline can quietly go wrong. ## What to say in an interview Name the assumption out loud — "pseudo" means nobody verified relevance — and name query drift as the failure it creates. Then give the controls: small feedback set, conservative weight, score floor, one iteration. That sequence shows you understand it as a bet with a known downside rather than a free improvement.
- What distinguishes pseudo-relevance feedback from true relevance feedback?True relevance feedback uses human judgments — a user marks which results were useful — so the feedback signal is verified. Pseudo-relevance feedback substitutes an assumption: the top-k of the first pass are treated as relevant with no confirmation. That substitution is what makes it deployable at scale and what creates query drift.
- Why is it usually wise to skip feedback when the first pass's top scores are all low?Low top scores are evidence the first query failed to find its neighbourhood. Expanding toward those documents imports vocabulary from material that is probably irrelevant, so the second pass is likely worse than the first. A score floor turns that case into a no-op, which is the safe default.
- Compared with generating a hypothetical passage, what is pseudo-relevance feedback's structural advantage?Its expansion terms come from documents that actually exist in the index, so it cannot introduce vocabulary the corpus has never used, and it needs no generation call — only a second index lookup. The trade is that it can only reinforce what the first pass already found; it cannot bridge a vocabulary gap the first pass never crossed.
saying these in an interview costs you the question
- Believing the feedback documents were verified as relevant
- Iterating the loop many times to 'converge'
- Weighting the feedback centroid above the original query
- Taking feedback from a weak, low-scoring first pass
- Confusing it with reordering the results you already have