A reranking stage is added to a RAG retrieval path — which planted passages does it promote?
answer
- two families, opposite outcomes
- one is judged as an answer now
- gibberish loses, good prose wins
- relevance is not provenance
basics
~20 sA reranking stage drops planted text shaped against an embedding model, which reads as nonsense, and promotes planted text written to read as the ideal answer. It scores relevance to the query, never who wrote the passage.
solid answer
~50 sThere are two families of planted passage, and the second stage treats them oppositely. One is text shaped against the embedding model so its vector lands near an expected query; it wins a distance contest but does not read as an answer, so a stage that rescores each candidate against the query pushes it out of the top few. The other is an ordinary, well-written passage composed to read as the confident, complete answer to a question the attacker expects someone to ask — and raising exactly that candidate is what the reranking stage is for. So the stage does not stop corpus poisoning; it selects which kind survives. Nothing in its score corresponds to who wrote the passage or how the record entered the corpus: relevance is not provenance. An attacker who learns a second stage is present simply stops using the first family.
go deeper
Be ready to say what a reranking stage scores on — how well a passage answers the query — and that it has no signal for who wrote the passage. Knowing that one plant family dies there and another is promoted is the whole answer at this level.
Explain why the two families come apart at that stage: one was built for a distance metric, the other for a reader. Say plainly that the stage changes which construction is worth writing rather than removing planted content.
Show you can state what a first-place rank proves and what it does not — candidate-set membership and apparent relevance, not integrity, truth or a passed source check. Interviewers listen for whether you would file a rank as evidence of a breach.
Own the framing when someone reports that a retrieval path is now covered because a second scoring stage exists. The useful contribution is naming the property the stage lacks rather than arguing about its accuracy.
## The two scoring steps, and what each one sees A retrieval path over a document corpus normally scores twice. The **first stage** embeds the query and every stored chunk separately and ranks by vector distance, handing back a candidate set of perhaps twenty to a hundred chunks. A **reranking stage** then rescores those candidates against the query and keeps the handful that are actually placed in the model's prompt. How that second scoring is implemented is a retrieval-engineering subject in its own right; for an attacker only its consequence matters — it judges each candidate *as an answer to this query*, rather than as a point that sits near the query in a vector space. ## Two families of planted passage An attacker who can add records to the corpus has two broad ways to make one of them land in the top few. - **Shaped against the encoder.** Text optimised so its embedding sits close to a family of query embeddings. It frequently reads as broken or padded prose, because a human reader was never the target. It wins distance contests it has no business winning. - **Written to read as the best answer.** An ordinary, fluent passage that states a confident, complete, directly responsive answer to a question the attacker expects will be asked, in the register the corpus itself uses. Against first-stage similarity alone both families work, and the first needs less guesswork about how the question will be phrased. ## What the second stage does to each The first family collapses. Judged as an answer rather than as a vector, text built to satisfy a distance metric does not read like a response to anything, and it falls out of the small set that reaches the model. The second family is *promoted*. The entire purpose of the stage is to raise the candidate that reads most like the best available answer to the query — which is precisely what that passage was written to be. The stage most effective against one family is the stage that most reliably surfaces the other. This is the sentence worth carrying into an interview: a reranking stage does not remove corpus poisoning, it **changes which construction is worth writing**. ## Relevance is not provenance | step | what its score is about | what it knows about the author | | --- | --- | --- | | first-stage similarity | distance between the query embedding and the passage embedding | nothing | | reranking | how well this passage answers this query | nothing | | the model reading the top chunks | the text placed in front of it | only what the text itself claims | A stored record often carries metadata — who created it, when, by which route it entered the corpus. Those fields exist; they are simply not terms in the relevance score. A passage can be first because it is the best-written answer in the candidate set and simultaneously be the newest, least-vouched-for record in the store. ## Getting the direction of the claim right When a planted passage ranks first after reranking, that proves two things and no more: it was in the candidate set, and it read as the most relevant answer. It does **not** prove the store was breached — in most of these pipelines the corpus takes writes through an ordinary, sanctioned path. It does not prove the passage is true. It does not prove any source check was performed and passed. A candidate who says "the reranker approved it" has already made the error the question is designed to find. ## Why this gets asked Because "we added a reranker" is the most common thing a team says when asked what happens to a poisoned corpus, and it is half right in a way that is worse than being wrong: the family it removes is the visibly weird one, so the corpus looks cleaner afterwards while the construction that survives is the one that is hardest to spot by reading. An engineer who can state which half the stage is blind to — and that the blindness is authorship, not quality — is reasoning about the pipeline rather than reciting its parts.
- Does a planted passage ranking first after reranking tell you anything about the store's integrity?No. It tells you the passage was in the candidate set and read as the most relevant answer to that query. Anyone with an ordinary write path into the corpus can produce that outcome, so first place is evidence about text and scoring, not about a compromised store.
- Why does encoder-shaped text survive the first stage and die at the second?The first stage ranks by distance between separately embedded query and passage, which such text is built to win. The second stage asks how well the passage answers the query, and text assembled to satisfy a distance metric does not read as an answer to anything, so it drops out of the few chunks that reach the model.
- A team says the reranking stage handles corpus poisoning. What single question exposes the gap?Ask which term in the score corresponds to who wrote the passage. There is none. The stage ranks answer quality relative to the query; authorship, age and route of entry may sit in metadata but they are not what it is scoring on.
A reading panel that only ever asks "which of these best answers the question" will reliably reject a page of keyword salad and reliably pick the most polished answer in the pile — without ever asking who wrote it.
saying these in an interview costs you the question
- Says a reranking stage blocks planted content
- Treats a high rerank score as evidence a passage is authoritative
- Assumes an internal corpus implies a trusted author
- Thinks both families of plant fail equally at reranking
- Confuses relevance ordering with source verification