What does a planted passage that survives a RAG reranking stage cost an attacker to write?
answer
- the plant is no longer alone
- presence is cheap, placement is not
- it must out-answer real precedent
- thin coverage is where it lands
basics
~20 sIt has to out-answer the corpus's genuine best passage for a query the attacker guessed in advance, judged by a stage that reads both. That is a quality contest, so the construction lands on thinly covered questions.
solid answer
~50 sThe surviving family is not free. Because the second stage rescores candidates by how well each answers the query, a plant is no longer competing on distance — it is competing with the best real passage the corpus already holds. That imposes three costs: the attacker must anticipate the question, and roughly its phrasing, because a passage only competes on queries it was written for; it must be written to actually out-answer genuine content, which means being more direct and more complete than the real material rather than merely emphatic; and it has to stay in the corpus and read as unremarkable to anyone who opens the record. The consequence is a coverage map: on questions the corpus answers well and often, the plant loses to real content. It lands on sparse, novel or awkwardly-phrased questions — where nothing good exists to beat.
go deeper
Know that a passage which survives a reranking stage has to read as a genuinely good answer, which means somebody guessed the question and wrote real content. It is work, not a trick string.
Explain the shift in currency: the first step buys entry to the candidate set, the second buys placement and is paid for in writing quality and prediction. Say why emphasis and length do not substitute for fit.
Demonstrate the coverage argument — where in a corpus this is affordable and where genuine material outranks it — and state a result as a query class with a placement rate rather than as a flat claim.
Be able to say what this family is worth over time compared with constructions tied to a specific component, and to describe corpus regions by how expensive they are to contest rather than by whether an attack succeeded once.
## Why the second stage changes the economics With only first-stage similarity in the path, getting retrieved is a distance problem: place a record whose embedding sits near the expected query embedding and it enters the candidate set. Once a reranking stage rescores those candidates on how well each answers the query, being *near* is only the entry fee. The plant now has to win a comparison against everything else that was retrieved — including the corpus's genuine best passage on the subject. ## The cost ledger | what the attacker must spend | why the second stage forces it | what ends the construction | | --- | --- | --- | | a guess about the question | a passage is only in the contest for queries it was written to answer | the question that actually gets asked is not the one anticipated | | writing that out-answers real content | the stage promotes the candidate that reads as the better answer | strong, well-maintained genuine coverage of that question | | plausibility to a reader | records are occasionally opened by a person, and an obviously odd one draws attention | somebody reads the record | | persistence in the corpus | the plant must still be there when the query arrives | routine pruning, expiry or re-authoring of the store | Notice what is *not* on the ledger: knowledge of a particular embedding model. Nothing about this family depends on which encoder the store uses, so rebuilding the index does not touch it. The thing that ends it is a question it was not written for, or genuine material that answers better. ## Emphasis is not relevance A common mistake is to assume a longer, more insistent, more repetitive passage ranks higher. A relevance judgement rewards directness and fit: a passage that reads as *the answer* to the query beats one that reads as an argument about the topic, and padding dilutes the fit. This is also why the surviving family reads as ordinary, competent documentation rather than anything dramatic — the property that gets it promoted is exactly the property that makes it hard to notice by skimming. ## Where it lands, and why that matters when reporting Because success is a comparison, the construction's yield is uneven across a corpus. Well-trodden questions — the ones with several strong, curated passages — are bad targets. The good targets are the edges: a rarely-asked procedure, a newly-arrived subject, an unusual phrasing of a common question, a question the corpus answers only obliquely. Those are also, awkwardly, the questions on which a reader is least able to notice that the answer is wrong, because they had no prior expectation of the answer. That unevenness has a direct consequence for anyone writing up a result. A plant that reaches the top for one phrasing and eighth for another has not failed — it has a scope. The honest statement is a query class and an observed placement rate, not "the passage ranks first". A claim of general placement collapses at the owner's first retest with a phrasing you did not try, and the argument is lost on the reproduction, not on the substance. ## The judgement being tested An interviewer asking this is checking whether a candidate treats retrieval as a single trick or as a pipeline with different currencies at each step. The first stage buys presence in the candidate set; the second stage buys placement, and it is paid for in writing quality and prediction rather than in vector geometry. Someone who can name what that costs — and can point at the corpus regions where the cost is affordable — has understood the stage better than someone who only knows that fluent text survives it.
- Why is a longer, more emphatic planted passage not automatically a better one?A relevance judgement rewards directness and fit to the query. A passage that reads as the answer beats one that reads as an argument about the topic, and padding dilutes the fit. Emphasis is a property of tone; the score is about how well the text responds to the question asked.
- The passage ranks first for one phrasing and eighth for another. How do you state that result?As an envelope: the query class it was written for, the phrasings tried, and the observed placement rate. The construction is query-specific by nature, so a general claim is not just overstated, it hands the owner a cheap refutation with one phrasing you did not test.
- Which parts of a corpus make this family unaffordable?Regions with several strong, current, well-written passages on the same question. There the plant must out-answer curated material to be promoted, which is a lot of writing for one query. Sparse or freshly-emerging subjects cost far less because there is little to beat.
saying these in an interview costs you the question
- Thinks a longer or more forceful passage always ranks higher
- Assumes the plant only competes with other planted records
- Ignores that the question has to be anticipated in advance
- Claims general placement from a single observed phrasing
- Believes rebuilding the index ends this family