For a planted passage to displace the docs chunk that wins a support query, what must it outscore?
answer
- you do not have to be first
- the budget is k, not one
- beat the last one inside the budget
- there is also a floor to clear
- presence and displacement are different wins
basics
~20 sTwo different bars. To reach the grounding set at all, the passage only has to beat the current last-placed candidate and clear any absolute score floor. To evict the incumbent docs chunk, it has to push that chunk out of the budget entirely — a much harder win.
solid answer
~40 sRetrieval hands the model a fixed budget of chunks, so the useful question is which bar you are aiming at. **Presence** costs the least: the passage must score above the current k-th candidate and, where the pipeline applies a minimum score, above that floor in absolute terms. Once it is inside the budget, it is in the model's context alongside the official page. **Displacement** — the grounding set holding the attacker's passage and *not* the vendor's — needs the incumbent pushed past position k, which for a well-matched documentation chunk is expensive. Both wins are bound to one anticipated phrasing: the passage was written against a guess at the question, so a paraphrase, a different user vocabulary, or a newly published official page can take the position back without anybody noticing a contest.
code
json · 12 lines{
"query": "how do I rotate an expired deploy key",
"top_k": 4,
"min_score": 0.62,
"candidates": [
{"chunk_id": "docs-4417#2", "score": 0.81, "source": "official-docs"},
{"chunk_id": "forum-90233#1", "score": 0.79, "source": "community-forum"},
{"chunk_id": "docs-4417#3", "score": 0.74, "source": "official-docs"},
{"chunk_id": "docs-2210#7", "score": 0.66, "source": "official-docs"},
{"chunk_id": "forum-88110#4", "score": 0.61, "source": "community-forum"}
]
}go deeper
Know that retrieval returns a fixed number of chunks, so a passage can reach the model's context without being the top result at all.
Be able to separate the two bars — beating the last in-budget candidate plus any minimum score, versus pushing the incumbent out of the budget — and say why the second costs far more.
Read a candidate list and state precisely what it establishes: which rows the budget and the floor cut, and whether you are looking at presence or displacement for that one phrasing.
Own the framing that a single winning query is a data point, not a capability claim; the number worth arguing about is coverage across the phrasings real users produce.
## Framing the contest correctly Somebody planting a passage in a corpus that a support assistant retrieves from inherits the ranking problem the pipeline's own authors have — but they get to pick which version of it they solve. Candidates in a real interview often answer as though there is one bar ('you have to rank first'). There are at least two, and they cost very different amounts. ## Bar one: presence Retrieval passes the model a fixed number of chunks — the retrieval budget, commonly written as k. Anything inside the budget is in the model's context. So to be *present*, a passage does not have to beat the best candidate; it has to beat the one currently sitting last inside the budget. On a query where the vendor's documentation contributes three strong chunks and the rest of the field is weak, the bar for fourth place can be far lower than the bar for first. A second bar sits underneath: many pipelines apply a minimum similarity score, below which a candidate is simply not returned. That is an absolute threshold, not a relative one. A passage can be the best of a weak field and still be dropped for failing to clear the floor. So presence needs both conditions: above the last in-budget candidate, and above the floor. ## Bar two: displacement Displacement is the stronger claim, and it is the one worth having: for this question, the grounding set contains the attacker's passage and does *not* contain the official one. That requires pushing the incumbent past position k — not merely outscoring it. If the vendor's documentation contributes several closely matching chunks to that query, every one of them has to fall out of the budget, and that is a much larger ask than slipping in at the bottom. This is why the honest description of most planted-passage results is 'it is in the context', not 'it replaced the documentation'. The two have different consequences: sharing the context means competing for the model's attention against a well-matched official page, while owning it means the model has nothing else on the subject to reconcile against. ## Where it stops working Both wins are narrow, and being precise about the narrowness is what separates a good answer from a hand-wave. - **It is tied to a phrasing.** The passage was written against an anticipated question. Users who ask the same thing in different words — a different noun for the same object, a symptom rather than a feature name — produce a different query and a different neighbourhood. - **It is tied to a snapshot of the corpus.** The vendor publishing a clearer page for that question can retake the position without anyone realising there was a contest. New content keeps arriving; the field the plant beat is not the field it faces next month. - **It is tied to the budget.** k is a configuration choice, and it can change. A larger budget makes presence easier and displacement harder at the same time. ## Reading a result set When you are handed a candidate list, read three things before you say anything: where the budget cuts, where the floor cuts, and how many of the in-budget rows come from the incumbent source. A planted chunk at rank two with three official chunks around it is a presence result. A planted chunk at rank one with the official chunks below the cut is a displacement result. Reporting the second when you have the first is the most common overstatement in this area. ## The claim you are entitled to A single observation of a planted chunk in a result set proves that the chunk scored inside the budget for that query at that moment. It does not prove the passage will rank for related questions, that the official page was removed, or that the retriever now prefers the source it came from. With a probabilistic pipeline and a corpus that keeps changing, one win is one data point about one query — which is precisely why the interesting number is not the score but how many distinct phrasings of the same user need still land on it.
- What is the practical difference between the planted chunk being present and it displacing the docs chunk?Presence puts the passage in the context beside a well-matched official page, so the model has two accounts to reconcile and the official one may win. Displacement removes the competing account entirely. Presence is cheap — beat the last in-budget candidate; displacement needs every incumbent chunk pushed past the budget, which on a well-documented question is expensive.
- Why does the win often fail on a paraphrase of the same question?The passage was written against an anticipated wording, so its advantage is measured against that query. A user who describes a symptom instead of naming the feature, or uses a synonym the documentation does not use, produces a different query vector and a different candidate neighbourhood. The narrowness is inherent to the method, not an implementation detail.
- A candidate list shows the planted chunk at rank one but the docs chunk at rank two. Is that displacement?No — both are inside the budget, so the model still sees the official account. That is a presence result with a favourable ordering. Calling it displacement overstates what the observation supports, and the distinction matters because the two lead to different downstream behaviour.
saying these in an interview costs you the question
- Thinks the plant must rank first to matter at all
- Ignores the absolute score floor and reasons only about rank
- Assumes one winning phrasing generalises to every phrasing
- Reports presence in the context as displacement of the docs
- Treats top-k as returning whole documents rather than chunks