skip to content

One Document or Many

One planted passage owns a query neighbourhood, not an assistant, and wider influence costs as many documents as it takes to crowd out the genuine ones. Interviewers want a cost, not a scare story.

on this pageshow

explore

questions

4

Why doesn't one planted review poison every answer an assistant gives from a marketplace review corpus?

level: juniorimportance: must knowfreq 58%

answer

  1. retrieval runs once per question
  2. the store ranks by distance, not by source
  3. a passage is only in play where it ranks
  4. the unit of compromise is a query neighbourhood
  5. breadth is bought in distinct documents

basics

~20 s

One planted passage wins only the questions whose wording lands near it in the retriever's vector space. First-stage search ranks passages by distance to each query, so a single document owns one narrow query neighbourhood, not the whole assistant.

solid answer

~50 s

Retrieval runs per question. For each buyer question the pipeline embeds the query, ranks stored passages by distance, and hands the top few to the generator. A planted review is therefore only in play for the questions it happens to rank highly for — the small neighbourhood of phrasings whose embeddings sit near it. Every other question pulls other passages and is untouched. So the unit of compromise is a query neighbourhood, not "the assistant". A red-teamer who moves one answer with one planted post has demonstrated exactly that: one phrasing, one retrieval, one generation. Widening it means covering more neighbourhoods, and since near-copies of the same text land in the same place — and get collapsed where a pipeline keeps one representative — that breadth has to be bought in genuinely different documents, each an attributable post under a listing the attacker controls.

go deeper

for a junior

Be ready to say that retrieval happens per question and ranks passages by similarity, so a planted passage only affects questions that land near it. Naming that limit correctly is most of the answer here.

for a middle

Explain the mechanics: query and passage are embedded, ranked by distance, and the top few are pasted into the prompt. From that, derive why the effect is per-neighbourhood and why copies of one passage do not widen it.

for a senior

Show how you would bound the claim in a real assessment: enumerate the question phrasings that matter, check how many moved, and refuse to state a result wider than what you measured.

for a principal

Own the framing an owner will act on. The useful statement is a count — how many documents moved how many questions — because that is what turns an alarming anecdote into something anyone can weigh.

## What actually happens when the assistant answers A product-research assistant over marketplace content does not "read the corpus". For each buyer question it runs a retrieval step: the question is turned into a vector by an embedding model, stored passages (chunks, not whole documents) are ranked by distance to that vector, and the top handful are pasted into the prompt as context. The generator then writes an answer from whatever ended up there. Everything about the scope of a poisoning write follows from that loop. The planted passage does not sit "in the assistant". It sits in a store, and it is only ever consulted for questions whose vectors are close to it. ## The query neighbourhood Call the set of question phrasings for which a passage lands in the retrieved top-k its **query neighbourhood**. That neighbourhood is a property of the passage's position in the embedding space and of the competition around it. A passage about the battery life of one specific product model is near the phrasings buyers use to ask about that model's battery, and far from questions about sizing, shipping, or a different product entirely. So the honest statement of what one planted document buys is: *for the phrasings in its neighbourhood, and only where it outranks the genuine passages competing there, it becomes part of the context the generator sees.* That is a real result and it can be a serious one — a false claim a buyer acts on is a payoff. It is not "the assistant is compromised". ## Why the wrong answer is so common The misconception — "one poisoned document owns the assistant" — comes from carrying an intuition across from systems where a single write really does have global effect: a modified configuration file, a compromised dependency, a changed system instruction. In those systems one artefact is consulted on every request. A corpus passage is not: it is one candidate among many, selected per query, and it competes with everything else the store holds. It is also fed by how a demonstration reads. A tester shows a screenshot of a moved answer. The screenshot proves a passage reached the context for that phrasing and that the generator used it. It says nothing about the other thirty-nine questions buyers actually ask, and with a probabilistic generator it is a single sample even for the one phrasing shown. ## What decides how narrow the neighbourhood is Two things, and both are attacker costs rather than defender knobs: 1. **How many genuine passages already answer the question.** A question dozens of detailed reviews already cover has a crowded neighbourhood: to be seen, the plant must outrank several established passages for each phrasing. A question almost nothing in the corpus answers is cheap — there is little competition and a single passage may hold the slot outright. 2. **How many distinct phrasings buyers use.** One posted passage does not automatically cover every wording of a question; buyers ask the same thing in several ways and those wordings may not all land in one neighbourhood. ## Breadth has a price, and the price is distinct documents The natural next move — post the same passage fifty times — does not widen anything. Identical or near-identical text embeds to nearly the same vector, so the copies occupy the same neighbourhood and compete with each other for the same slots rather than reaching new questions; and where the pipeline keeps one representative of near-duplicate text, most of the copies never exist as separate candidates at all. Reach across a family of questions therefore costs *genuinely differing* documents, each of which is an attributable write under a listing the attacker controls. That is what makes the scope question the first one worth asking: it converts a vague claim ("the corpus is poisoned") into a countable one ("n documents moved m of the questions we enumerated"). ## Directions to keep straight - A retrieved planted passage proves the passage **reached the context**, not that the store was breached; the write went through an ordinary user-content channel the pipeline ingests. - A high similarity score proves the passage **matched the query**, not that it is true, authoritative, or from a trusted seller. The store ranks by distance and has no notion of who wrote the text. - A moved answer on one phrasing proves that phrasing moved **once**. Repeat it before calling it a finding. ## The one case where a single document is enough When the target question is narrow, rarely asked, and barely covered by genuine content, one passage can own it — and that is precisely the situation worth looking for, because it is the cheapest result available. The mistake is generalising from it to the assistant as a whole.

  • A tester moved one buyer question's answer with a single planted post. What has that actually demonstrated?
    That the passage reached the retrieved context for that phrasing and outranked the genuine passages competing there, and that the generator used it once. It says nothing about other phrasings, other questions, or repeatability — with a probabilistic generator a single success is one sample, not a rate.
  • When is a single planted document genuinely enough?
    When the target question is narrow and the corpus barely covers it. With little genuine competition in that neighbourhood, one passage can hold a retrieved slot on its own. That is why sizing starts by asking how many real sources already answer the question, not by counting how many posts you can afford.
  • Does the planted passage need to look authoritative to be retrieved?
    Not for first-stage retrieval. A similarity search embeds query and passage separately and ranks by vector distance; it has no notion of author, reputation or truth. Authority-shaped wording matters later, for whether the generator repeats the claim and whether a reader believes the answer — not for getting into the candidate list.

A planted signpost misdirects only the travellers standing at that one junction. Everyone on the other roads never sees it, however convincing the sign is.

saying these in an interview costs you the question

  • Claims one poisoned document compromises the whole assistant
  • Treats a high similarity score as evidence of authority or truth
  • Says the vector store was breached when a passage was merely retrieved
  • Assumes every stored passage is consulted on every question
  • Generalises one moved answer into a corpus-wide result

context

open as a page

Why does mass-duplicating one planted passage buy no reach when a pipeline collapses near-copies?

level: middleimportance: should knowfreq 44%

basics

~10 s

Near-identical text embeds to nearly the same vector, so copies chase the same questions rather than new ones, and ingest that keeps one representative removes most of them. Reach follows position, not count.

open as a page

How would you estimate how many distinct planted posts a family of buyer questions requires?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Enumerate the phrasings buyers use, then gauge genuine competition for each by probing the assistant itself. Well-covered questions need several distinct passages each; thinly covered ones need one. The count is the sum, not a constant.

open as a page

When does planting more documents in a review corpus stop being worth what it costs?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

When the next document costs more than the question it would move is worth. Cheap, thinly covered questions go first; each further one is more crowded, more attributable, and needs upkeep as genuine content keeps arriving.

open as a page