How does an attacker who cannot see a RAG index pick which ingested source to poison?
answer
- the review is invisible from outside
- read behaviour as a proxy
- turnaround, verbatim, reversions
- rank the sources, pick the loosest
- signal is about the review, not the index
basics
~20 sThey cannot see the index or the review roster, so they read observable proxies of each feeding source's pre-publication review: how fast a submission appears, whether prior submissions show up verbatim, and whether corrections are ever reverted. The source with the loosest review is the pick.
solid answer
~50 sThe attacker stands outside and cannot see the index, the build, or who reviews each source. What they can see is each source's public behaviour, and they use it as a proxy for how strict that source's review is. Fast turnaround suggests little or no human step. Prior submissions appearing verbatim - unedited, unrewritten - suggests the review does not substantively read content. Corrections that stick, rather than being reverted, suggest nobody is watching for tampering. A source that publishes quickly, keeps text as submitted, and never rolls anything back is one whose review is weak, and that is where a poisoning write is cheapest and most likely to survive. The attacker is comparing sources on these signals, not attacking the strongest one; the whole move is choosing the weakest review among several, none of whose review they can directly observe. What they get past is that specific source's editorial step - guessed at, never seen.
go deeper
Know that the attacker cannot see the review and must guess it from a source's public behaviour. Be able to name a couple of observable signals.
Explain the specific proxies - turnaround, verbatim survival, reversion - and why each correlates with weak review, and that the attacker ranks sources on them.
Show you understand the scoping discipline: a proxy measures the source's review only, not whether content reached the index or any answer, and would resist over-reading a signal.
Discuss what it means that the strength of every feeding source's review is inferable from outside - a property of the sources, not something a downstream index control changes.
## The problem the attacker faces A corpus is built from several ingested sources, and each source has its own pre-publication review - an editor, an owner, a moderator - of its own, unknown strictness. The attacker wants to write into the source whose review is weakest, because that write is cheapest to place and most likely to survive to the next build. But they cannot see any of it: not the index, not the build schedule, not the review roster, not who reads a submission before it publishes. Everything about the review is invisible from where they stand. So the move is inference: read observable, public behaviour of each source and treat it as a proxy for the hidden review. ## The proxies Several external signals correlate with weak review: - **Turnaround.** How long between a submission and its public appearance? Near-instant publication is a strong hint that no human is in the loop, or that the human step is a rubber stamp. A long, variable delay suggests a real editorial queue. - **Verbatim survival.** When prior contributions appear, are they unchanged from how they were submitted, or rewritten, trimmed, reformatted? Content that appears verbatim indicates the review is not substantively reading and editing - it is passing text through. Heavy editing indicates a reviewer who actually engages with content. - **Reversion behaviour.** Do corrections and edits stick, or do they get rolled back? A source where changes persist indefinitely is one nobody is auditing after the fact; a source that reverts anomalies has someone - or something - watching. None of these is a certainty. They are correlations an attacker reads across candidate sources to rank them, cheapest-and-loosest first. ## Why this is a comparison, not a single guess The attacker usually has more than one source they could write to, all feeding the same corpus. The skill is comparative: which of these several has the weakest review? A source might be attractive on one proxy (instant publication) and unattractive on another (aggressive reversion). Reading them together produces a ranking, and the top of that ranking is the target. This is why the leaf is *choosing where to write*: the write itself is trivial once the source is chosen; the judgment is picking the source. ## What the signals are and are not A crucial discipline: an observable proxy tells you about the *review*, not about the *index*. Fast turnaround means the source publishes quickly - it does not tell you the build has run, that the chunk reached the index, or that any query retrieves it. The attacker who conflates "my submission published on the source" with "my content is in the model's answers" has over-read the signal. All the proxy establishes is that the source's review was, or was not, an obstacle. Equally, verbatim survival on the source proves the review did not rewrite the text; it does not prove the text will survive later stages such as deduplication or a build's own filtering. The proxy is scoped to the one obstacle it measures. ## Why interviewers ask it It separates candidates who can name "poison a source" from those who can explain *how a blind attacker chooses which source*. The answer requires understanding that the review is invisible and must be inferred from behaviour - which is exactly the reasoning a defender needs to anticipate, and exactly the mechanism a red-teamer executes. A weak answer treats source selection as obvious or random; the real move is reading proxies to rank the hidden reviews.
- Why is fast turnaround on a source a signal about its review but not about the index?Turnaround measures how quickly the source itself publishes a submission, which reflects whether a human review step exists. It says nothing about when the build runs, whether the chunk reached the index, or whether any query retrieves it. Reading it as proof of end-to-end reach over-extends a signal that only measures the source's own review.
- If prior submissions appear verbatim, what has the attacker actually learned?That the source's review does not substantively rewrite content - text passes through roughly as submitted, so an editor is not closely reading it. That is a strong proxy for weak review. It does not guarantee later stages leave the text intact; deduplication or build-side filtering can still change or drop it. The signal is scoped to the editorial step it observed.
saying these in an interview costs you the question
- Assumes source selection is obvious or random
- Reads publication speed as proof of reaching the index
- Thinks the attacker can see each source's review
- Confuses surviving a source's review with surviving the whole pipeline