RAG & Knowledge-Base Poisoning
You will learn how attackers plant malicious content in the documents a RAG pipeline retrieves so the model faithfully reproduces or acts on it, distinct from poisoning the model's weights. Interviewers probe it because RAG is the most common enterprise LLM pattern and its ingestion path is a wide-open trust boundary.
on this pageshowhide
explore
- Getting into the Corpus11 questions
- Choosing Where to Write4 questions
- One Document or Many4 questions
- Waiting for the Reindex3 questions
- Being Retrieved19 questions
- A Passage That Ranks4 questions
- Tuned to the Encoder4 questions
- The Reranker's Blind Half3 questions
- Halves the Reviewer Reads4 questions
- Fields the Filter Trusts4 questions
- Taken as Authority8 questions
questions
page 2 of 2For a planted passage to displace the docs chunk that wins a support query, what must it outscore?
basics
~20 sTwo different bars. To reach the grounding set at all, the passage only has to beat the current last-placed candidate and clear any absolute score floor. To evict the incumbent docs chunk, it has to push that chunk out of the budget entirely — a much harder win.
An attacker with no query access into a RAG index plants a document - what feedback do they get?
basics
~20 sEffectively none. From outside, a file discarded at ingest and a file waiting for the next build look identical - the share shows it either way - and no surface reports which happened. The activation date is inferred, never observed.
In a corpus whose status flags were copied from submitted files, what can a retrospective review prove?
basics
~20 sVery little from inside the index. A planted row and a legitimate one carry self-declared values that pass the same predicate identically, so separation depends on records made outside the documents - and those establish submission events, not correctness.
An ingestion pipeline splits wiki pages into overlapping chunks. Why does overlap lower the precision an attacker needs at a split point?
basics
~20 sOverlap duplicates text across adjacent units instead of reconciling it. A paragraph near a boundary is emitted both with its framing sentence and without it, and both units are embedded and separately retrievable, so a coarse aim suffices.
How do you argue the value of encoder-tuned passages against a pipeline that re-embeds twice a year?
basics
~20 sBudget it as dated evidence, not capability. The durable product is the finding that a first stage ranks text nobody read; the tuned passage is a perishable demonstration whose expiry is set by someone else's release schedule.
When does planting more documents in a review corpus stop being worth what it costs?
basics
~20 sWhen the next document costs more than the question it would move is worth. Cheap, thinly covered questions go first; each further one is more crowded, more attributable, and needs upkeep as genuine content keeps arriving.
Triage fixed your corpus-poisoning finding by editing the one bad row. What do you tell the owner?
basics
~20 sEditing the row closes one instance, not the class. The only defensible claim is that this sentence is gone; anything broader would require re-verifying every passage that can be retrieved, which nobody has funded or scheduled.
A wrongly attributed answer circulated for weeks with no record of who saw the source chip: what can you honestly claim about impact?
basics
~20 sOnly a bounded upper limit and a window, labelled as such: the period the planted page was reachable and the number of answers that could have carried the chip. Reach, belief and repetition are unmeasured, and unmeasured is not zero.
showing 31–38 of 38