skip to content

RAG & Knowledge-Base Poisoning

You will learn how attackers plant malicious content in the documents a RAG pipeline retrieves so the model faithfully reproduces or acts on it, distinct from poisoning the model's weights. Interviewers probe it because RAG is the most common enterprise LLM pattern and its ingestion path is a wide-open trust boundary.

on this pageshow

explore

questions

page 1 of 2

An attacker's forum post and the vendor's official docs sit in one index — what decides which the retriever returns?

level: juniorimportance: must knowfreq 72%

answer

  1. the store ranks, it does not vouch
  2. what input does the score actually take?
  3. chunks compete, documents do not
  4. similarity is a distance, not authority

basics

~20 s

Closeness to the query decides. A first-stage retriever scores each chunk by how near it sits to the question and has no notion of who wrote it, so a forum post and an official page compete on similarity alone.

solid answer

~40 s

Nothing in the ranking step looks at the source. First-stage retrieval turns the question into a vector, compares it against the stored vector for each chunk, and returns the nearest ones; where a lexical signal is also in the path, it scores shared wording. Both signals are functions of the text and the query, never of authorship. The index usually does store a `source` field, but a metadata field changes results only if the query path applies it, and the similarity score itself never does. So an attacker who writes a community answer in the same vocabulary the official documentation uses is competing on exactly the axis the retriever measures, with the vendor's own page as one more candidate. The retriever ranks; it does not vouch.

go deeper

for a junior

Be ready to say plainly what a first-stage retriever compares: the question against each stored passage, by distance. Know that it returns chunks, not documents, and that no part of that number knows who wrote the text.

for a middle

Explain where provenance actually lives — a metadata field the score never reads — and why that makes an outsider's passage a peer candidate to the vendor's own page rather than a lesser one.

for a senior

An interviewer expects you to turn this into a property of the pipeline: the security question about a corpus is which surfaces an outsider can write into, and what the assistant does with whatever comes back.

for a principal

Own the framing that ingest breadth is a trust decision, not a relevance decision. Indexing a community surface buys coverage and simultaneously admits authors to the grounding set; be able to state that trade in one sentence.

## The question behind the question A developer-support assistant for a cloud vendor is grounded on an index that holds two very different kinds of text: the vendor's own product documentation, and the public community forum where anyone with an account can post an answer. An interviewer asking this wants to hear whether you know what the retrieval step actually computes — because the whole family of corpus-poisoning attacks rests on the answer. ## What is stored, and what is scored Three facts, in order. **The unit is a chunk, not a document.** Ingestion splits each source into passages and stores one entry per passage. A retriever returns chunks. Nobody retrieves 'the official docs page' — they retrieve a few hundred words that were cut out of it. **The score is a distance.** A first-stage similarity search embeds the query and each stored passage *separately* — that is what a bi-encoder does — and ranks candidates by a vector distance such as cosine similarity. Where a lexical signal is also part of the path, it scores overlap of the actual terms. Either way the number is computed from two pieces of text: the query, and the passage. There is no third input. **Provenance is metadata, and metadata is not the score.** The index almost certainly stores a `source` field, a URL, an ingest timestamp. Those fields exist for filtering, for citations and for debugging. They only affect what comes back if the query path explicitly uses them — as a predicate, or as an extra term in a scoring formula somebody wrote. The similarity number itself has no slot for them. An index that contains a forum post and a docs page has, at the ranking layer, two rows of numbers. ## Why this matters to somebody planting text An attacker writing into the forum inherits a very ordinary problem: to be used, their passage has to be one of the nearest neighbours of the question a user will ask. They do not have to break into anything, guess a credential, or find a defect. They have to write a passage that sits close to the anticipated query in the space the retriever measures — which, for most encoders and most lexical scorers, means restating the question in the vocabulary the answerer uses and answering it directly. Notice what this costs and what it does not. It costs domain knowledge (you must know the terms the vendor's docs use) and a guess at the question. It does not cost any evasion: the winning passage is a well-written, on-topic answer. And it is bounded — the win is per-query, not global. ## The direction of the claim Getting this backwards is the classic junior error, and interviewers listen for it. | Observation | What it proves | What it does **not** prove | |---|---|---| | The forum chunk was returned | It scored inside the retrieval budget for that query | That the store was breached or tampered with | | The forum chunk ranked above the docs chunk | It matched the query wording more closely | That it is more accurate, more current, or endorsed | | The assistant cited the forum post | That chunk was in the grounding set | That anything verified it | A high similarity score means *a passage matched a query*. It is not a truth signal, not an authority signal, and not a provenance signal. The store measured a distance and sorted. ## The consequence for how you read an index Once you accept that ranking is blind to authorship, the security property of a knowledge base stops being 'is the content good' and becomes **who can put text into it, and what does the assistant do with what comes back**. Every ingest surface an outsider can influence — a community forum, a docs pull request, a support ticket that gets archived into the corpus, a partner page that gets crawled — is a place where somebody can compete for the top of a result set on equal terms with the vendor's own writing. That is not a defect in the retriever. It is what a retriever is. ## What to say in an interview Say: the retriever ranks chunks by similarity to the query; similarity is computed from text alone; source is a metadata field that the score never consults; therefore an outsider's passage and the vendor's passage are the same kind of candidate. Then add the sentence that shows you have thought about it: *similarity is not authority, so 'the assistant retrieved it' is a statement about distance, not about trust.*

  • If the assistant answers from the forum chunk and cites it, what has that proved?
    Only that the chunk was in the grounding set for that query — it scored inside the retrieval budget. A citation records which text was placed in front of the model, not that anything checked the text. Reading a citation as an endorsement is the same mistake as reading a similarity score as a truth score.
  • Indexing only vendor-written pages would remove the outsider from this picture. Does it remove the class?
    It removes anonymous authorship from one surface, not the property that ranking is blind to source. Any ingest path an outsider can influence restores it: a documentation pull request, an archived support ticket, a crawled partner page. The question to ask of any corpus is who can write into it, not who nominally owns it.
  • Does storing a source field on every chunk change the ranking at all?
    Not by itself. A metadata field affects results only where the query path uses it — for a filter, a citation, or a scoring term somebody deliberately added. The first-stage similarity number is computed from the query text and the passage text; there is no slot in it for who wrote the passage.

A tape measure is not a referee. It will tell you which of two passages sits closer to the question, and it has nothing at all to say about which one you should believe.

saying these in an interview costs you the question

  • Says the vector store verifies or vouches for its sources
  • Assumes official pages get a ranking boost by default
  • Thinks a high similarity score means the passage is correct
  • Believes retrieval returns whole documents rather than chunks
  • Reads a retrieved forum chunk as evidence the store was breached

context

open as a page

In RAG retrieval, why doesn't a metadata filter reading a document's own status field exclude a planted passage?

level: juniorimportance: must knowfreq 58%

basics

~20 s

A metadata predicate is only as trustworthy as where its value comes from. When status, date or type are copied out of the submitted document, the passage's author sets them too, so narrowing removes honest rivals and keeps the plant.

open as a page

In a wiki assistant that retrieves chunks, why does approving a page's diff not mean a human read what the model reads?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A review and a retrieval consume different artefacts. The reviewer reads an added hunk inside the whole page; the model receives one chunk cut from that page, alone. Approval evidences that a person read the file, not the fragment.

open as a page

A reranking stage is added to a RAG retrieval path — which planted passages does it promote?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A reranking stage drops planted text shaped against an embedding model, which reads as nonsense, and promotes planted text written to read as the ideal answer. It scores relevance to the query, never who wrote the passage.

open as a page

Why does an attacker's encoder-tuned passage rank in first-stage similarity search despite reading badly?

level: juniorimportance: must knowfreq 65%

basics

~20 s

First-stage similarity search embeds query and passage separately and ranks on vector distance alone. That score never reads the text for fluency, authorship or truth, so text fitted to the encoder's geometry can outrank prose written for a human reader.

open as a page

Why does restricting who can write to a RAG index not stop an attacker adding a poisoned document?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A poisoning document does not enter at the index; it enters at one of the sources the corpus ingests. Restricting index writers leaves every feeding source's own writers untouched, and the attacker writes to whichever source reviews least.

open as a page

Why doesn't one planted review poison every answer an assistant gives from a marketplace review corpus?

level: juniorimportance: must knowfreq 58%

basics

~20 s

One planted passage wins only the questions whose wording lands near it in the retriever's vector space. First-stage search ranks passages by distance to each query, so a single document owns one narrow query neighbourhood, not the whole assistant.

open as a page

An attacker plants a passage in a nightly-rebuilt RAG corpus - when does the poisoning take effect?

level: juniorimportance: must knowfreq 58%

basics

~20 s

Not at write time. Retrieval searches an index, not the file share, so a planted passage is inert until the next scheduled build parses, chunks and embeds it - a delay the attacker neither sees nor controls.

open as a page

Why does an instruction-shaped-text screen over retrieved chunks miss a planted false fact?

level: juniorimportance: must knowfreq 66%

basics

~20 s

The strongest corpus plant contains no instruction at all: it is an ordinary declarative sentence that happens to be false. A screen matching directive-shaped text has nothing to match, and the model is reading, not disobeying.

open as a page

A research assistant renders an attacker's page-declared byline and publisher in a source chip: what does glancing at it verify?

level: juniorimportance: must knowfreq 68%

basics

~10 s

Nothing about authorship. The title, byline and publisher in a source chip are values the fetched page declared about itself, so a page the attacker wrote supplies its own credentials alongside its claim.

open as a page

Why is a planted passage that ranks first usually the least strange-looking text in an index?

level: middleimportance: must knowfreq 58%

basics

~20 s

Ranking rewards clarity. A passage wins by restating the anticipated question in the answerer's own vocabulary and answering it directly, so the highest-scoring plant reads like the best-written reference entry in the corpus, not like anything anomalous.

open as a page

Why can narrowing a retrieval query on document-supplied metadata raise a planted chunk's odds of being returned?

level: middleimportance: should knowfreq 45%

basics

~20 s

Narrowing removes candidates, and a planted chunk is never among them: it satisfies the predicate by construction, while legitimate documents with blank or non-conformant metadata are dropped. Fewer rivals means a better chance at the top-k slots.

open as a page

What property lets one wiki paragraph read as background on the full page and as standing guidance in a single retrieved chunk?

level: middleimportance: should knowfreq 45%

basics

~20 s

Context-dependence, placed deliberately. The material framing the paragraph — a heading, a date, a scoping sentence, a pronoun's antecedent — sits next to it, not inside it, so a split that separates them emits only the assertive half.

open as a page

What does a planted passage that survives a RAG reranking stage cost an attacker to write?

level: middleimportance: should knowfreq 40%

basics

~20 s

It has to out-answer the corpus's genuine best passage for a query the attacker guessed in advance, judged by a stage that reads both. That is a quality contest, so the construction lands on thinly covered questions.

open as a page

What happens to an encoder-tuned planted passage when a corpus is re-embedded under a new model?

level: middleimportance: should knowfreq 45%

basics

~20 s

The fit dies while the document survives. A re-embed recomputes every vector with a different function, so a passage shaped against the old encoder's geometry lands somewhere unremarkable. What still ranks is whatever the passage plainly says.

open as a page

How does an attacker who cannot see a RAG index pick which ingested source to poison?

level: middleimportance: should knowfreq 50%

basics

~20 s

They cannot see the index or the review roster, so they read observable proxies of each feeding source's pre-publication review: how fast a submission appears, whether prior submissions show up verbatim, and whether corrections are ever reverted. The source with the loosest review is the pick.

open as a page

Why does mass-duplicating one planted passage buy no reach when a pipeline collapses near-copies?

level: middleimportance: should knowfreq 44%

basics

~10 s

Near-identical text embeds to nearly the same vector, so copies chase the same questions rather than new ones, and ingest that keeps one representative removes most of them. Reach follows position, not count.

open as a page

What makes a plainly false claim planted in a catalog description durable rather than a one-off?

level: middleimportance: should knowfreq 48%

basics

~20 s

Durability comes from the plant being a stored text artefact rather than a decoding accident. It is retrieved again for every future asking of the same question, so it reproduces exactly where a hallucination does not.

open as a page

A planted source chip must survive a reader's click-through: what does that force the attacker to build?

level: middleimportance: should knowfreq 46%

basics

~20 s

A whole coherent artefact, not a string. The destination page has to keep standing up to inspection: self-declared metadata matching the chip, a layout and register that read like the impersonated publication, and continuous availability while the answer circulates.

open as a page

A report says any forum user can outrank the official docs in a support assistant's index — is that a finding?

level: seniorimportance: should knowfreq 42%

basics

~20 s

It depends on steerability, not on the score. Ranking above the docs is ordinary retrieval behaviour; a finding needs evidence that an untrusted author can choose which question their passage wins, and a consequence downstream when it does.

open as a page

In a service-bulletin retrieval index, how do you separate submitter-authored metadata from pipeline-assigned metadata?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Trace each field back to the moment it entered the index payload. Fields parsed out of the artefact - front matter, embedded properties, filename - were written by the submitter. Fields stamped at receipt or derived from the caller were not.

open as a page

An automated ticket-triage workflow acts on the top reranked precedent — an attacker's own closed ticket ranks first. Is that a reranking defect?

level: seniorimportance: should knowfreq 36%

basics

~20 s

No. The reranking stage promoted the passage that read most like the best answer, which is its job. The defect is an unattended workflow treating relevance order as authority over a corpus anyone who files a ticket can write into.

open as a page

An encoder-tuned passage ranks in your offline harness but never retrieves live — what explains the gap?

level: seniorimportance: should knowfreq 38%

basics

~20 s

An offline harness proves a property of a function you hold, not of a deployment. The live path differs in the unit indexed, the text extracted, the encoder pin, the scoring mix and the competition — any one ends the fit.

open as a page

Scoping a RAG poisoning test, how do you price which of several feeding sources to attempt?

level: seniorimportance: should knowfreq 40%

basics

~20 s

The weakest-reviewed source is not automatically the best target. Price each on three things together: how loose its review is, how likely its owner is to notice and revert, and whether its content actually reaches readers. A loose source nobody retrieves, or one whose owner reverts within a day, is a poor foothold.

open as a page

How would you estimate how many distinct planted posts a family of buyer questions requires?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Enumerate the phrasings buyers use, then gauge genuine competition for each by probing the assistant itself. Well-covered questions need several distinct passages each; thinly covered ones need one. The count is the sum, not a constant.

open as a page

A filed RAG-poisoning finding will not reproduce on demand - how do you tell a build-window artefact from a dead finding?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Compare clocks, not attempts. Repeated queries inside one index generation settle nothing; establish when the document was written, which builds have run since, whether the source still holds it, and whether the current index contains a chunk from it.

open as a page

An assistant repeats a false metric definition and triage closed it as a hallucination. What reopens it?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Determinism plus provenance reopens it. The same wording returns across fresh sessions, phrasings and model changes, and that wording is present verbatim in a retrieved chunk and in the stored row behind it - a decoding accident reproduces neither.

open as a page

An assistant attributed a fabricated statistic to a named institute: how do you tell an invented citation from a fetched page's own metadata?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Work from the turn's fetch record, not from a retry. If a fetched page declared that publisher and the chip reproduces it, the attribution was supplied by that page; if no fetched text carried the name, the generation produced it.

open as a page

A wiki owner says every diff is reviewed and the assistant is told to ignore page instructions. What do those two buy?

level: principalimportance: should knowfreq 38%

basics

~10 s

The review buys attacker cost, not coverage: the whole-page reading must stay unremarkable, but the reviewed and consumed artefacts differ. The prompt line buys less, because the construction need not read as an instruction.

open as a page

A poisoning write's weakest link is a partner-run intake queue you don't control - whose finding is it?

level: principalimportance: should knowfreq 33%

basics

~20 s

It is a shared finding whose root cause sits outside the team that runs the assistant. The assistant owner cannot fix the partner's review, and their instinct to lock down the index buys nothing against a write that never targeted the index. The real decision is whether to keep ingesting a source whose review they neither see nor set.

open as a page

showing 1–30 of 38