skip to content

Being Retrieved

A planted document that never reaches the top-k is inert, so an attacker inherits the ranking problem the pipeline's own authors have. Interviewers probe it because the two ways out fail differently.

on this pageshow

explore

questions

19

An attacker's forum post and the vendor's official docs sit in one index — what decides which the retriever returns?

level: juniorimportance: must knowfreq 72%

answer

  1. the store ranks, it does not vouch
  2. what input does the score actually take?
  3. chunks compete, documents do not
  4. similarity is a distance, not authority

basics

~20 s

Closeness to the query decides. A first-stage retriever scores each chunk by how near it sits to the question and has no notion of who wrote it, so a forum post and an official page compete on similarity alone.

solid answer

~40 s

Nothing in the ranking step looks at the source. First-stage retrieval turns the question into a vector, compares it against the stored vector for each chunk, and returns the nearest ones; where a lexical signal is also in the path, it scores shared wording. Both signals are functions of the text and the query, never of authorship. The index usually does store a `source` field, but a metadata field changes results only if the query path applies it, and the similarity score itself never does. So an attacker who writes a community answer in the same vocabulary the official documentation uses is competing on exactly the axis the retriever measures, with the vendor's own page as one more candidate. The retriever ranks; it does not vouch.

go deeper

for a junior

Be ready to say plainly what a first-stage retriever compares: the question against each stored passage, by distance. Know that it returns chunks, not documents, and that no part of that number knows who wrote the text.

for a middle

Explain where provenance actually lives — a metadata field the score never reads — and why that makes an outsider's passage a peer candidate to the vendor's own page rather than a lesser one.

for a senior

An interviewer expects you to turn this into a property of the pipeline: the security question about a corpus is which surfaces an outsider can write into, and what the assistant does with whatever comes back.

for a principal

Own the framing that ingest breadth is a trust decision, not a relevance decision. Indexing a community surface buys coverage and simultaneously admits authors to the grounding set; be able to state that trade in one sentence.

## The question behind the question A developer-support assistant for a cloud vendor is grounded on an index that holds two very different kinds of text: the vendor's own product documentation, and the public community forum where anyone with an account can post an answer. An interviewer asking this wants to hear whether you know what the retrieval step actually computes — because the whole family of corpus-poisoning attacks rests on the answer. ## What is stored, and what is scored Three facts, in order. **The unit is a chunk, not a document.** Ingestion splits each source into passages and stores one entry per passage. A retriever returns chunks. Nobody retrieves 'the official docs page' — they retrieve a few hundred words that were cut out of it. **The score is a distance.** A first-stage similarity search embeds the query and each stored passage *separately* — that is what a bi-encoder does — and ranks candidates by a vector distance such as cosine similarity. Where a lexical signal is also part of the path, it scores overlap of the actual terms. Either way the number is computed from two pieces of text: the query, and the passage. There is no third input. **Provenance is metadata, and metadata is not the score.** The index almost certainly stores a `source` field, a URL, an ingest timestamp. Those fields exist for filtering, for citations and for debugging. They only affect what comes back if the query path explicitly uses them — as a predicate, or as an extra term in a scoring formula somebody wrote. The similarity number itself has no slot for them. An index that contains a forum post and a docs page has, at the ranking layer, two rows of numbers. ## Why this matters to somebody planting text An attacker writing into the forum inherits a very ordinary problem: to be used, their passage has to be one of the nearest neighbours of the question a user will ask. They do not have to break into anything, guess a credential, or find a defect. They have to write a passage that sits close to the anticipated query in the space the retriever measures — which, for most encoders and most lexical scorers, means restating the question in the vocabulary the answerer uses and answering it directly. Notice what this costs and what it does not. It costs domain knowledge (you must know the terms the vendor's docs use) and a guess at the question. It does not cost any evasion: the winning passage is a well-written, on-topic answer. And it is bounded — the win is per-query, not global. ## The direction of the claim Getting this backwards is the classic junior error, and interviewers listen for it. | Observation | What it proves | What it does **not** prove | |---|---|---| | The forum chunk was returned | It scored inside the retrieval budget for that query | That the store was breached or tampered with | | The forum chunk ranked above the docs chunk | It matched the query wording more closely | That it is more accurate, more current, or endorsed | | The assistant cited the forum post | That chunk was in the grounding set | That anything verified it | A high similarity score means *a passage matched a query*. It is not a truth signal, not an authority signal, and not a provenance signal. The store measured a distance and sorted. ## The consequence for how you read an index Once you accept that ranking is blind to authorship, the security property of a knowledge base stops being 'is the content good' and becomes **who can put text into it, and what does the assistant do with what comes back**. Every ingest surface an outsider can influence — a community forum, a docs pull request, a support ticket that gets archived into the corpus, a partner page that gets crawled — is a place where somebody can compete for the top of a result set on equal terms with the vendor's own writing. That is not a defect in the retriever. It is what a retriever is. ## What to say in an interview Say: the retriever ranks chunks by similarity to the query; similarity is computed from text alone; source is a metadata field that the score never consults; therefore an outsider's passage and the vendor's passage are the same kind of candidate. Then add the sentence that shows you have thought about it: *similarity is not authority, so 'the assistant retrieved it' is a statement about distance, not about trust.*

  • If the assistant answers from the forum chunk and cites it, what has that proved?
    Only that the chunk was in the grounding set for that query — it scored inside the retrieval budget. A citation records which text was placed in front of the model, not that anything checked the text. Reading a citation as an endorsement is the same mistake as reading a similarity score as a truth score.
  • Indexing only vendor-written pages would remove the outsider from this picture. Does it remove the class?
    It removes anonymous authorship from one surface, not the property that ranking is blind to source. Any ingest path an outsider can influence restores it: a documentation pull request, an archived support ticket, a crawled partner page. The question to ask of any corpus is who can write into it, not who nominally owns it.
  • Does storing a source field on every chunk change the ranking at all?
    Not by itself. A metadata field affects results only where the query path uses it — for a filter, a citation, or a scoring term somebody deliberately added. The first-stage similarity number is computed from the query text and the passage text; there is no slot in it for who wrote the passage.

A tape measure is not a referee. It will tell you which of two passages sits closer to the question, and it has nothing at all to say about which one you should believe.

saying these in an interview costs you the question

  • Says the vector store verifies or vouches for its sources
  • Assumes official pages get a ranking boost by default
  • Thinks a high similarity score means the passage is correct
  • Believes retrieval returns whole documents rather than chunks
  • Reads a retrieved forum chunk as evidence the store was breached

context

open as a page

In RAG retrieval, why doesn't a metadata filter reading a document's own status field exclude a planted passage?

level: juniorimportance: must knowfreq 58%

basics

~20 s

A metadata predicate is only as trustworthy as where its value comes from. When status, date or type are copied out of the submitted document, the passage's author sets them too, so narrowing removes honest rivals and keeps the plant.

open as a page

In a wiki assistant that retrieves chunks, why does approving a page's diff not mean a human read what the model reads?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A review and a retrieval consume different artefacts. The reviewer reads an added hunk inside the whole page; the model receives one chunk cut from that page, alone. Approval evidences that a person read the file, not the fragment.

open as a page

A reranking stage is added to a RAG retrieval path — which planted passages does it promote?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A reranking stage drops planted text shaped against an embedding model, which reads as nonsense, and promotes planted text written to read as the ideal answer. It scores relevance to the query, never who wrote the passage.

open as a page

Why does an attacker's encoder-tuned passage rank in first-stage similarity search despite reading badly?

level: juniorimportance: must knowfreq 65%

basics

~20 s

First-stage similarity search embeds query and passage separately and ranks on vector distance alone. That score never reads the text for fluency, authorship or truth, so text fitted to the encoder's geometry can outrank prose written for a human reader.

open as a page

Why is a planted passage that ranks first usually the least strange-looking text in an index?

level: middleimportance: must knowfreq 58%

basics

~20 s

Ranking rewards clarity. A passage wins by restating the anticipated question in the answerer's own vocabulary and answering it directly, so the highest-scoring plant reads like the best-written reference entry in the corpus, not like anything anomalous.

open as a page

Why can narrowing a retrieval query on document-supplied metadata raise a planted chunk's odds of being returned?

level: middleimportance: should knowfreq 45%

basics

~20 s

Narrowing removes candidates, and a planted chunk is never among them: it satisfies the predicate by construction, while legitimate documents with blank or non-conformant metadata are dropped. Fewer rivals means a better chance at the top-k slots.

open as a page

What property lets one wiki paragraph read as background on the full page and as standing guidance in a single retrieved chunk?

level: middleimportance: should knowfreq 45%

basics

~20 s

Context-dependence, placed deliberately. The material framing the paragraph — a heading, a date, a scoping sentence, a pronoun's antecedent — sits next to it, not inside it, so a split that separates them emits only the assertive half.

open as a page

What does a planted passage that survives a RAG reranking stage cost an attacker to write?

level: middleimportance: should knowfreq 40%

basics

~20 s

It has to out-answer the corpus's genuine best passage for a query the attacker guessed in advance, judged by a stage that reads both. That is a quality contest, so the construction lands on thinly covered questions.

open as a page

What happens to an encoder-tuned planted passage when a corpus is re-embedded under a new model?

level: middleimportance: should knowfreq 45%

basics

~20 s

The fit dies while the document survives. A re-embed recomputes every vector with a different function, so a passage shaped against the old encoder's geometry lands somewhere unremarkable. What still ranks is whatever the passage plainly says.

open as a page

A report says any forum user can outrank the official docs in a support assistant's index — is that a finding?

level: seniorimportance: should knowfreq 42%

basics

~20 s

It depends on steerability, not on the score. Ranking above the docs is ordinary retrieval behaviour; a finding needs evidence that an untrusted author can choose which question their passage wins, and a consequence downstream when it does.

open as a page

In a service-bulletin retrieval index, how do you separate submitter-authored metadata from pipeline-assigned metadata?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Trace each field back to the moment it entered the index payload. Fields parsed out of the artefact - front matter, embedded properties, filename - were written by the submitter. Fields stamped at receipt or derived from the caller were not.

open as a page

An automated ticket-triage workflow acts on the top reranked precedent — an attacker's own closed ticket ranks first. Is that a reranking defect?

level: seniorimportance: should knowfreq 36%

basics

~20 s

No. The reranking stage promoted the passage that read most like the best answer, which is its job. The defect is an unattended workflow treating relevance order as authority over a corpus anyone who files a ticket can write into.

open as a page

An encoder-tuned passage ranks in your offline harness but never retrieves live — what explains the gap?

level: seniorimportance: should knowfreq 38%

basics

~20 s

An offline harness proves a property of a function you hold, not of a deployment. The live path differs in the unit indexed, the text extracted, the encoder pin, the scoring mix and the competition — any one ends the fit.

open as a page

A wiki owner says every diff is reviewed and the assistant is told to ignore page instructions. What do those two buy?

level: principalimportance: should knowfreq 38%

basics

~10 s

The review buys attacker cost, not coverage: the whole-page reading must stay unremarkable, but the reviewed and consumed artefacts differ. The prompt line buys less, because the construction need not read as an instruction.

open as a page

For a planted passage to displace the docs chunk that wins a support query, what must it outscore?

level: middleimportance: nice to knowfreq 34%

basics

~20 s

Two different bars. To reach the grounding set at all, the passage only has to beat the current last-placed candidate and clear any absolute score floor. To evict the incumbent docs chunk, it has to push that chunk out of the budget entirely — a much harder win.

open as a page

In a corpus whose status flags were copied from submitted files, what can a retrospective review prove?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Very little from inside the index. A planted row and a legitimate one carry self-declared values that pass the same predicate identically, so separation depends on records made outside the documents - and those establish submission events, not correctness.

open as a page

An ingestion pipeline splits wiki pages into overlapping chunks. Why does overlap lower the precision an attacker needs at a split point?

level: seniorimportance: nice to knowfreq 31%

basics

~20 s

Overlap duplicates text across adjacent units instead of reconciling it. A paragraph near a boundary is emitted both with its framing sentence and without it, and both units are embedded and separately retrievable, so a coarse aim suffices.

open as a page

How do you argue the value of encoder-tuned passages against a pipeline that re-embeds twice a year?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Budget it as dated evidence, not capability. The durable product is the finding that a first stage ranks text nobody read; the tuned passage is a perishable demonstration whose expiry is set by someone else's release schedule.

open as a page