skip to content

Waiting for the Reindex

A plant answers nothing until a build picks it up, and the attacker has no channel telling them which. Interviewers probe it because the write is cheap and the blind wait is the real cost.

on this pageshow

explore

questions

3

An attacker plants a passage in a nightly-rebuilt RAG corpus - when does the poisoning take effect?

level: juniorimportance: must knowfreq 58%

answer

  1. the write and the effect are different events
  2. retrieval searches an index, not the share
  3. somebody else's job schedule owns the start date
  4. removal runs on that same clock, backwards

basics

~20 s

Not at write time. Retrieval searches an index, not the file share, so a planted passage is inert until the next scheduled build parses, chunks and embeds it - a delay the attacker neither sees nor controls.

solid answer

~50 s

The write and the effect are two different events on two different clocks. A retriever does not read the document share; it searches an index of embedded chunks that a pipeline builds from those sources on a schedule - here, nightly. Until that build runs, no vector exists for the planted passage, so no similarity search can return it and no answer can change. That means corpus poisoning ships with a built-in latency owned by somebody else's job scheduler, and the attacker has no way to shorten it or confirm it elapsed. The same clock runs in reverse: pulling the file back off the share removes it from the source, not from the index, so the passage can keep being retrieved until the following build. In both directions, the source and the index disagree for a whole window.

go deeper

for a junior

Recall that retrieval runs over an index built from the sources on a schedule, so a newly written document cannot be returned until that build runs. Say plainly that the effect is deferred, not immediate.

for a middle

Be ready to walk the stages between a file landing on a share and a chunk being retrievable - crawl, extract, chunk, embed, index - and to say which of them run only at build time.

for a senior

Show that you check the index and the source separately during an incident. An interviewer wants to hear you refuse to call a deletion a fix until a build has regenerated the index without the document.

for a principal

Own the framing that the build cadence sets both the exposure window and the remediation window, and that any claim about what the corpus contains is really a claim about the last build.

## Two events, not one The common wrong answer here is that corpus poisoning takes effect when the attacker writes. It does not. Writing puts bytes into a **source system** - a document-management share, a wiki, a ticket queue. Answering uses an **index**: a store of vectors, each built from a chunk of text that some pipeline extracted, split and embedded earlier. A first-stage similarity search compares a query vector against the vectors that exist *right now* in that index. A document that has never been through the pipeline has no vector, occupies no position in that space, and cannot be returned at any rank. So the interesting quantity is the gap between those two events, and who owns it. ## What sits between the write and the answer A typical batch pipeline over a shared drive walks roughly this path: | Stage | What it does | What it can drop | |---|---|---| | Crawl | enumerates files in the folders in scope | anything outside scope | | Extract | turns bytes into text | files it cannot parse | | Chunk | splits text into passages | nothing, but it moves boundaries | | Embed | encodes each chunk into a vector | nothing normally | | Index | writes vectors and metadata into the store | superseded or filtered records | Every one of those stages runs at build time, not at write time. A nightly cadence means the passage's earliest possible activation is the next run of that job; a weekly cadence means up to a week. If the pipeline only rebuilds folders it detects as changed, or only re-embeds when a model or configuration changes, the window can be longer still. None of this is a defence somebody designed against poisoning - it is ordinary batch-pipeline mechanics that happens to govern the attack's delivery date. ## The attacker's clock is not the attacker's The consequence for the person doing the writing is severe and specific: the activation date belongs to an operations schedule they cannot read. They cannot bring it forward, and unless they can query the assistant they cannot tell when it arrived. A plant written on Tuesday afternoon might answer on Wednesday morning, or never - because extraction failed on the file type, because the folder was out of the crawler's scope, because a curation step in front of the corpus discarded it. From outside, **a file that was silently discarded and a file that is merely waiting look exactly the same**: both are still sitting on the share. This is why the write is a bet rather than an operation. The attacker pays up front - a document that exists, is attributable to whatever account uploaded it, and sits in a place people can find - in exchange for a possible effect at an unknown later time. ## The same latency running backwards The second half of the point is the one that surprises people, and it matters more than the first. When somebody finds the planted document and deletes it from the share, they have changed the *source*. The index still holds the chunk and its vector until the next build regenerates the index without it. So for the length of that window, the assistant can keep quoting a passage whose file no longer exists anywhere a person can look at. "We removed the document" is a statement about the share; it is not yet a statement about what retrieval can return. Anyone triaging the incident has to say which of the two they checked. The symmetric framing is the thing to remember: **the index reflects the sources as of the last build, in both directions.** Additions are late; deletions are late by the same amount. ## What each observation actually proves - The document is on the share -> the write landed on the share. Nothing more. - The assistant quoted it -> a chunk from it was in the index and was retrieved for that query. It does not tell you when it was ingested. - The assistant did not quote it -> the current index returned no matching chunk for the query that was tried. It does not prove the document was rejected, and it does not prove it will never be indexed. - The file was deleted -> the source no longer holds it. The index may still. Getting these directions right is most of the value of the topic. A candidate who says "we deleted it, so it is fixed" has skipped a whole build cycle, and a candidate who says "the poisoning starts as soon as you upload" has skipped the pipeline.

  • The morning after the finding is filed, the file is pulled off the share. When does the assistant stop quoting it?
    At the next build that regenerates the index without it. Deleting from the share changes the source; the chunk and its vector stay in the index until a rebuild drops them. Inside that window the assistant can still return a passage from a file nobody can find any more, so a removal alone is not yet evidence that retrieval has changed.
  • Does a planted file sitting on the share for a week mean it is in the index?
    No. Presence on the share proves only that the write landed there. The crawl may not cover that folder, extraction may have failed on the file type, or a curation step may have discarded it - and none of those leave a trace the writer can see. Only a retrieval that returns a chunk from the document shows it made it through.
  • How does the picture change on a corpus with continuous rather than nightly ingestion?
    The window shrinks but does not vanish, and the reasoning is identical: there is still a moment before the document has a vector and a moment after removal when its vector is still there. Continuous ingestion mostly changes the size of the disagreement between source and index, not its existence.

Writing the document is leaving a book on the loading dock. Nothing on the shelves changes until the next stocking run - and taking it back off the dock does not un-shelve the copy that already went out.

saying these in an interview costs you the question

  • Says corpus poisoning takes effect the moment the file is written
  • Assumes the retriever reads the document share directly
  • Treats deleting the source file as immediate remediation
  • Thinks no change in answers today proves the write was rejected
  • Forgets that the deletion latency equals the addition latency

context

open as a page

A filed RAG-poisoning finding will not reproduce on demand - how do you tell a build-window artefact from a dead finding?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Compare clocks, not attempts. Repeated queries inside one index generation settle nothing; establish when the document was written, which builds have run since, whether the source still holds it, and whether the current index contains a chunk from it.

open as a page

An attacker with no query access into a RAG index plants a document - what feedback do they get?

level: middleimportance: nice to knowfreq 31%

basics

~20 s

Effectively none. From outside, a file discarded at ingest and a file waiting for the next build look identical - the share shows it either way - and no surface reports which happened. The activation date is inferred, never observed.

open as a page