skip to content

An attacker with no query access into a RAG index plants a document - what feedback do they get?

level: middleimportance: nice to knowfreq 31%

answer

  1. one-way write, no receipt
  2. discarded and pending look the same from outside
  3. the loss is the loop, not the write
  4. each retry costs a visible file and returns nothing

basics

~20 s

Effectively none. From outside, a file discarded at ingest and a file waiting for the next build look identical - the share shows it either way - and no surface reports which happened. The activation date is inferred, never observed.

solid answer

~50 s

Nothing comes back. Batch ingestion over a shared drive publishes no receipt: the crawl may not cover that folder, extraction may have failed on the file type, a curation step may have dropped it, or the build may simply not have run yet - and all of those outcomes present as the same thing, a file still sitting where it was written. Without a query into the assistant there is no channel that distinguishes them. The practical cost is the loop, not the write: an attacker who can observe the effect can adjust the passage and try again cheaply, and this one cannot. So a blind write has to be right on the first attempt, must survive whatever the pipeline does to it, and must still match the query it was written for whenever the build eventually lands.

go deeper

for a junior

Know that uploading a file to a corpus source gives no confirmation that anything ingested it, and that the uploader sees the same picture whether it was dropped or is merely pending.

for a middle

Explain the several stages that can quietly discard a document and why they are indistinguishable from outside. Be able to say what the loss of a feedback loop costs in attempts and in information.

for a senior

Demonstrate that you reason about what an observation supports rather than what it suggests, and that you can say which of several explanations for silence a given piece of evidence actually rules out.

for a principal

Be ready to argue how much confidence a one-way channel can ever justify, and why findings of this class need evidence from inside the pipeline rather than from the writer's side of it.

## Writing into a pipeline you cannot see the far end of This leaf's situation is narrow and worth stating precisely: someone holds upload rights to a file share that a corpus is batch-built from, and holds no access at all to the assistant that answers over it. Everything they do is one-way. That single constraint - the absence of a feedback channel - is what changes the shape of the work. ## The outcomes that all look the same Once the file is written, several very different things may happen, and none of them is reported anywhere the writer can see: - The crawl does not cover that folder, so the file is never a candidate. - Extraction fails on the format, so no text is produced from it. - A curation or approval step in front of the corpus discards it. - Chunking splits it so that the passage is divided across boundaries. - It is ingested cleanly and simply awaits the next build. - It is ingested, indexed, and never ranks highly enough for any query anyone asks. From the share, every one of those is indistinguishable from every other: the file is there, unchanged, with the same timestamp. This is what the phrase "the wait is inferred, never observed" means. The writer has a belief about a date and no evidence for it. ## What the loss of iteration actually costs Most attack work against a language-model application is iterative. You try something, look at what came back, adjust, try again; the information gained per attempt is what makes the work tractable. A blind write into a batch corpus removes that entirely, and the consequences compound: - **No convergence.** A passage cannot be tuned toward the queries it needs to win, because no result set is ever visible. - **No failure attribution.** If nothing happens, the writer cannot tell whether the document was discarded, whether it was indexed but outranked, or whether the build simply has not run. Each of those would call for a different next move, and there is no way to choose between them. - **Attempts are expensive and visible.** Each retry is another file on a share that people browse and that records who uploaded it. Repetition raises the chance of being noticed while producing no information in return. - **Timing is uncontrolled.** The effect, if there is one, appears whenever the schedule says. It may land when nobody is asking the relevant question at all. So the write is a bet placed once, on credit, with the settlement date and the settlement itself both unobservable. ## What the writer can and cannot infer A careful attacker reasons from proxies rather than feedback, and it is worth being honest about how weak those proxies are: | Observation | What it supports | What it does not support | |---|---|---| | The file is still on the share | the write landed | that anything ingested it | | Other people's documents are answered about | the pipeline runs on some sources | that it covers this folder | | The file's format matches existing corpus documents | extraction is plausible | that extraction succeeded | | A published refresh cadence | a rough window | that a build ran, or included this file | Every row is an inference about somebody else's pipeline made from outside it, and the strongest of them supports only a probability. ## Why this matters on the defending side of the table too An interviewer asking this is usually probing whether a candidate understands that findings in this class are cheap to file and expensive to confirm. The engineer who inherits the report faces the mirror image of the attacker's problem: the report says a document was written, the assistant is not currently quoting it, and nothing in either the share or the answer log settles whether the write ever reached the index. Both parties are reasoning about the same invisible boundary from opposite sides. The correct summary is not that the attack is impossible - documents do get ingested, and passages do get retrieved. It is that this variant trades away every advantage that interactive iteration gives, and buys in return a plant whose activation the writer neither controls nor witnesses.

  • If nothing changes for a month, what is the single most defensible conclusion?
    Only that no answer observed so far included a chunk from that document. It does not establish that ingestion rejected the file, that the folder is out of scope, or that the build never ran. Each of those would explain the silence equally well, and no observation available from the share separates them.
  • Why does this shape make the passage's wording matter more, not less?
    Because there is only one attempt. A passage that can be adjusted against feedback can start vague and get sharper; a blind one has to already match the queries it needs to win against, survive chunking intact, and stay plausible to anyone who opens the file. Everything that would normally be discovered by iteration has to be guessed up front.

saying these in an interview costs you the question

  • Assumes an ingest failure surfaces somewhere the uploader can see
  • Treats silence as proof that the document was rejected
  • Says the attacker can retry cheaply until it works
  • Confuses presence on the share with presence in the index
  • Ignores that every retry leaves another attributable file

context