An attacker's forum post and the vendor's official docs sit in one index — what decides which the retriever returns?
answer
- the store ranks, it does not vouch
- what input does the score actually take?
- chunks compete, documents do not
- similarity is a distance, not authority
basics
~20 sCloseness to the query decides. A first-stage retriever scores each chunk by how near it sits to the question and has no notion of who wrote it, so a forum post and an official page compete on similarity alone.
solid answer
~40 sNothing in the ranking step looks at the source. First-stage retrieval turns the question into a vector, compares it against the stored vector for each chunk, and returns the nearest ones; where a lexical signal is also in the path, it scores shared wording. Both signals are functions of the text and the query, never of authorship. The index usually does store a `source` field, but a metadata field changes results only if the query path applies it, and the similarity score itself never does. So an attacker who writes a community answer in the same vocabulary the official documentation uses is competing on exactly the axis the retriever measures, with the vendor's own page as one more candidate. The retriever ranks; it does not vouch.
go deeper
Be ready to say plainly what a first-stage retriever compares: the question against each stored passage, by distance. Know that it returns chunks, not documents, and that no part of that number knows who wrote the text.
Explain where provenance actually lives — a metadata field the score never reads — and why that makes an outsider's passage a peer candidate to the vendor's own page rather than a lesser one.
An interviewer expects you to turn this into a property of the pipeline: the security question about a corpus is which surfaces an outsider can write into, and what the assistant does with whatever comes back.
Own the framing that ingest breadth is a trust decision, not a relevance decision. Indexing a community surface buys coverage and simultaneously admits authors to the grounding set; be able to state that trade in one sentence.
## The question behind the question A developer-support assistant for a cloud vendor is grounded on an index that holds two very different kinds of text: the vendor's own product documentation, and the public community forum where anyone with an account can post an answer. An interviewer asking this wants to hear whether you know what the retrieval step actually computes — because the whole family of corpus-poisoning attacks rests on the answer. ## What is stored, and what is scored Three facts, in order. **The unit is a chunk, not a document.** Ingestion splits each source into passages and stores one entry per passage. A retriever returns chunks. Nobody retrieves 'the official docs page' — they retrieve a few hundred words that were cut out of it. **The score is a distance.** A first-stage similarity search embeds the query and each stored passage *separately* — that is what a bi-encoder does — and ranks candidates by a vector distance such as cosine similarity. Where a lexical signal is also part of the path, it scores overlap of the actual terms. Either way the number is computed from two pieces of text: the query, and the passage. There is no third input. **Provenance is metadata, and metadata is not the score.** The index almost certainly stores a `source` field, a URL, an ingest timestamp. Those fields exist for filtering, for citations and for debugging. They only affect what comes back if the query path explicitly uses them — as a predicate, or as an extra term in a scoring formula somebody wrote. The similarity number itself has no slot for them. An index that contains a forum post and a docs page has, at the ranking layer, two rows of numbers. ## Why this matters to somebody planting text An attacker writing into the forum inherits a very ordinary problem: to be used, their passage has to be one of the nearest neighbours of the question a user will ask. They do not have to break into anything, guess a credential, or find a defect. They have to write a passage that sits close to the anticipated query in the space the retriever measures — which, for most encoders and most lexical scorers, means restating the question in the vocabulary the answerer uses and answering it directly. Notice what this costs and what it does not. It costs domain knowledge (you must know the terms the vendor's docs use) and a guess at the question. It does not cost any evasion: the winning passage is a well-written, on-topic answer. And it is bounded — the win is per-query, not global. ## The direction of the claim Getting this backwards is the classic junior error, and interviewers listen for it. | Observation | What it proves | What it does **not** prove | |---|---|---| | The forum chunk was returned | It scored inside the retrieval budget for that query | That the store was breached or tampered with | | The forum chunk ranked above the docs chunk | It matched the query wording more closely | That it is more accurate, more current, or endorsed | | The assistant cited the forum post | That chunk was in the grounding set | That anything verified it | A high similarity score means *a passage matched a query*. It is not a truth signal, not an authority signal, and not a provenance signal. The store measured a distance and sorted. ## The consequence for how you read an index Once you accept that ranking is blind to authorship, the security property of a knowledge base stops being 'is the content good' and becomes **who can put text into it, and what does the assistant do with what comes back**. Every ingest surface an outsider can influence — a community forum, a docs pull request, a support ticket that gets archived into the corpus, a partner page that gets crawled — is a place where somebody can compete for the top of a result set on equal terms with the vendor's own writing. That is not a defect in the retriever. It is what a retriever is. ## What to say in an interview Say: the retriever ranks chunks by similarity to the query; similarity is computed from text alone; source is a metadata field that the score never consults; therefore an outsider's passage and the vendor's passage are the same kind of candidate. Then add the sentence that shows you have thought about it: *similarity is not authority, so 'the assistant retrieved it' is a statement about distance, not about trust.*
- If the assistant answers from the forum chunk and cites it, what has that proved?Only that the chunk was in the grounding set for that query — it scored inside the retrieval budget. A citation records which text was placed in front of the model, not that anything checked the text. Reading a citation as an endorsement is the same mistake as reading a similarity score as a truth score.
- Indexing only vendor-written pages would remove the outsider from this picture. Does it remove the class?It removes anonymous authorship from one surface, not the property that ranking is blind to source. Any ingest path an outsider can influence restores it: a documentation pull request, an archived support ticket, a crawled partner page. The question to ask of any corpus is who can write into it, not who nominally owns it.
- Does storing a source field on every chunk change the ranking at all?Not by itself. A metadata field affects results only where the query path uses it — for a filter, a citation, or a scoring term somebody deliberately added. The first-stage similarity number is computed from the query text and the passage text; there is no slot in it for who wrote the passage.
A tape measure is not a referee. It will tell you which of two passages sits closer to the question, and it has nothing at all to say about which one you should believe.
saying these in an interview costs you the question
- Says the vector store verifies or vouches for its sources
- Assumes official pages get a ranking boost by default
- Thinks a high similarity score means the passage is correct
- Believes retrieval returns whole documents rather than chunks
- Reads a retrieved forum chunk as evidence the store was breached