Why does an instruction-shaped-text screen over retrieved chunks miss a planted false fact?
answer
- no command was ever given
- the screen scores mood, not truth
- an ordinary on-topic declarative sentence
- detection has nothing to inspect
- reads as an error, not an attack
basics
~20 sThe strongest corpus plant contains no instruction at all: it is an ordinary declarative sentence that happens to be false. A screen matching directive-shaped text has nothing to match, and the model is reading, not disobeying.
solid answer
~50 sA screen of that kind scores retrieved text for the features of a command - imperative mood, a role switch, a claim that the task has changed, an address to the assistant rather than to a reader. The strongest corpus plant has none of them. Picture an internal assistant that answers from a data catalog, where a metric's description field is a stored row rendered into the prompt as prose and anyone with catalog write access can edit it. Rewrite that description so it states the metric's definition incorrectly and you have added a fluent, on-topic, grammatically ordinary sentence. Nothing was commanded, so instruction detection has nothing to inspect, and the assistant is not misbehaving: it grounds its answer in the text it retrieved, exactly as designed. What lands is a false operational answer that someone quotes, and it looks like an everyday model error.
go deeper
Be ready to state plainly that the strongest corpus plant contains no instruction, so a detector looking for command-shaped text finds nothing. Do not answer by listing detection tools.
Explain what such a screen is actually scoring - imperative mood, role markers, off-topic structure - and why a fluent on-topic false sentence exhibits none of it while still reaching the model.
Show you would classify the failure correctly under pressure: the assistant grounded its answer faithfully, so the defect sits with the material it was handed, not with the model that read it.
Own the framing consequence: if the only thing inspecting retrieved content is an instruction detector, an entire class of corpus failure is invisible by construction and will keep being reported as model quality.
## Two attacks that share one vocabulary **Prompt injection** aims at the *application's* instructions. The attacker gets text in front of the model - directly in their own turn, or indirectly through content the application retrieves, fetches or is handed - and that text tries to change what the assistant is doing. **Jailbreaking** aims at something else again: the model's *trained* refusal. Both are attacks on behaviour, and both leave a characteristic trace, because a directive has to read like a directive to work. This leaf is about the third thing, and it is the one that gets confused with the first. The attacker does not tell the assistant to do anything. The attacker changes what the assistant will read as true, and then lets the ordinary pipeline do the rest. ## What the screen actually inspects A screen placed over retrieved text before it reaches the model - whether it is a small classifier, a set of patterns, or a generative screening model asked to judge a span - is scoring for the shape of an instruction. The signals available to it are things like: - imperative verbs addressed to an assistant rather than to a human reader - an assertion that earlier guidance is superseded, or that the task has changed - role or speaker markers embedded inside body text - structure that mimics a system or developer message - content that is off-topic for the passage it appears in A false declarative sentence has none of these. It is on-topic by construction - it has to be, or nobody would retrieve it for the question it is meant to answer - and its register is identical to every other passage in the corpus, because it was written to sit beside them. | Stage | What it can see | What it cannot see | |---|---|---| | Screen over retrieved text | Whether the span looks like a command | Whether the span is true | | The assistant | The chunks it was handed | Which of them anybody vouched for | | A comparison of answer to retrieved text | Whether the answer follows from the source | Whether the source follows from reality | That middle row is the whole point. Similarity ranking is a distance computation; it is not a judgment about authority, provenance or truth. Being retrieved is not being vouched for. ## Why the model is behaving correctly This matters for how the failure gets classified. When an assistant repeats a planted claim it has not disobeyed anything. It was asked to answer from the retrieved material, and it did. The answer is faithful to its source. The defect is upstream: a store that renders anyone's edit into an answering context treated that edit as material worth grounding on. So the failure surfaces as a wrong answer with no anomalous features anywhere in the trace - no blocked span, no refusal, no odd instruction sitting in the transcript. It looks exactly like a model getting something wrong on its own. ## What the attacker pays, and what they get The cost is write access to one field that gets rendered into prompts, plus knowing which question people ask routinely. There is no wording to tune, no screen to probe, no refusal to work around - which also means there is nothing about the construction that becomes obsolete when a model version changes or a screening layer is retrained. A directive-shaped span is in an arms race with whatever inspects text; a false sentence is not in that race at all, because the thing that would have to inspect it is a fact check, and no stage of the pipeline is one. The payoff in an ask-the-data setting is a false operational answer that a person carries into a document, a decision or a number somebody else quotes. And there is a second, quieter payoff: when the error is noticed, the natural diagnosis is that the model hallucinated. Triage that closes it as a model error leaves the artefact in place, so the same answer comes back tomorrow. ## Where it stops working It stops working where somebody actually reads the source and disagrees with it - which is why plausibility, not cleverness, is the binding constraint on this construction. It stops working where the corpus is small enough that a contradiction with a neighbouring passage is obvious. And it stops working where the claim is checkable against something other than text, because the moment a reader recomputes the number the sentence describes, the sentence loses. None of that is detection of an attack. It is somebody noticing a fact is wrong - which is why this class of finding is so frequently filed against the wrong owner.
- Does this mean a screen over retrieved text buys nothing?It buys what it was built for: spans that try to redirect the assistant, which are still the common case and still arrive through retrieved content. What it cannot do is adjudicate a claim. Its output is a score for instruction-likeness, and reporting that score as if it meant the retrieved material is safe is a direction-of-claim error - passing means the text scored low on one feature, nothing more.
- If nothing was commanded, in what sense is this an attack at all?Somebody authored a specific false sentence, placed it in a field they knew is rendered into answering context, and chose wording that matches the question people ask. That is authored, placed and aimed - the same three things any injection requires. Only the last step differs: instead of asking for behaviour, it supplies a premise and lets the normal pipeline convert it into an answer.
- How is this different from the assistant hallucinating?A hallucination is produced at decoding time and is not reliably reproducible; rephrase the question, open a fresh session, or change sampling and it usually moves. A plant is a stored artefact: it is retrieved again for the same question, so the same wrong claim comes back with the same wording, and it survives session boundaries and model changes because it never lived in the model.
A forged order in a company's mail is detectable because orders have a shape. A wrong number quietly typed into the reference table nobody re-reads has no shape at all - it just gets used.
saying these in an interview costs you the question
- Claims injection detection also covers corpus poisoning
- Treats text from the company's own store as trusted
- Says the model misbehaved, when it grounded the answer correctly
- Calls every false model output a hallucination
- Reads a high similarity score as a claim about truth
- Assumes an attack must contain an instruction to count