In an email an unattended invoice workflow reads, why is an outright override span the weakest injection?
answer
- command versus premise
- the loudest shape is the trained one
- it never mentions the earlier rules
- a reason the task changed, not a demand
basics
~20 sAn outright override announces disobedience, which is the one shape refusal training and screening layers have seen most. Spans that carry supply a reason the task legitimately changed, so the model reclassifies its situation instead of defying its orders.
solid answer
~50 sAn override span states that the model should stop doing what it was told, and that shape loses three ways: it is the densest region of refusal training data, it is the densest region of the corpus an input screen was built from, and it forces the model to notice a conflict and pick a side — a comparison the application's own instructions usually win. A span that carries does the opposite. It never mentions the earlier orders. It presents as ordinary content of the kind the run was built to read as fact, such as a counterparty asserting that an order was amended, and supplies a reason the task in front of the model is now a different task. The model is not disobeying; it has reclassified its situation. That is the difference between a command and a premise.
go deeper
Be ready to state the contrast cleanly: a span that commands the model to abandon its orders versus a span that supplies a reason the job changed, and why the second is the one that carries.
Expect to explain the three separate failure modes of the override shape — trained resistance, screening-corpus density, and forcing an explicit conflict — rather than lumping them into 'models are trained against it'.
Show that you can say where the construction stops working: it needs a reason plausible inside the corpus the run reads, and a task that admits an alternative coherent reading. Name the obstacle it got past without proposing the fix.
Own the reading of the evidence. One successful run is a data point about one deployment, and a span that was acted on tells you it reached the context and cohered — not that anything upstream was missing.
## The setting An unattended supplier-invoice and purchase-order workflow reads inbound email. Every run opens by restating the buyer's standing routing and approval rules, then states the run's own task: extract the amount, match it to an open purchase order, file a verdict. Nobody reads the correspondence on the routine path; the filed verdict is the output, and later readers treat it as already checked because a machine signed it. Anything an attacker writes into an email body or signature block on an open thread arrives as data the run was built to read as fact. ## Two shapes an injected span can take **Shape one — the override.** Text addressed to the model, referring to its instructions, telling it to stop following them. This is the shape almost every candidate names first. **Shape two — the premise.** Text that never mentions the model or its instructions at all, that presents as ordinary content of the kind the run consumes, and that supplies a reason the task in front of the model is a different task from the one it assumed — a counterparty stating that the order was amended and that a different line now applies. ## Why the override shape loses Three things work against it, and they fail at different stages, so keep them apart. **It is the densest region of refusal training data.** That phrasing family has been public for years and is exactly what instruction-following adversarial post-training aims at. Models resist it better than almost anything else in this domain. **It is the densest region of screening data.** An input screen, whether it matches patterns or scores with a classifier, was built from a corpus of this shape, because that is the corpus that exists. A span in this family is the easiest thing such a layer has to score. **It poses the conflict out loud.** It requires the model to notice that two sets of instructions disagree and to choose between them. The application's instructions are the ones a model prefers in that comparison, and they were supplied in the turn it treats as most authoritative. Forcing an explicit choice is asking to lose the comparison you most reliably lose. ## What a carrying span supplies instead A span that carries never enters that comparison. It supplies three things. **A source worth crediting.** It arrives on a thread the workflow already corresponds with, from a party the run has reason to read, and it reads like that party. Nothing about it is anomalous to a model reading correspondence. **A reason the task changed.** Not a demand that the rules be dropped — an assertion about the world from which a different task follows. The standing rules remain true; the situation they govern is claimed to be different. **A job that is coherent on its own.** The model can carry the alternative reading through to an output that satisfies the run's stated task. A reading that leaves the model unable to finish gets discarded in favour of the original one. Notice what is absent: the span never contradicts the standing rules out loud. Compliance and obedience are made to coincide, which is why there is no disobedience for a trained preference to resist. ## Where the construction stops working It is target-specific. The reason has to be plausible inside the corpus the run actually reads, which means it depends on the buyer's vocabulary, the state of the thread, and what the workflow is for. That is reconnaissance cost, not model-trick cost, and it does not port to the next buyer. It also depends on the run's task admitting an alternative reading. If the job is narrow enough that no other coherent task exists — return one field from one document and nothing else is an output — a premise has nothing to attach to. And it says nothing about position or precedence. Where the span sits governs whether it survives to be read at all, which is a different question with a different answer. What decides adherence here is semantic: whether the model finds the alternative frame more coherent than the original one. ## Reading the result correctly A span the model acted on proves the span reached the context and was found coherent. It does not prove that a screening layer was absent, that the mailbox or the store was compromised, or that the model's instruction ordering was inverted. And a construction that worked once against one deployment has worked once against one deployment — with a probabilistic model, that is a data point, not a reliability claim.
- Does the model have to believe the claim for this to work?Belief is the wrong frame. The model weighs readings of the material in front of it, and the span only has to make the alternative reading the more coherent one — a claim consistent with the thread, the vocabulary and the run's stated job. Nothing verifies it, because the run's task is to read correspondence as fact, so plausibility inside that corpus is the whole bar.
- What if the claimed reason is checkable and demonstrably false?It only matters if something in the pipeline checks it. An unattended run that derives its verdict from the correspondence it was handed has no second source to contradict the claim, and the model has no notion of who wrote a passage or whether they were entitled to say it. Checkability becomes relevant only once a later reader re-derives the claim, which on the routine path nobody does.
- So does bare override phrasing never work?It works where nothing is in the way — a deployment with no input screen and a model that happens to comply. It is the cheapest thing to try, which is why it is tried first, and its failure is informative: it tells you something resisted, though not whether the model refused or a screening layer blocked. Those are different events and telling them apart is a separate measurement.
A forged order tells the clerk to break the rule. A forged fact tells the clerk the rule now applies to a different case — and never asks them to break anything.
saying these in an interview costs you the question
- Says the payload must tell the model to ignore previous instructions
- Treats it as a precedence problem the system prompt can settle
- Assumes anything a model obeyed must have looked like an instruction
- Thinks a longer or more emphatic command performs better
- Believes the span has to contradict the standing rules to beat them