Which prompt-injection controls should sit outside the model rather than in the prompt?
answer
- ask: does it hold if the model is fooled
- probabilistic prompt vs deterministic code
- assume the prompt layer fails
- failures correlate, they do not multiply
- removing a capability has no failure rate
basics
~20 sAnything that must hold after the model is persuaded belongs in code: approval before side effects, the set of capabilities the agent can invoke at all, and isolation of untrusted text from the privileged component. Prompt-level measures are a first filter, never load-bearing.
solid answer
~50 sSplit the catalogue by what a control depends on. Prompt-level measures — instruction/data discipline, spotlighting, telling the model to report embedded imperatives — all depend on the model cooperating, so they reduce how often the system is tested but cannot be cited as the reason a feature is safe. Deterministic controls run in the harness and hold regardless of what the model concluded: gating side effects behind human approval, constraining which capabilities exist in the loop, and keeping untrusted prose away from the component that holds them. Design the system on the assumption that the prompt layer fails, then ask what the model could do at that moment; if the answer is unacceptable the fix is capability or architecture, never stronger wording. And do not expect stacked prompt-level measures to multiply — their failures are correlated, because one adaptive payload can defeat several at once.
go deeper
Know that prompt wording helps but is not the safety net, and that the important protections live in the code around the model — what it is allowed to call and what stops for a human.
Sort the catalogue yourself: name which controls depend on the model cooperating and which run in the harness, and explain why the second group is what you cite when asked why a feature is safe.
Design from the side-effect inventory rather than the prompt: enumerate what the agent can cause, decide per capability whether it exists, whether it needs review, and whether the component needs raw prose at all.
Own the residual-risk conversation, including declining features that have no safe configuration, and make the posture durable through declared tool classifications, per-request mitigation logging and an inventory of untrusted-content entry points.
## The sorting question For each control in the catalogue ask one thing: does it still hold if the model is fully persuaded by something it read? That single question sorts the taxonomy into two piles, and the sorting is the whole of the design judgment. **Probabilistic, model-dependent.** Instruction/data separation in the prompt. Spotlighting — delimiting, datamarking, encoding. Instructions to report embedded imperatives rather than act on them. Classifier passes that score input or output for injection attempts. All of these reduce the *rate*. None of them survive a persuaded model, and a classifier is itself a model with the same failure mode. **Deterministic, model-independent.** Human approval enforced at the side-effect boundary. The set of capabilities present in the loop at all. Isolation of untrusted prose from the privileged component. Policy checks on resolved arguments before execution. These run in code; they hold whether or not the model was fooled. The design rule follows directly: put on the deterministic side everything whose failure you could not accept, and use the probabilistic side to reduce how frequently the deterministic side is exercised. ## Why stacking prompt-level measures disappoints A reasonable-sounding argument says three mitigations at 90% each leave 0.1% through. That arithmetic assumes independence, and these failures are not independent. They share a mechanism — the model choosing to honour a convention described in the same context as the attacker's text — so a payload crafted against one frequently defeats the others in the same pass. Worse, an adversary iterates: a static defence facing repeated attempts loses most of its per-attempt margin. Layering is right; expecting the layers to multiply is not. Layers should be *diverse in mechanism*, and the diversity that matters is prompt-level versus code-level, not two more paragraphs of prompt. ## Making the decision concrete Start from the side effects, not the prompt. Enumerate what the agent can cause: what it can write, send, spend, delete, or transition. For each, ask what the worst plausible outcome is if it fires with attacker-chosen arguments, and whether it is reversible. That inventory drives everything. Then decide, per capability: - Should this exist in the loop at all? Removing a capability is the only control with no residual failure rate. Many agents hold tools present "for completeness" that no user journey needs. - If it exists, does it fire without a human? Irreversible and externally visible actions should not. - Does the component holding it need to read untrusted prose, or can a constrained record serve? If a schema suffices, the channel closes. - What does the harness verify about the arguments before executing, independent of the model's intent? Only after those four does prompt-level hardening enter, as hygiene applied uniformly to every untrusted span including tool results. ## Residual risk, stated honestly A principal-level answer says out loud that this is risk reduction, not elimination, and names what remains. Approval bounds side effects but not information flow. Isolation shrinks the channel to the size of the schema. Capability limits do not help against misuse of a capability the feature genuinely requires. Some feature requests — an autonomous agent reading arbitrary third-party content and taking irreversible external actions without review — do not have a safe configuration in the current state of the field, and the correct output of the design review is a narrower feature, not a longer prompt. ## Organisational shape Three things make the posture durable. First, make the classification of a tool as side-effecting a property declared in code and reviewed like a schema change, so it cannot drift. Second, log which mitigations were applied per request, so an incident review can distinguish "the layer failed" from "the layer was not there." Third, keep an inventory of untrusted-content entry points; the usual regression is a new tool that fetches external pages into a context that was previously fed only vetted material. As of mid-2026 the industry reference taxonomy for this class of risk in agent systems is the OWASP Top 10 for Agentic Applications, whose leading entry is goal hijack — useful shared vocabulary in a design review, and a reminder that the field treats these as architectural rather than prompt-quality concerns. ## Interview framing Give the sorting question first, name both piles, then state the assumption you design under: the prompt layer fails, so the question is what the model can do at that moment. Close with the honest limit — some capabilities should not be behind an agent at all — because a candidate who claims a complete defence is the one who has not run this in production.
- How would you argue against a proposal to add a third prompt-level defense instead of a code-level one?On correlation. The existing prompt measures already share the mechanism the new one uses — the model honouring a convention stated alongside attacker text — so the marginal reduction is far below what its standalone rate suggests, and it is measured against payloads that predate it. The same effort spent narrowing a tool's scope or adding an approval yields a control whose failure rate does not depend on the model at all.
- What would you monitor to know whether this posture is holding in production?Per-request records of which mitigations were applied, so incidents separate "layer failed" from "layer absent"; approval rate against time-to-decide, which detects the fatigue that hollows out the strongest layer; the inventory of untrusted-content entry points, since new tools quietly add them; and the count of capabilities reachable without human review, which should trend down, not up.
- Is there a class of feature you would decline to build rather than mitigate?Yes — the combination of reading arbitrary third-party content, holding access to sensitive data, and taking irreversible externally visible actions with no review. Each control shaves the risk without closing it, and stacking them does not compose to a guarantee. The design output should be a narrower feature: review on the irreversible step, or a constrained record instead of raw content.
saying these in an interview costs you the question
- Treating a longer system prompt as a security control
- Assuming layered prompt mitigations multiply their effectiveness
- Adding an injection classifier and calling the risk closed
- Never asking what the model could do once persuaded
- Claiming a configuration that eliminates prompt injection