skip to content

Why is a prompt-injection detection classifier not a security boundary?

level: seniorimportance: must knowfreq 58%

answer

  1. a percentage is not a boundary
  2. the defender publishes, then the attacker adapts
  3. static suites overstate robustness
  4. one miss is enough when retries are free
  5. ask what happens when the filter says benign

basics

~20 s

Detection is probabilistic and only has to be wrong once, while an attacker can retry and tune against it. Adaptive attacks against a dozen published defenses succeeded over 90% of the time. Filters reduce noise; a deterministic architectural limit is the boundary.

solid answer

~50 s

A classifier gives you a probability, and an attacker who can retry converts any nonzero miss rate into a success. That is not theoretical: work published in 2025 ("The Attacker Moves Second", Nasr, Carlini and colleagues) re-attacked twelve published jailbreak and injection defenses with attacks tuned against each one and pushed success rates above 90%, against defenses whose original papers reported near-total robustness on static suites. The same asymmetry shows up in vendor reporting — a model may sit near 0.1% success on a single attempt but several percent after a hundred adaptive attempts. So detection, instruction-hierarchy training and content tagging are all *nudges*: they lower per-attempt success and they are worth having. The boundary has to be something deterministic underneath — a capability the session simply does not hold, a fixed control flow, or a policy check on values — so that a successful injection still cannot reach a consequential action.

go deeper

for a junior

Know that a detector can be wrong and that attackers retry, so filtering text is not the same as preventing an action. Say plainly that you would also limit what the agent is able to do.

for a middle

Explain the difference between per-attempt success rate and success under a retry budget, and why a defense evaluated on a fixed attack suite reports an optimistic number. Name at least one deterministic control you would put underneath the filter.

for a senior

Argue from adaptive-attack evidence rather than intuition, and be ready to redesign a feature so a full bypass is survivable. Demonstrate the review habit of asking what the worst outcome is if the classifier returns benign on every input.

for a principal

Set the organisational rule about what probabilistic controls may and may not license, and decide where the deterministic boundaries live across many teams. Be candid with stakeholders that residual risk is nonzero and that the mitigation is blast radius, not a better filter.

## What a detector actually gives you An injection detector — a fine-tuned classifier, a small guard model, a heuristic over suspicious phrasing — outputs a score. You pick a threshold and trade false negatives against false positives. On a benign corpus with a fixed attack suite this looks excellent: reported detection rates in the high nineties are normal. The number is real, and it is also the wrong number to design around. ## Why the attacker moves second Security evaluations of defenses are usually run against a **static** set of attacks that existed before the defense did. The attacker in production does not work that way. They see the deployed system, probe it, and adapt — rephrasing, encoding, splitting an instruction across documents, burying it in a format the classifier scores as benign. The 2025 study "The Attacker Moves Second" made this concrete by re-testing twelve published defenses with attacks optimised specifically against each, using human red teams and automated search. Most defenses that had reported robustness collapsed past 90% attack success. The generalisable lesson is not that those twelve were bad work; it is that a static-suite number is an upper bound on robustness, and adaptive success rate is the number that matters. ## The arithmetic of retries Suppose your stack fails on 0.1% of attempts — an excellent per-attempt figure. An attacker who can send a hundred emails, open a hundred issues, or publish a hundred pages faces roughly a 10% chance overall, and they only need one. Vendor containment write-ups report exactly this shape: single-attempt success near a tenth of a percent, rising to several percent under a hundred adaptive attempts. Wherever the attacker controls the retry budget, per-attempt probabilities are close to meaningless. ## Two structural gaps detectors cannot close First, **direct injection through the user's own input** bypasses model-layer training entirely: instruction-hierarchy training teaches the model to privilege the developer's instructions over content, but the user turn is a legitimate high-trust channel, and a malicious user is inside it. Second, **the detector sees text, while the damage is done by actions**. A payload can be split so that no single document looks hostile, and the harmful behaviour only emerges from the composition of several benign-looking inputs across a long session. ## What a boundary looks like instead The distinguishing property of a boundary is that it holds *when the model has already been fooled*. Concretely: - **Absent capability.** The session that reads untrusted content has no credential for the private store and no outbound path. Nothing the model decides can change that. - **Fixed control flow.** The sequence of actions is committed before untrusted content enters, so injected text cannot introduce a step that was not planned. - **Policy on values.** Every value carries provenance and a permitted audience, and the check happens at the point of action, not in the model. - **Isolation.** Untrusted items are processed in separate contexts whose outputs are constrained to a narrow, typed shape. All four are deterministic: they are code, not inference. The useful formulation is that the deterministic boundary is what gets hit when everything probabilistic misses — so the design question is never "is the filter good?" but "what happens on the run where the filter is wrong?" ## Where filters still belong None of this means deleting your classifier. Detection is genuinely valuable as (a) noise reduction against the large volume of low-effort attacks, (b) a telemetry signal that tells you an attacker is probing, and (c) a rate-limiting trigger that raises the cost of the retry loop that makes adaptive attacks work. What you must not do is let a detector license a design that would be unsafe without it. The test is simple and worth asking in review: if this classifier returned "benign" for every input, what would the worst outcome be? If the answer is "our customer database leaves the building," the classifier was carrying weight it cannot carry. ## How to talk about this in an interview Do not overclaim in either direction. "Prompt injection is unsolved at the model layer" is the current honest position; "therefore nothing helps" is wrong. Layer the probabilistic defenses, expect them to be partially bypassed, and size the blast radius so that being bypassed is survivable. Interviewers are listening for whether you treat a percentage as a mitigation or as a boundary.

  • If detection is unreliable, why keep an injection classifier in production at all?
    Because it does useful work that is not boundary work. It filters the high volume of low-effort attempts, it gives you a telemetry signal that someone is probing a surface, and it can trigger rate limiting, which directly attacks the retry budget that adaptive attacks depend on. Keep it — just never let its presence justify an architecture that would leak data the moment it returns a false negative.
  • Does instruction-hierarchy training, where the model is taught to rank system over user over content, change your answer?
    It lowers per-attempt success, which is worth having, but it is still learned behaviour with a nonzero failure rate, so the retry arithmetic applies unchanged. It also has a structural gap: a hostile user occupies the user turn, which the hierarchy is trained to trust, so direct injection routes around it. Treat it as another nudge that sits above the deterministic boundary.
  • What would make you accept a probabilistic control as sufficient for a given feature?
    When the worst outcome of a full bypass is acceptable. If the session holds no private data and no irreversible action — say a public-content summariser — then a bypass yields a bad summary, and a classifier plus monitoring is a proportionate response. The judgement is about blast radius, not about the classifier's reported accuracy.

saying these in an interview costs you the question

  • Quotes a static-benchmark detection rate as proof the system is safe
  • Believes a sufficiently good system prompt makes injection impossible
  • Assumes attackers use published payloads rather than adapting to your defense
  • Thinks instruction-hierarchy training closes injection at the model layer
  • Stacks more probabilistic filters instead of removing a capability

context