Why does a fixed red-team prompt suite overstate an LLM feature's real safety?
answer
- fixed list, moving target
- the attacker moves second
- regression test, not a safety measure
- passing proves patched, not safe
- published defenses fell to adaptive re-attack
basics
~20 sA fixed suite only measures attacks someone already wrote down. A real attacker sees your defense and adapts, so passing the suite proves the known attacks are patched — not that the system resists new ones.
solid answer
~50 sA static suite is a list of prompts plus a check for whether the system did something it should not. Passing it answers a narrow question — do the attacks I already catalogued still work? — which is a **regression** result, not a safety result. Attackers move second: they see the deployed behaviour, probe it, and shape a new payload around whatever blocked the old one, so the fixed list is a target that never moves while the threat does. This is why the 2025 adaptive-attack literature was so damaging: defenses that reported near-zero success against static evaluations were driven past 90% attack success once the attacker was allowed to adapt to them. Keep the suite as a cheap regression gate, but report safety as an attack-success rate measured by an adaptive attacker with a stated budget, and supplement it with autonomous auditing that explores behaviours nobody wrote a probe for.
go deeper
Know the difference between a test suite you wrote and an adversary who writes new inputs. Be able to say that passing a red-team suite means the listed attacks were blocked, nothing more.
Explain why an attack distribution is not stationary the way a quality eval set is, and why that makes a fixed suite a regression gate rather than a safety measurement. Name resampling and adaptive attacks as what you add.
Show you would restructure reporting: attack-success rate against a stated attacker budget, per severity category, plus a permanent regression case for every past incident. Be ready to say which defenses you treat as probabilistic rather than as boundaries.
Own the argument that no static number justifies a launch decision, and set what evidence does: an ASR-versus-budget curve, a bounded worst case if it fails, and a response plan. Expect to defend spending scarce red-team time on adaptive and autonomous exploration instead of growing the list.
## What a static suite actually measures A static red-team suite is a fixed collection of adversarial inputs — jailbreak prompts, injected documents, poisoned tool outputs — each paired with a checker that decides whether the system produced the forbidden behaviour. Running it yields one number: the fraction of listed attacks that succeeded. That number is meaningful, but it answers a narrow question: *do the attacks I already know about still work against this build?* Interviews go wrong when a candidate reports that number as an answer to the much larger question, *is this feature safe to ship?* The gap between those two questions is the entire subject. A benefits-eligibility assistant that answers questions about a household's case file might be red-teamed for three weeks before launch against a curated suite of a few hundred prompts. If it passes every one, what has been established is that those few hundred inputs do not produce a leak or a wrong eligibility claim today. Nothing has been established about the input the first motivated attacker writes next week. ## Why the target moves Security evaluation differs from quality evaluation in one structural way: the distribution of inputs is chosen by an adversary who observes your system. Quality evals sample from something like the real user distribution, and that distribution is roughly stationary, so a fixed golden set stays representative. An attack distribution is not stationary — it is a best response to whatever you deployed. Every published defense becomes, on release, a description of what the next payload has to avoid. This is what the field means by *the attacker moves second*. A defense evaluated against attacks written before the defense existed is being graded on a test the attacker never has to sit. The 2025 adaptive-attack work made the effect concrete: taking twelve published jailbreak and prompt-injection defenses, each of which reported low attack-success rates against its own static evaluation, and re-attacking each one with methods tuned specifically to it, pushed success above 90% in most cases. The resulting consensus as of 2026 is that detection classifiers, delimiter-based spotlighting and instruction-hierarchy training are probabilistic nudges that raise attacker cost, not boundaries that hold. Vendor numbers show the same shape from the other direction: a frontier assistant reported around 0.1% success against a single attempt, rising to roughly 5–6% once the attacker was allowed a hundred adaptive attempts against the same target. ## What the suite is still worth None of this makes static suites useless — it makes them the wrong thing to *report*. They are excellent as: - **Regression gates.** Anything that ever succeeded in production becomes a permanent case, so the same failure cannot ship twice. - **Policy coverage.** One case per category in your written policy proves you at least exercise each category. - **Cheap per-commit signal.** They are fast and deterministic enough to run on every release candidate, which adaptive campaigns are not. The honest framing is: the static suite is a floor test. Failing it is decisive; passing it is not evidence of safety. ## What to add on top Three things carry the actual evidence. First, **adaptive attacks with a declared budget** — a human or automated attacker allowed to iterate against this specific build, reporting success as a curve against attempts rather than a single pass/fail. Second, **resampling**: rerunning the same nominal attack many times with perturbations, because a defense that blocks an attack 95% of the time is not blocking it at all when the attacker can retry. Third, **autonomous auditing agents** such as Petri, which explore the deployed system on their own and surface behaviours no written probe anticipated — the failure you did not think to test for is precisely the one a fixed list cannot contain. ## How to say it in an interview Do not claim a pass rate. Say what the attacker was allowed to do: "no adaptive attacker achieved data exfiltration within N attempts against this build, and the categories where success did occur are X and Y at these rates." That sentence is falsifiable, it names the budget, and it is the only form of safety claim that survives contact with someone who has read the adaptive-attack literature. Then say what happens when it fails anyway — staged rollout, monitoring, containment — because a residual rate you have planned for is a much stronger position than a pass rate you cannot defend.
- What would make you keep a static suite in the pipeline at all?Regression value and cost. It is fast, deterministic and cheap enough to run on every release candidate, and it guarantees that any attack which once succeeded in production can never ship again. It also gives coverage evidence that each written policy category is exercised. What it cannot do is generalise to attacks it does not contain, so it gates rather than certifies.
- An autonomous auditing agent finds a behaviour none of your probes anticipated. What do you do with it beyond fixing it?Treat the finding as two artefacts. The fix addresses the instance; a new permanent regression case addresses recurrence. Then ask what class the finding belongs to and whether your probe categories had a blind spot — a behaviour nobody wrote a probe for usually means a missing category in the policy, not just a missing prompt. Auditing agents are worth most when their output feeds category design.
- How does this change if the attacker cannot see your system prompt or defenses?Less than people hope. Attackers infer defenses from behaviour — which inputs get refused, which get sanitised, which produce a stock message — and iterate from there, so obscurity buys attempts, not immunity. It is also a poor assumption for any product with a public interface, and it collapses entirely once the system prompt leaks. Plan for a defense whose description is known.
saying these in an interview costs you the question
- Reports a 100% suite pass rate as proof the feature is safe
- Believes adding enough static prompts eventually covers the attack space
- Treats a detection classifier's score as a security boundary
- Assumes the model vendor's safety training removes the need to test the product
- Never states how many attempts the attacker was allowed