You run a long-published, fixed-attack jailbreak benchmark against a hosted chat endpoint that the vendor has patched many times, and it reports a very low attack-success rate. What are the reasons that number can badly understate your real exposure, and what would you run alongside it?
answer
- public prompts get patched first
- contamination: model saw the exam
- high score meaningful, low score weak
- exact vs paraphrase gap
- pair with fresh + application-level attacks
basics
~20 sIts prompts are public and old, so vendors have likely patched exactly those strings and they may sit in training or filter data. A low score then proves those specific attempts fail, not that the model resists new ones. Pair it with freshly generated attacks and attacks aimed at your own application.
solid answer
~50 sThe freeze that makes a suite comparable is the same freeze that makes it age. Three mechanisms drive the optimistic drift: 1. **Direct patching.** Public attack strings are the cheapest thing for a vendor to block, by filter rule or by fine-tuning. The suite becomes a list of things already fixed. 2. **Contamination.** The prompts appear in papers, repos and scrapes, so they can end up in training or safety-tuning data. The model has effectively seen the exam. 3. **Paraphrase brittleness.** A defence that blocks the exact frozen phrasings often fails on lightly reworded variants, and the suite never asks that question. So a low score licenses one narrow claim: *these* attempts, under *this* judge, mostly fail. Alongside it, run automatically generated fresh attacks, mutated variants of the suite's own behaviours, and attacks against your own application surfaces — system prompt, tools, retrieved content — which no public suite covers.
go deeper
Says the prompts are public and old, so the vendor has probably fixed those specific ones and the low score may not generalise.
Separates direct patching, training-data contamination and paraphrase brittleness, and proposes running fresh attacks too.
States the asymmetry — high score is strong evidence, low score is weak — and designs the variant-versus-exact comparison that exposes surface patching.
Decides what the org may claim externally from a frozen number, and funds a continuously refreshed attack corpus so the tripwire is not the only signal.
## The freeze that buys comparability is the same freeze that causes ageing A public, fixed attack suite is a **stationary target** that the entire ecosystem gradually optimises against — partly on purpose, mostly as a side effect. Nothing dishonest has to happen for its number to drift optimistic, and the drift never appears in the number itself. A score of 2% looks identical whether the model genuinely became robust to that class of attack or whether the exam simply leaked. ## Three mechanisms, and they are distinct - **Direct patching.** Published attack strings are the cheapest possible thing for a vendor to defend against: a filter rule, a blocklist, a small round of safety fine-tuning on those exact phrasings. This is rational behaviour, and it turns the suite into a list of things already fixed. - **Contamination.** The prompts live in papers, public repositories and dataset releases, so they are scraped and can end up inside pretraining, instruction-tuning or safety-tuning corpora. The model has effectively seen the exam paper. This one is invisible from the outside and usually unfalsifiable without vendor cooperation. - **Paraphrase brittleness.** A defence that matches surface phrasings will block the frozen strings and fail on lightly reworded variants of the same behaviour. The frozen suite never asks the paraphrase question, so this weakness is structurally outside its field of view. ## The asymmetry that is the whole practical rule A **high** score is strong evidence: known public attacks still work against your endpoint, which is a hard floor on your exposure and normally an immediate, easy-to-defend finding. A **low** score is weak evidence, and it gets weaker every month the suite stays public. Stated compactly: a stale suite is an **excellent failure detector** and a **poor safety certificate**. Almost every misuse of one of these numbers is somebody treating the second half as if it were the first. ## How to tell whether the freeze, rather than the model, is doing the work 1. Run a *variant arm*: the same behaviours the suite targets, expressed in phrasings its authors and the vendor never saw — reworded, restructured, translated, or reformatted. Compare it against the exact frozen phrasings, judged by the same judge so the grader is held constant. The interesting output is not either raw rate but the **gap** between them. A large frozen-to-variant gap is the signature of surface-level patching; a small gap suggests the defence is operating on the behaviour rather than the string. 2. Second, run a generation-based attack loop that searches fresh attacks against the same behaviours, and compare its success rate against the frozen suite's. 3. Third, track the frozen suite's score across model updates over time: a number that only ever drops while your own freshly generated attacks stay flat is a strong indication the suite is being fitted rather than the model getting safer. ## What the extra arms cost, and why teams skip them - **The frozen run is the cheap arm** — a few hundred to a few thousand target calls, minutes of wall clock, single-digit dollars, and it needs no expert once the adapter exists. - **The variant arm** roughly doubles that, which is negligible. - **The generation-based arm** is the one with real cost: an optimisation loop typically spends tens to low hundreds of queries per behaviour, so a hundred behaviours becomes thousands to tens of thousands of calls, hours of wall clock, and a bill one or two orders of magnitude above the frozen run — plus rate limits, plus the account and legal clearance to run sustained adversarial traffic against a hosted endpoint. That gap in cost is exactly why the cheap, misleading arm is the one that gets run and quoted, and it is worth naming when you ask for the budget. ## Where the number misleads Beyond reading a low score as robustness: - a **flat score across releases** can mean the bar fell rather than the product held, since the suite ages while the model is patched; - a **bare-model figure** quoted for a deployed product ignores the surfaces that actually carry your risk — the system prompt, the tool-calling path, retrieved documents, multi-turn state; - and "we scored better than last quarter" is not a safety improvement if the only thing that changed is that more of the published strings are now blocked by name. ## What you run alongside it, and how you report it - Freshly generated attacks against the same harm behaviours; - application-surface attacks you own and version yourself; - and a small human red-team pass on the highest-consequence flows. Report the frozen suite's number *with* its revision and its age, next to the variant-versus-exact gap and the fresh-attack result, and state in the same sentence that the frozen figure is a **regression signal** and not a coverage claim.
- The frozen suite reports near zero, but your generated attacks succeed often. What does that tell you?That the defence is fitted to known public strings rather than to the behaviour. Report the gap; it is the finding.
- Is a stale suite worth running at all?Yes, as a cheap regression tripwire and as a floor: if public attacks still work, that is an immediate, high-confidence finding.
It is a past exam paper that has been in circulation for years. A high score on it proves very little, but a failing mark still tells you something real — and the class average rising every year mostly means the paper leaked, not that the students got smarter.
saying these in an interview costs you the question
- Reading a low frozen-suite score as evidence the model is robust.
- No awareness that public prompts can leak into safety-tuning data.
- Never testing reworded variants of the suite's own behaviours.
- Reporting the suite figure without its revision or age.