A model card reports a near-zero attack-success rate on AdvBench, a public corpus of harmful-behaviour prompts. What could produce that number without the model being any harder to jailbreak?
answer
- public list becomes refusal-tuning data
- guard and target share training text
- errors stack, do not cancel
- rewording delta is the real signal
- score = recall of a list
basics
~20 sNothing prevents the corpus being inside safety tuning. Those exact prompts appear in refusal training data, and often in the guard classifier that scores the run, so the model refuses memorised strings. Reword the same requests and the rate typically climbs. The number measures recall of a public list.
solid answer
~50 sTwo absorption paths, and they compound. **Safety tuning.** A public attack corpus is a ready-made set of prompts labelled 'should refuse'. It is the cheapest refusal-training data available, so it ends up in the tuning mix. The model then refuses those surface forms specifically — not the underlying behaviour class. **The judge.** Whatever decides an attempt succeeded — a guard classifier, a harm-judge model — is usually trained on the same public labelled text. It flags those prompts with unusual confidence, so borderline compliant responses get scored as refusals. Both errors push the attack-success rate down, so they do not cancel; they stack. The diagnostic is a rewording test: express the same harmful behaviours in fresh phrasings and re-measure. If the rate jumps, the original score was memorisation. A model whose refusal generalises shows a small gap. Treat the gap itself, not the headline number, as the interesting quantity.
go deeper
Says the prompts are public so the model was probably trained to refuse them, and the score therefore overstates safety.
Separates the two absorption paths — refusal tuning and the scoring guard — and proposes the rewording delta as the test.
Also re-scores the original transcripts with an independent judge, reads per-prompt refusal shape, and reports the delta rather than the headline.
Sets the org rule for which number may appear in a claim, and treats contamination as expected engineering rather than misconduct when negotiating with vendors.
Start by naming what the number is: the fraction of prompts from one public list, judged by one scorer, on which the model produced something that scorer called compliant. That is the whole content of an **attack-success rate** (ASR). Three clauses, and each is attackable. ## Why this corpus in particular gets absorbed Capability-benchmark contamination is usually accidental — a web crawl swallowed a test set that happened to be online. Attack-corpus contamination is mostly **deliberate and correct**. A safety team that finds a published list of harmful requests has an obligation to make the model refuse them; AdvBench's `harmful_behaviors` split is a few hundred pre-labelled *should refuse* examples, which is the cheapest refusal-tuning data available anywhere. So the overlap is not merely likely, it is close to certain, and it is not a scandal. That is exactly why the resulting score cannot be read as robustness: the right engineering decision destroyed the measurement. ## Why the guard makes it worse Whatever decides an attempt succeeded — a shipped guard classifier, a harm-judge model, HarmBench's released classifier — was itself trained on labelled harmful text, and public safety corpora are where that text comes from. If the vendor's refusal tuning and the vendor's guard drew on the same lists, the evaluation is closed-loop: the thing under test and the thing doing the testing share priors. Where that bites is the interesting middle of the distribution. A response that is subtly compliant — hedged, partial, wrapped in a fictional frame, correct in outline but vague in detail — is precisely the case where an independently trained judge and the vendor's own guard will disagree, and the vendor's guard will read low. So the target's memorisation pushes ASR down and the judge's confidence pushes ASR down again. Two biases in the same direction stack; they do not cancel. ## What it costs to find out The diagnostic is cheap relative to how much weight the number carries. Re-running a few hundred behaviours in two arms — verbatim and re-authored — at a handful of samples each is a few thousand short completions plus one grader call apiece: single-digit to low-tens of dollars on a hosted endpoint, under an hour of wall clock, and a scanner such as garak scales it linearly through `--generations`. The expensive input is human: authoring paraphrases that hold the behaviour fixed while sharing no distinctive phrasing is skilled work, on the order of days, and re-scoring with an independent judge only means anything if a human-adjudicated sample calibrates that judge too. ## The specific misreading The wrong sentence is *"ASR on AdvBench was 0.4%, so the model is hard to jailbreak."* What the evidence supports is *"on strings the model was almost certainly trained to refuse, graded by an instrument trained on the same strings, it refused."* Those are different claims, and only the second is defensible. Two adjacent misreadings are just as common: | Claim made | What the data actually licenses | |---|---| | "The vendor cheated." | Nothing. Training on published harmful prompts is the expected engineering choice. | | "A bigger public corpus fixes it." | Nothing. More contaminated prompts narrow the confidence interval around a systematically biased estimate. | | "It scored well on our rewrite too, so it is robust." | Only if the rewrite shares no surface cues — a synonym swap preserves the very thing under test. | ## What you would check, in order 1. **Rewording delta.** Re-express the same behaviours in phrasings that share no distinctive n-grams with the published prompts, interleave the two arms in one run window so unannounced endpoint changes hit both equally, and report the delta. 2. **Independent re-scoring.** Freeze the transcripts from the original run and re-grade them with a judge the vendor did not train, plus a human-labelled sample as tiebreak. A large one-directional disagreement is a judge-contamination signal, separate from the model. 3. **Refusal shape.** Near-instant, template-identical refusals across an entire published list look like recall. Refusals that vary with how severe the request actually is look like generalisation. 4. **Dates.** A corpus published comfortably before the model release had time to be absorbed; one released after it did not. ## What you may honestly claim afterwards Not "the model is robust" and not "the vendor cheated". The defensible write-up is: on prompts drawn verbatim from a public corpus the attack-success rate was A; on matched behaviours re-authored by us it was B; and B is the number that describes what a real attacker faces. The gap B minus A is your contamination estimate, and it is more informative to a decision-maker than either figure alone.
- The vendor argues that refusing published harmful prompts is exactly what safety tuning should do, so the score is legitimate. What is your reply?Agree on the engineering and refuse the inference. Training on them is correct behaviour; it just makes those prompts unusable as a measurement. Both things are true at once.
- You reword the prompts and the attack-success rate barely moves. What does that tell you?Either the refusal genuinely generalises over that behaviour class, or your rewording preserved the surface cues the model keys on. Check the second before believing the first.
- Which number would you put in the summary line of a report?The held-out or reworded rate, with the public-corpus rate beside it and a one-line note that the target and its guard were likely trained on that corpus.
It is like grading a spelling test on the exact twenty words the class was drilled on that morning, with an examiner who was drilled on the same twenty. A perfect score tells you the drill took hold, not that anyone can spell.
saying these in an interview costs you the question
- Reads a low attack-success rate on a public corpus as a robustness claim.
- Assumes contamination implies vendor misconduct rather than ordinary safety engineering.
- Only considers the model, never that the scoring guard saw the same corpus.
- Proposes 'just use a bigger public corpus' as the fix.
- Cannot name a concrete check that distinguishes memorisation from generalisation.