AdvBench and HarmBench are publicly released attack-prompt corpora used to score how often a model refuses. Why do red teams keep part of their own attack prompt set unpublished?
answer
- published set becomes training data
- public corpus for continuity, private for truth
- guard classifiers train on the same lists
- held-out costs authoring plus labels
- score moves on rewording
basics
~20 sBecause anything published gets trained on. Vendors add public attack prompts to safety-tuning data and to guard classifiers, so models learn to refuse those exact prompts. The score rises without the model becoming harder to attack. An unpublished set stays a fair test, since nobody could have fitted to it.
solid answer
~50 sA published attack corpus stops being a test the moment it becomes training data. Two things absorb it. First, safety tuning: the prompts are cheap, labelled negative examples, so they land in refusal fine-tuning sets. Second, the classifiers that sit in front of and behind a model — the input/output guards — are trained on the same public lists, because that is where labelled harmful text is. The consequence is that a low attack-success rate on that corpus mostly proves the target recognises those strings. Reword the same request and the number moves. So teams split their material: public corpora for continuity and cross-vendor comparison, a private held-out set for the number they actually believe. The tradeoff is that the private set is expensive to author, cannot be compared against anyone else's results, and decays as it leaks through logs and reports.
go deeper
Says that published prompts get trained on, so a score against them overstates robustness, and that some prompts must stay private.
Names both absorption paths — safety-tuning data and the guard classifier's training set — and knows the rewording test that exposes memorisation.
Runs both sets, reports the gap between them as the contamination estimate, and owns the hygiene rules that keep the private set usable.
Budgets the authoring, labelling and refresh cost, decides who inside the org may see the held-out set, and sets the policy that keeps it out of vendor logs.
## What a released attack corpus actually contains A public attack corpus is two artefacts bolted together. The first is a list of request strings a well-behaved assistant is expected to decline: AdvBench's `harmful_behaviors` split is on the order of 500 of them, and HarmBench and JailbreakBench publish curated behaviour lists of similar shape. The second is a **grading convention** — a refusal-phrase heuristic, or, for HarmBench and JailbreakBench, a released classifier that decides whether a response actually carried the behaviour out. Both halves are downloadable. Both halves are what gets absorbed. The number these produce, the **attack-success rate** (ASR), is simply: of the prompts on that list, the fraction on which the grader said the model complied. Every word of that definition is attackable, because the list is public and so is the grader. ## The two absorption paths **Safety tuning.** A list of prompts already labelled *should refuse* is the cheapest refusal-training data that exists. It is pre-written, pre-labelled, and free. It lands in the refusal fine-tuning mix, and what the model then learns is a mapping from those particular surface forms to a refusal — a narrower thing than learning to decline the behaviour class behind them. **The guard.** The input and output moderation classifiers that sit around a deployed model need labelled harmful text to train on, and public safety corpora are where labelled harmful text lives. So the same strings end up inside the instrument that later scores the run. Neither move is misconduct. A safety team that finds a published list of harmful requests and does *not* make the system refuse them has failed at its job. That is the uncomfortable part: doing the correct engineering is precisely what destroys those prompts as a measurement instrument. ## Why the two errors stack rather than cancel A contaminated target refuses prompts it has memorised, which pushes ASR down. A contaminated grader is unusually confident on those same prompts and reads hedged or partially compliant answers as refusals, which also pushes ASR down. Two biases pointing the same way compound; there is no averaging-out to hope for. ## What a run costs, and what the replacement costs The measurement is cheap. Roughly 500 prompts at, say, five samples each is about 2,500 completions per arm — a scanner such as garak multiplies exactly this way through its `--generations` flag — plus one grader call per completion. On a hosted endpoint that is single-digit to low-tens of dollars and well under an hour of wall clock. Nothing about the price signals that the number is worthless. The expensive artefact is the replacement. A held-out set costs: - **Authoring.** Someone who knows the attack families has to cover the *behaviour* space, not the phrasing space. Order of one to two engineer-weeks for a few hundred usable items, and it recurs. - **Ground truth.** Each behaviour needs a written definition of what a successful response looks like, or nothing can score it consistently. - **Adjudication.** A human-labelled sample each cycle to keep whatever automated scorer you use calibrated. - **Decay.** Every engagement erodes the set — prompts sent to endpoints that retain inputs, examples pasted into tickets, the tuning team gradually optimising against it. Budget a refresh fraction per cycle, not a one-off build. ## Where the number misleads The specific wrong reading is: *ASR near zero on AdvBench, therefore the model is hard to jailbreak.* What the number supports is far weaker — on strings the system was very likely trained to refuse, graded by an instrument trained on the same strings, it refused. Rewording the same requests typically moves the rate, sometimes a lot. A second wrong reading is treating a private set as automatically clean: a held-out set built by paraphrasing the public one inherits the same surface forms and measures the same memorisation under a new filename. ## What you check Interleave a re-authored arm with the verbatim arm in one run window and report the gap. Re-score the original transcripts with a grader the vendor did not train. Look at refusal *shape* — flat, template-identical refusals across an entire published list read as recall; refusals that vary with how severe the request is read as generalisation. Check the corpus publication date against the model release date. Then report both numbers, labelled by provenance, and let the gap carry the argument.
- If your held-out set cannot be compared against other vendors' published numbers, what is it actually for?Internal decisions: does this release regress against the last one, and does this deployment clear our bar. Comparability across vendors is what the public corpus is for, with its caveat attached.
- How does a held-out attack set leak in practice?Through prompts sent to hosted endpoints that retain and train on inputs, through examples pasted into reports and tickets, and through the tuning team getting access and fitting to it.
saying these in an interview costs you the question
- Treats a low score on a public attack corpus as proof of robustness.
- Says contamination does not matter because 'the model still refused'.
- Proposes building the held-out set by paraphrasing the public one, which keeps the same surface forms.
- Wants to publish the held-out prompts to be transparent, burning the set.
- Never considers that the judge or guard classifier is contaminated too.