skip to content

Contaminated Corpora

Once a published attack corpus sits in safety-tuning data and in a guard's own training set, the score improves while the model does not. Interviewers ask how you would spot that from a report alone.

on this pageshow

explore

questions

5

A model card reports a near-zero attack-success rate on AdvBench, a public corpus of harmful-behaviour prompts. What could produce that number without the model being any harder to jailbreak?

level: middleimportance: must knowfreq 62%

answer

  1. public list becomes refusal-tuning data
  2. guard and target share training text
  3. errors stack, do not cancel
  4. rewording delta is the real signal
  5. score = recall of a list

basics

~20 s

Nothing prevents the corpus being inside safety tuning. Those exact prompts appear in refusal training data, and often in the guard classifier that scores the run, so the model refuses memorised strings. Reword the same requests and the rate typically climbs. The number measures recall of a public list.

solid answer

~50 s

Two absorption paths, and they compound. **Safety tuning.** A public attack corpus is a ready-made set of prompts labelled 'should refuse'. It is the cheapest refusal-training data available, so it ends up in the tuning mix. The model then refuses those surface forms specifically — not the underlying behaviour class. **The judge.** Whatever decides an attempt succeeded — a guard classifier, a harm-judge model — is usually trained on the same public labelled text. It flags those prompts with unusual confidence, so borderline compliant responses get scored as refusals. Both errors push the attack-success rate down, so they do not cancel; they stack. The diagnostic is a rewording test: express the same harmful behaviours in fresh phrasings and re-measure. If the rate jumps, the original score was memorisation. A model whose refusal generalises shows a small gap. Treat the gap itself, not the headline number, as the interesting quantity.

go deeper

for a junior

Says the prompts are public so the model was probably trained to refuse them, and the score therefore overstates safety.

for a middle

Separates the two absorption paths — refusal tuning and the scoring guard — and proposes the rewording delta as the test.

for a senior

Also re-scores the original transcripts with an independent judge, reads per-prompt refusal shape, and reports the delta rather than the headline.

for a principal

Sets the org rule for which number may appear in a claim, and treats contamination as expected engineering rather than misconduct when negotiating with vendors.

Start by naming what the number is: the fraction of prompts from one public list, judged by one scorer, on which the model produced something that scorer called compliant. That is the whole content of an **attack-success rate** (ASR). Three clauses, and each is attackable. ## Why this corpus in particular gets absorbed Capability-benchmark contamination is usually accidental — a web crawl swallowed a test set that happened to be online. Attack-corpus contamination is mostly **deliberate and correct**. A safety team that finds a published list of harmful requests has an obligation to make the model refuse them; AdvBench's `harmful_behaviors` split is a few hundred pre-labelled *should refuse* examples, which is the cheapest refusal-tuning data available anywhere. So the overlap is not merely likely, it is close to certain, and it is not a scandal. That is exactly why the resulting score cannot be read as robustness: the right engineering decision destroyed the measurement. ## Why the guard makes it worse Whatever decides an attempt succeeded — a shipped guard classifier, a harm-judge model, HarmBench's released classifier — was itself trained on labelled harmful text, and public safety corpora are where that text comes from. If the vendor's refusal tuning and the vendor's guard drew on the same lists, the evaluation is closed-loop: the thing under test and the thing doing the testing share priors. Where that bites is the interesting middle of the distribution. A response that is subtly compliant — hedged, partial, wrapped in a fictional frame, correct in outline but vague in detail — is precisely the case where an independently trained judge and the vendor's own guard will disagree, and the vendor's guard will read low. So the target's memorisation pushes ASR down and the judge's confidence pushes ASR down again. Two biases in the same direction stack; they do not cancel. ## What it costs to find out The diagnostic is cheap relative to how much weight the number carries. Re-running a few hundred behaviours in two arms — verbatim and re-authored — at a handful of samples each is a few thousand short completions plus one grader call apiece: single-digit to low-tens of dollars on a hosted endpoint, under an hour of wall clock, and a scanner such as garak scales it linearly through `--generations`. The expensive input is human: authoring paraphrases that hold the behaviour fixed while sharing no distinctive phrasing is skilled work, on the order of days, and re-scoring with an independent judge only means anything if a human-adjudicated sample calibrates that judge too. ## The specific misreading The wrong sentence is *"ASR on AdvBench was 0.4%, so the model is hard to jailbreak."* What the evidence supports is *"on strings the model was almost certainly trained to refuse, graded by an instrument trained on the same strings, it refused."* Those are different claims, and only the second is defensible. Two adjacent misreadings are just as common: | Claim made | What the data actually licenses | |---|---| | "The vendor cheated." | Nothing. Training on published harmful prompts is the expected engineering choice. | | "A bigger public corpus fixes it." | Nothing. More contaminated prompts narrow the confidence interval around a systematically biased estimate. | | "It scored well on our rewrite too, so it is robust." | Only if the rewrite shares no surface cues — a synonym swap preserves the very thing under test. | ## What you would check, in order 1. **Rewording delta.** Re-express the same behaviours in phrasings that share no distinctive n-grams with the published prompts, interleave the two arms in one run window so unannounced endpoint changes hit both equally, and report the delta. 2. **Independent re-scoring.** Freeze the transcripts from the original run and re-grade them with a judge the vendor did not train, plus a human-labelled sample as tiebreak. A large one-directional disagreement is a judge-contamination signal, separate from the model. 3. **Refusal shape.** Near-instant, template-identical refusals across an entire published list look like recall. Refusals that vary with how severe the request actually is look like generalisation. 4. **Dates.** A corpus published comfortably before the model release had time to be absorbed; one released after it did not. ## What you may honestly claim afterwards Not "the model is robust" and not "the vendor cheated". The defensible write-up is: on prompts drawn verbatim from a public corpus the attack-success rate was A; on matched behaviours re-authored by us it was B; and B is the number that describes what a real attacker faces. The gap B minus A is your contamination estimate, and it is more informative to a decision-maker than either figure alone.

  • The vendor argues that refusing published harmful prompts is exactly what safety tuning should do, so the score is legitimate. What is your reply?
    Agree on the engineering and refuse the inference. Training on them is correct behaviour; it just makes those prompts unusable as a measurement. Both things are true at once.
  • You reword the prompts and the attack-success rate barely moves. What does that tell you?
    Either the refusal genuinely generalises over that behaviour class, or your rewording preserved the surface cues the model keys on. Check the second before believing the first.
  • Which number would you put in the summary line of a report?
    The held-out or reworded rate, with the public-corpus rate beside it and a one-line note that the target and its guard were likely trained on that corpus.

It is like grading a spelling test on the exact twenty words the class was drilled on that morning, with an examiner who was drilled on the same twenty. A perfect score tells you the drill took hold, not that anyone can spell.

saying these in an interview costs you the question

  • Reads a low attack-success rate on a public corpus as a robustness claim.
  • Assumes contamination implies vendor misconduct rather than ordinary safety engineering.
  • Only considers the model, never that the scoring guard saw the same corpus.
  • Proposes 'just use a bigger public corpus' as the fix.
  • Cannot name a concrete check that distinguishes memorisation from generalisation.

context

open as a page

You are handed only a published safety evaluation — a low attack-success rate on a public attack corpus, scored by a guard classifier the same vendor ships — and no access to any training data. What evidence would tell you the corpus is inside the safety tuning or inside that classifier's training set?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Re-run the same harmful behaviours in fresh wording and compare. A large gap between verbatim prompts and paraphrases points at memorisation. Score the responses with a judge the vendor did not train, check whether the corpus predates the model release, and see whether the guard's own documentation lists it as training data.

open as a page

AdvBench and HarmBench are publicly released attack-prompt corpora used to score how often a model refuses. Why do red teams keep part of their own attack prompt set unpublished?

level: juniorimportance: should knowfreq 52%

basics

~20 s

Because anything published gets trained on. Vendors add public attack prompts to safety-tuning data and to guard classifiers, so models learn to refuse those exact prompts. The score rises without the model becoming harder to attack. An unpublished set stays a fair test, since nobody could have fitted to it.

open as a page

Your quarterly red-team report has cited an AdvBench attack-success rate for several quarters. You now learn the deployed guard classifier was trained on that same public corpus. Do you drop the number, and what goes in the report instead?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Do not silently drop it. Keep the public number as a regression floor, clearly labelled as a set the guard was trained on, and make a held-out attack set the headline. Report both, since the gap between them is your best estimate of how much the public score is inflated.

open as a page

Replacing a contaminated public attack corpus means building and maintaining a held-out attack set of your own. What does that cost an organisation over time, and what rules stop the replacement from becoming contaminated too?

level: principalimportance: should knowfreq 36%

basics

~20 s

It costs skilled authoring time, labelled ground truth, and a refresh cadence as the set ages. Keeping it clean is policy: never publish examples, only aggregates; send it only to endpoints under no-training terms; hold a sequestered slice nobody iterates against; and keep the tuning team from seeing it.

open as a page