You built the attack half of your guardrail test corpus by downloading a public jailbreak-prompt dataset, and the guardrail's vendor lists that same dataset among its training sources. Why is the resulting catch rate uninformative, and what do you build instead?
answer
- train/test overlap flatters the number
- memorisation scored as detection
- held-out, authored or transformed
- public equals contaminated by default
- publishing burns the corpus
basics
~20 sThe classifier was fitted on those exact strings, so a high catch rate scores memorisation, not detection. It tells you nothing about a phrasing it has never seen. Build a held-out attack half you author or transform yourself, keep it unpublished, and record where every item came from so contamination can be checked later.
solid answer
~60 sA guardrail trained on a public corpus has effectively seen the answer key. Testing it on that corpus measures recall of memorised strings, and the number will be high whatever the model's real generalisation is. The failure is one-directional and always flattering: contamination inflates the score, so you will ship believing the layer is stronger than it is. What you build instead is a held-out attack half with stated provenance. Author items yourself against the technique families you care about, or take public items and transform them enough that the surface form is new — paraphrase, change domain, change language, change the wrapper — while the policy label stays the same. Keep the whole corpus out of public repos, issue trackers and prompts sent to third-party services, because a published test set becomes next quarter's training data. Then record, per item, where it came from, who labelled it and when. That record is what lets you answer "could the guard have seen this?" a year later, and what lets you retire items whose provenance has gone stale.
go deeper
Recognises that testing on data the classifier trained on is circular and inflates the result.
Explains that the bias is one-directional, and builds a held-out half by authoring or transforming items while keeping the label fixed.
Carries per-item provenance, uses the original-versus-paraphrase gap to quantify memorisation, and enforces hygiene that keeps the corpus unpublished.
Sets the organisation's rule for which numbers may be quoted to whom, budgets for periodic corpus refresh as contamination erodes it, and makes provenance a required field in any guardrail evaluation report.
**The mechanism.** A moderation or prompt-injection classifier — Llama Guard, ShieldGemma, Prompt Guard, WildGuard, a hosted omni-moderation endpoint — is a supervised model fitted on labelled examples. Fitting means its parameters were adjusted until it produced the right label on those specific strings. If your test items are drawn from that same collection, the quantity you measure is training accuracy wearing a test result's clothes. The critical property is that the bias is *one-directional*: overlap can only raise the measured catch rate, never lower it. A defect that made your guard look worse than it is would be discovered in a week; a defect that makes it look better than it is ships, and the first person to find out is an attacker with a phrasing nobody wrote down. **Memorisation versus generalisation, concretely.** What you want to know is whether the guard recognises the *shape* of a coercive instruction. What a contaminated corpus tells you is whether it recognises strings it has already been fitted to. The two come apart sharply: it is routine for a guard to block a well-known public item at 90%+ and miss a paraphrase of that same item that preserves the policy label. That gap is the size of the illusion, and it is measurable — pair each public original with a transformed twin, run both, and report the delta. A large delta means most of the flattering number was recall. **Why vendor silence is not reassurance.** Some guard vendors publish training sources; most publish a summary at best. Absence of a disclosure is not evidence of cleanliness — public jailbreak collections are widely scraped, mirrored and re-uploaded, and any item that has sat on the open web for a year should be assumed seen. Set the prior accordingly: public provenance is contaminated by default, and it takes a positive reason, not a missing denial, to treat it otherwise. The same applies to your own harvested traffic if your inference contract permits training on it. **Building the held-out half.** Three sources, in descending order of trust and ascending order of cost: - *Harvested* — items from your own traffic and your own engagements, scrubbed. Freshest, and drawn from the distribution you actually care about, but only usable at the volume your product generates. - *Transformed* — public items rewritten so the surface changes while the policy label does not: paraphrase, re-domain, re-frame in another document type. Cheap per item, but every transform needs a human to confirm the label did not drift, because a paraphrase that softens intent has quietly become a benign item. - *Authored* — written from scratch by your own team against the technique families in scope, never published. Most trustworthy, most expensive. **What it costs.** Authoring is the bill people underestimate: an experienced red-teamer produces perhaps ten to twenty usable, labelled attack items an hour, so a 200-item held-out attack half is several engineer-days before anyone runs anything. Transformation with human review lands somewhere around a third of that. And the asset depreciates — every guard release, every model swap, and every accidental disclosure erodes it, so budget a refresh cadence rather than a one-off build. There is also a friction cost that surprises teams: a held-out corpus cannot be pasted into a vendor support ticket, an external evaluation service, or a public issue without burning the items. Keep a separate, sacrificial repro set built from public material for exactly those conversations. **Where the number misleads.** "88% blocked" is not one claim; it is at least two. On a held-out authored set it is evidence about unseen phrasings. On a public set the vendor trained on it is evidence about that vendor's fitting procedure and nothing else. Two further traps: a *low* score on your held-out set is not automatically a guard failure, because hand-authored items can be harder or subtly off-policy compared with real traffic — check the items before condemning the guard. And a vendor's own published catch rate was measured on their held-out data and their distribution, so it transfers to your deployment only as a weak prior. **What to check.** Carry per-item provenance — source class, upstream identifier, author, labeller, date, and whether the item has ever left your control — and quote the provenance mix beside every rate. Where the public source is downloadable, do a literal-overlap check of your items against it. Prefer items authored after the guard build's training cut-off, and record that date. Then run the paired original-versus-transform test at least once per guard version, because the delta is the cheapest contamination estimate you will ever get.
- You paraphrase a public attack item and the guardrail now misses it, though it blocked the original. What have you learned?That the block on the original was substantially surface memorisation rather than generalisation. The size of the gap between original and paraphrase is a usable measure of how much of the reported catch rate is contamination.
- The vendor publishes no training sources at all. How do you set your prior?Assume anything long-published on the open web has been seen, and rely on authored, transformed or self-harvested items for any number you intend to act on.
- Does the same contamination worry apply to the benign half?Yes, in mirror image: benign items drawn from a public set the guard was tuned on will pass too easily and understate over-blocking, so the benign half needs stated provenance as well.
Testing a guard on the corpus it was trained on is like marking a student using the exact questions they revised from the answer sheet: the high mark is real but it measures recall of that sheet, and the error can only ever run in the flattering direction.
saying these in an interview costs you the question
- Treating a public benchmark score as evidence for a specific deployment.
- Assuming no disclosure means no contamination.
- Publishing the test corpus, or pasting it into an external service, then reusing it.
- Believing contamination could bias the score in either direction.
- Keeping no record of where items came from, so contamination cannot be audited.