You are handed only a published safety evaluation — a low attack-success rate on a public attack corpus, scored by a guard classifier the same vendor ships — and no access to any training data. What evidence would tell you the corpus is inside the safety tuning or inside that classifier's training set?
answer
- interleave verbatim and paraphrase
- re-score old transcripts, second judge
- refusal shape: flat vs severity-tracking
- corpus date before release date
- step change at one release
basics
~20 sRe-run the same harmful behaviours in fresh wording and compare. A large gap between verbatim prompts and paraphrases points at memorisation. Score the responses with a judge the vendor did not train, check whether the corpus predates the model release, and see whether the guard's own documentation lists it as training data.
solid answer
~50 sYou cannot inspect the training mix, so you infer it from behaviour, with two independent probes. **Probe the model.** Hold the behaviour class fixed and change the surface form. Interleave verbatim and reworded prompts in one session so endpoint drift hits both. A wide, consistent gap is memorisation. Also look at the *shape* of refusals: identical boilerplate fired instantly across an entire published list reads differently from refusals that vary with how harmful the request actually is. **Probe the judge.** Re-score the same transcripts with a scorer the vendor did not train — a second classifier, or human adjudication on a sample. If the vendor's guard calls things refusals that an independent scorer calls partial compliance, the measurement instrument is contaminated independently of the model. **Paper evidence.** Corpus publication date versus model release, whether the guard's own documentation names public safety corpora as training sources, and whether the reported score jumped discontinuously at one release rather than improving gradually. Report the deltas, not a verdict.
go deeper
Suggests rewording the prompts and seeing whether the score changes.
Designs the rewording test properly — matched behaviours, interleaved runs — and knows the corpus date matters.
Separates target contamination from judge contamination, re-scores frozen transcripts independently, reads refusal shape, and writes up deltas rather than accusations.
Decides what evidence bar the org needs before it will refuse to accept a vendor's published safety number, and what it will ask vendors to disclose instead.
You have a PDF and no training data. That makes this a detection problem under no visibility, and the discipline that separates a useful answer from a hunch is insisting on *which component* is contaminated, because the remedies differ. A contaminated target needs different prompts. A contaminated judge needs a different scorer. A report that says only "this looks contaminated" tells the reader neither. ## Signal 1 — the verbatim/paraphrase delta (tests the target) Take the behaviours the published corpus covers, re-author each one in phrasing that shares no distinctive n-grams with the original, and run both arms **interleaved** against the same endpoint in the same window. Interleaving is not fussiness: hosted endpoints change without announcement, and a sequential A-then-B design confounds contamination with drift, which is a different leaf's problem entirely. A wide, consistent gap on matched behaviours is the strongest available evidence that the published rate reflects surface recall rather than a refusal that generalises. The construction rule that people get wrong: synonym substitution is not re-authoring. Contamination attaches to phrasing patterns, not only to exact strings, so a set derived by swapping words keeps the cues under test and will show a falsely small delta. ## Signal 2 — judge disagreement (tests the scorer) Freeze the transcripts from the *original* run and re-score them with an independently trained judge, plus a human-labelled sample as tiebreak. Because the prompts and responses are unchanged, any movement in the rate is attributable to the instrument alone. Guard contamination shows up as systematic, one-directional disagreement: the vendor's classifier reads hedged or partially compliant responses as refusals on precisely the prompts it was trained around. This is the signal people forget, and it is why a score can survive a rewording test and still be wrong. ## Signal 3 — refusal shape Memorised refusals tend to be fast, template-identical and insensitive to what was actually asked. Generalised refusals track severity: more hedging on the genuinely dangerous items, less on the mild ones. A flat, uniform refusal profile across an entire public list is a cheap, run-free tell. ## Signal 4 — timeline and documentation A corpus published comfortably before the model release had time to be absorbed; one released afterwards did not. And vendor documentation for a shipped guard sometimes names public safety datasets among its training sources. When it does, that is the cleanest evidence obtainable with no experiment at all, and it is the first thing to look for because it is free. ## Signal 5 — discontinuity across releases Robustness that improves gradually across many attack families over successive releases looks like real work. A rate that steps to near-zero on one public list at one release while other families barely move looks like fitting to that list. ## What the investigation costs Two arms of a few hundred behaviours at a handful of samples each is a few thousand short completions plus a grader call apiece — single-digit to low-tens of dollars on a metered endpoint, under an hour of wall clock, and a scanner such as garak scales it linearly through `--generations`. Re-scoring frozen transcripts costs one more grader pass over text you already have, which is the cheapest evidence in the whole exercise. The expensive line items are human: authoring the paraphrase arm without inheriting surface forms, and adjudicating a labelled sample so the independent judge is itself calibrated. Budget days of skilled time, not dollars, and do not let the low API cost tempt you into skipping the adjudication — an uncalibrated second judge produces a disagreement number nobody can interpret. ## Where these numbers themselves mislead Three traps, in the order people fall into them. First, a **near-zero delta** is read as proof the model generalises; it is equally consistent with a paraphrase arm that preserved the cues, so verify the rewrite before believing the result. Second, **judge disagreement is read as model behaviour** — it cannot be, when the transcripts are frozen. Third, and most damaging, a delta is read as an accusation. You cannot observe a training set, and "they trained on the test set" is a claim your evidence does not support. ## How to write it up Never assert misconduct. Write: on prompts drawn verbatim from a public corpus the attack-success rate was A; on matched behaviours re-authored by us, run interleaved in the same window, it was B; independent re-scoring of the original transcripts moved A to A-prime; the guard's documentation does or does not name public safety corpora among its training sources; the corpus predates the release by N months. Let A, B and A-prime carry the argument, and state plainly which of them you would put behind a claim.
- Your rewording delta is near zero but the independent judge disagrees with the vendor's guard on a third of transcripts. What do you conclude?The target's refusal may generalise fine; the measurement is the problem. Re-report the whole evaluation under the independent scorer before making any robustness claim.
- Why interleave the verbatim and reworded prompts rather than running them back to back?Hosted endpoints change without announcement. Interleaving inside one window means any drift hits both arms equally, so the gap you measure is attributable to the prompts.
- What is the strongest evidence you can get without running anything?Vendor documentation for the shipped guard naming public safety corpora among its training sources, plus a corpus publication date comfortably before the model release.
saying these in an interview costs you the question
- Concludes 'they trained on the test set' from the score alone, with no experiment.
- Tests only the model and never re-scores with an independent judge.
- Runs verbatim and reworded sets in separate windows, confounding contamination with endpoint drift.
- Builds the paraphrase set by swapping synonyms, preserving the surface cues under test.
- Reports a verdict instead of the measured deltas.