skip to content

Your quarterly red-team report has cited an AdvBench attack-success rate for several quarters. You now learn the deployed guard classifier was trained on that same public corpus. Do you drop the number, and what goes in the report instead?

level: seniorimportance: should knowfreq 44%

answer

  1. demote, do not delete
  2. public rate = regression floor
  3. held-out set becomes headline
  4. publish the gap as the estimate
  5. document the series break

basics

~20 s

Do not silently drop it. Keep the public number as a regression floor, clearly labelled as a set the guard was trained on, and make a held-out attack set the headline. Report both, since the gap between them is your best estimate of how much the public score is inflated.

solid answer

~50 s

Dropping it costs you the one thing it was good for: continuity. Several quarters of the same measurement is a regression detector — if that rate ever rises, something broke badly, because those are the prompts the system is most trained to handle. Keep it as a floor. What changes is its rank in the report. It stops being the headline and becomes a labelled line item: 'public corpus, in the guard's training data, retained for trend only'. The headline moves to a held-out set you authored, covering the same behaviour classes. Then report the gap explicitly. The difference between the public rate and the held-out rate is the most defensible contamination estimate you can produce, and it is more informative to a decision-maker than either number alone. The risk to manage is a reader who has quoted the old headline in a compliance artefact. Restating the series with the caveat attached, rather than deleting it, is what keeps the report trusted.

go deeper

for a junior

Recognises the number is inflated and should be flagged rather than presented as a robustness claim.

for a middle

Keeps it as a labelled regression metric and knows a held-out set has to take the headline.

for a senior

Structures the report — headline, retained line item, explicit gap, migration note — and manages readers who already quoted the old number.

for a principal

Sets the org policy for which numbers may back an external safety claim and what provenance disclosure the org asks of vendors.

This is a reporting-integrity decision more than a technical one, and both instinctive answers are wrong in ways that cost the team something real. ## Extreme one: keep citing it unchanged You now know the figure is close to a tautology. The deployed guard was trained on those prompts and it catches those prompts; the attack-success rate on them measures the guard's recall of a list it studied. Continuing to headline it means the report asserts robustness it cannot support, and reports get quoted. Someone downstream pastes the figure into a risk register or a customer questionnaire, where it will outlive the caveat and become the org's position. ## Extreme two: delete it A quarterly series that silently loses a metric is worse than one that keeps a caveated metric. Readers who quoted it get no signal that it was withdrawn, the trend break is unexplained, and you throw away a genuine alarm: a **rise** in the attack-success rate on the set the system is most heavily trained against is a loud signal that something regressed badly — a guard rollback, a tuning-data mistake, a routing change to an unguarded model. That alarm is worth keeping wired up precisely because the number is saturated. ## The structure that works - **Headline:** the held-out attack set, authored in-house, matched to the same behaviour classes. This is the number the report defends. - **Retained line item:** the public-corpus rate, with a one-line provenance note naming it as training data for the shipped guard, and its role stated explicitly as *regression floor, not evidence of robustness*. - **Derived line:** the gap between the two, presented as the contamination estimate. This is the sentence a decision-maker actually needs — on prompts the system has seen, X; on matched behaviours it has not, Y. - **Migration note** in the quarter the change lands, explaining why the headline moved, so a series break is documented rather than mysterious. ## What the change costs The held-out set must exist before you can promote it, and it is not free: skilled authoring of matched behaviours (order of one to two engineer-weeks per few hundred items), a written ground truth per behaviour, and a human-adjudicated sample each cycle to calibrate the scorer. Running both arms is trivially cheap by comparison — a few hundred prompts at a handful of samples each, plus a grader call per completion, is single-digit to low-tens of dollars and under an hour on a metered endpoint. Then there is a recurring hygiene cost that has no line in most budgets: the held-out prompts may never be published, may only go to endpoints under no-training and retention-limited terms, and must stay away from the team tuning the system. Every quarter of use erodes the set, so budget a refresh fraction, not a one-off build. If you have not paid that cost yet, say so. The honest interim report carries the public rate with its provenance caveat, a plain statement that no uncontaminated measurement exists yet, and the held-out set as a dated deliverable. ## Where these numbers mislead The retained public rate misleads in exactly one direction and one magnitude: **downwards, by an unknown amount**. It is a floor on measured refusal, not a bound on attacker success. Two more traps sit nearby. Re-thresholding the same guard — tightening its decision boundary and reporting the stricter figure as though it were cleaner — removes nothing; the overlap between the guard's training data and the prompts is untouched, you have only moved the operating point of a contaminated instrument. And promoting a "held-out" set that was produced by paraphrasing the public one reproduces the original contamination under a private filename, which is worse than the public number because it looks trustworthy. ## What to check before publishing Confirm the held-out set covers the same behaviour classes as the public one, or the gap you report is measuring coverage rather than contamination. Confirm both arms were run in the same window against the same endpoint version. Confirm the two arms were graded by the same scorer, or the gap absorbs judge disagreement as well. And check who has already quoted the old headline, so the migration note reaches them. ## What not to do Do not accuse the vendor. Training a guard on published harmful-prompt corpora is the expected engineering choice, not misconduct; it merely destroys those prompts as a measurement. Framed as measurement provenance, the conversation with the vendor stays productive and may even get you the disclosure you actually want.

  • A stakeholder asks why you keep reporting a number you have just called uninformative. What is the one-line answer?
    It is a floor, not a claim: those are the prompts the system is most trained to refuse, so if that rate ever rises something has broken badly.
  • You do not yet have a held-out set. What goes in this quarter's report?
    The public rate with the provenance caveat, an explicit statement that no uncontaminated measurement exists yet, and the held-out set as a committed deliverable with a date.

saying these in an interview costs you the question

  • Silently removes the metric, breaking a quarterly series readers have quoted.
  • Keeps it as the headline with no provenance caveat.
  • Frames vendor guard training on public safety corpora as misconduct.
  • Promotes a held-out set that does not exist yet, or is just a paraphrase of the public one.
  • Reports only one of the two numbers, so the reader cannot see the gap.

context