Replacing a contaminated public attack corpus means building and maintaining a held-out attack set of your own. What does that cost an organisation over time, and what rules stop the replacement from becoming contaminated too?
answer
- asset with depreciation, not a file
- authoring, ground truth, adjudication
- aggregates in reports, never prompts
- no-training terms or treat as published
- measurement/tuning access boundary
basics
~20 sIt costs skilled authoring time, labelled ground truth, and a refresh cadence as the set ages. Keeping it clean is policy: never publish examples, only aggregates; send it only to endpoints under no-training terms; hold a sequestered slice nobody iterates against; and keep the tuning team from seeing it.
solid answer
~60 s**The cost is recurring, not one-off.** Authoring needs people who understand the attack families, because a set built by paraphrasing a public corpus inherits its surface forms. Every behaviour needs a ground-truth definition of what counts as a successful response, or you cannot score it. Then it decays — through vendor logs, through examples quoted in tickets, through the team gradually tuning against it — so budget a refresh every cycle. **The hygiene rules are organisational.** Report aggregates, never prompts. Send it only to endpoints covered by no-training, no-retention terms, and treat any endpoint without them as a publication event. Sequester a slice that is never used for iteration, only for the final read. Keep an access boundary between the people who measure and the people who tune, because the fastest way to contaminate your own set is to let it become a target for optimisation. The honest tradeoff: you gain a number you can defend and lose cross-vendor comparability, which is why you still run the public corpus alongside.
go deeper
Knows the replacement set has to stay unpublished to be worth anything.
Names the concrete costs — authoring, labels, refresh — and the obvious leaks through logs and reports.
Runs the rotation and sequestration discipline in practice and calibrates the scorer with human adjudication.
Funds it as a depreciating asset, sets the access boundary and no-training procurement terms, and decides when an honestly caveated public number beats a half-maintained private one.
Treat a held-out attack set as an asset with a depreciation schedule, not as a file that someone checks in once. ## Build cost - **Authoring.** The people writing it must know the attack families well enough to cover the *behaviour* space rather than the *phrasing* space. Rewriting a public corpus with synonym substitutions is the cheap trap: contamination attaches to phrasing patterns as much as to exact strings, so a derived set inherits the very surface cues you were trying to escape and quietly measures the same memorisation. Realistic figure: one to two engineer-weeks of skilled time for a few hundred usable items, and that number recurs every refresh. - **Ground truth.** Every item needs a written definition of what a successful response looks like, at the granularity your scorer can actually apply. Ambiguous items generate judge disagreement that later gets misread as model drift. - **Scoring.** You still need something to grade responses. If you grade with a hosted moderation classifier you have imported someone else's blind spots — including, plausibly, training on the same public safety corpora you left behind. Budget human adjudication of a sample every cycle to calibrate whatever automated scorer you use. Note what is *not* expensive: the runs. A few hundred prompts at a handful of samples each, plus one grader call per completion, is a few thousand short calls — single-digit to low-tens of dollars, under an hour of wall clock, and a scanner such as garak scales exactly linearly through its `--generations` flag. Anyone who budgets this work by API spend will conclude it is nearly free and will not fund the part that actually costs money. ## Depreciation The set loses value through every exposure: prompts sent to endpoints that retain inputs, examples pasted into tickets and slide decks, and — the largest by far — the internal team optimising against it. Every other leak is slow; a tuning team with access contaminates the set within one cycle, because a private eval that is also a target is just a public benchmark with a smaller audience. Plan a refresh fraction per cycle and a full rotation horizon. ## Hygiene rules worth writing down 1. Reports carry aggregate rates and behaviour-class names, never prompt text. 2. Any endpoint without contractual no-training and retention limits is treated as **publication**; prompts sent there move to a burned pool and are rotated out. 3. A sequestered slice is never used during iteration. It is opened only for the release read, and once a decision has been argued against it, it rotates. 4. An access boundary separates measurement from tuning. 5. Provenance is recorded per item: author, date, behaviour class, and whether it has ever left the boundary. ## Where the private number misleads A held-out rate feels authoritative because you built the set, and that feeling is the risk. Four specific misreadings: - **"It is private, therefore it is clean."** Only until someone tunes against it. Privacy is a property of the process, not of the file. - **Trend read across refreshes.** When you rotate items, this quarter's rate is measured on a partly different set. A movement can be composition change rather than model change; keep an anchor slice unchanged across at least two cycles so you can separate the two. - **Cross-vendor comparison.** A private rate is not comparable with anyone's published number and cannot be independently reproduced, which is a real cost when a regulator or customer asks for evidence. That is the argument for keeping a public corpus in the report as a labelled, comparable, contaminated line. - **Coverage mistaken for robustness.** A low rate on a set that covers four behaviour classes says nothing about the fifth. ## What you check each cycle Re-run a small anchor slice unchanged and confirm its rate is stable, so a moving headline can be attributed to new items rather than to the endpoint or the scorer. Adjudicate a human-labelled sample and record the automated scorer's agreement rate; a drop there is scorer drift, not model change. Audit exposure: how many items left the boundary this cycle, and were they rotated. And compare rates on reused versus fresh items — if reused items score systematically better, the set has begun to leak into whatever the tuning team is doing. ## When not to build one If the organisation cannot fund authoring, adjudication and rotation, a half-maintained held-out set that everyone quietly tunes against is **worse** than an honestly caveated public number, because it looks trustworthy and is not. That judgement — whether the recurring budget exists — is the actual decision here, and it is made once a year, not once.
- What is the single highest-value hygiene rule if you can only enforce one?The access boundary between measurement and tuning. Every other leak is slow; a team optimising directly against the set contaminates it in one cycle.
- How do you decide the refresh fraction per cycle?From exposure: how many items left the boundary, how many were argued over in a release decision, and how much the reported rate moved on reused versus fresh items. Rotate at least the exposed ones.
- A customer asks for reproducible evidence and you cannot share the set. What do you give them?Aggregates plus methodology — behaviour classes covered, scoring definition, adjudication rate — and the caveated public-corpus number for comparability, stating plainly which one backs the claim.
A held-out attack set behaves like a fleet vehicle rather than a filing cabinet: it depreciates every time it is driven, and the budget line that matters is the replacement schedule, not the purchase price.
saying these in an interview costs you the question
- Treats the held-out set as a one-time build with no refresh budget.
- Builds it by paraphrasing a public corpus, inheriting the same surface forms.
- Sends it to endpoints with no retention or training terms and calls it private.
- Gives the tuning team access, turning the eval into an optimisation target.
- Quotes prompt text in reports and slides.
- Drops the public corpus entirely, losing all comparability.