skip to content

Your organisation wants to track one safety benchmark's attack-success rate as a quarterly figure across model upgrades. What has to be frozen for that series to mean anything, and what silently breaks it?

level: principalimportance: nice to knowfreq 30%

answer

  1. one thing varies: the model
  2. vendor and pin the item list
  3. judge drift needs re-scoring overlap
  4. public list contamination flatters
  5. private held-out set as control line

basics

~20 s

Freeze everything except the model: the pinned item list, the attack method and its budget, attempts per item and the aggregation rule, decoding settings, the target wrapper, and the exact judging procedure. Silent breakers are a refreshed item list, an updated judge, a changed default sampling setting, and the list leaking into training data.

solid answer

~60 s

A trend line is only readable if exactly one thing varied. So pin the instrument as an artefact under version control: the item list at a specific commit, the attack configuration, attempts per item, the aggregation rule, decoding parameters, the target wrapper, and the ruling procedure including the judge's own weights or prompt. The silent breakers are the ones that do not raise an error: - **Item refresh** — the suite adds or edits behaviours and your rate jumps for population reasons. - **Judge drift** — the ruling model or classifier is updated, or a hosted judge changes underneath you, so old and new quarters are scored differently. - **Harness defaults** — a tool upgrade changes a default temperature, max tokens or retry count. - **Contamination** — a public list becomes widely mirrored and later models have effectively seen it, so the rate improves without any defence changing. Mitigations: re-score an archived quarter with the current judge whenever the judge changes, and keep a held-back private item set alongside the public one so contamination shows up as a divergence between the two lines.

go deeper

for a junior

Should say the test must be kept the same each quarter so only the model changes.

for a middle

Should list the concrete things to pin — items, attack, attempts, decoding, judge — and name judge changes as a breaker.

for a senior

Should add re-scoring archived runs on a judge change, insist on overlap windows for re-baselines, and record the configuration tuple with every point.

for a principal

Should treat the instrument as an owned versioned artefact, run a private held-out control line for contamination, and bound what the series is allowed to claim in reporting.

A trend line is readable only if exactly one thing varied between points. In a quarterly ASR series the one thing is meant to be the model, and everything else in the instrument has to be pinned as hard as a dependency lockfile. ### Treat the instrument as a versioned artefact Vendor the behaviour list into your own repository at a fixed revision rather than pulling it from upstream at run time. Pin the harness version, the attack configuration and its budget, attempts per item and the aggregation rule, decoding settings (temperature, top-p, max tokens, seed where honoured), the target wrapper, and the ruling procedure including the judge's own weights or its exact prompt. Then record that whole tuple *with every data point*: ``` 2026-Q3 list=jbb@a1b4c9 n=25 agg=any-of-n attack=tmpl-v3 budget=1 turn target=chat-endpoint+sysprompt-v7 temp=1.0 judge=harm-clf@2026-02 ASR=11.4% ``` A point without its tuple cannot be re-checked, and in a series that runs for years you will need to re-check. ### The breakers that raise no error - **Item refresh.** The suite adds, edits or retires behaviours; your rate steps for population reasons and nothing logs it. - **Judge drift.** The classifier is retrained or the judge model behind a hosted API is updated underneath you, so quarters are scored by different rulings while the chart pretends they are one series. - **Harness defaults.** A tool upgrade silently changes a default temperature, max-token cap or retry count — retries in particular quietly change the effective attempts per item. - **Contamination.** A public list becomes widely mirrored and later models have effectively seen it. ### Contamination is the hardest one, because it flatters The failure mode is that the rate falls and everyone congratulates the safety work. Defend with a private held-out set of comparable items, written and maintained by your own team, run every quarter beside the public list. Divergence — the public line improving while the private line is flat — is the contamination signal. Keep the private items out of any external system, rotate a fraction of them periodically, and treat a leak of the private set as an incident, because it destroys the only control you have. ### What a quarter costs, and what that implies Each data point is a full sweep: `behaviours x attempts` generations plus a judging call per transcript, doubled if you run a private control list beside the public one — for a few hundred behaviours at a couple of dozen samples, tens of thousands of calls, hours of wall clock and a bill in the tens to low hundreds of dollars per model, per quarter. Small enough to run, large enough that nobody re-runs the whole history on a whim. That cost asymmetry drives one concrete design decision: **archive every transcript**. Re-scoring stored transcripts under a new judge costs only the judging half and needs no model access at all, so a judge change is recoverable; re-running old models is often impossible once a provider retires a version. Storage for a few years of transcripts is trivial next to the inference you already paid for. ### Re-baselining discipline When something must change — you adopt a refreshed list, or the judge is replaced — do not splice. Run the new instrument over one or two archived quarters to create an overlap window, publish both series through it, and annotate the changeover on the chart itself. Without overlap a step in the line is unattributable and the history becomes decoration. ### Where the number misleads The series can support a regression alarm: this quarter's model is materially worse on a fixed instrument than last quarter's, investigate. It cannot support an absolute safety claim, because the denominator is one curated list under one attack at one budget, and the figure would move under any other instrument. Say that in the chart caption, not a footnote — a single tracked percentage travels into slide decks stripped of its tuple faster than anything else a red team produces, and a quarter later someone will compare it to a vendor's published number. Finally, name an owner for the instrument, separate from whoever consumes the trend. That owner approves re-baselines, maintains the private control set, and refuses out-of-band changes during a reporting window — which is the only mechanism that actually stops a well-meaning upgrade from silently breaking a year of history.

  • The suite you track publishes a substantially revised behaviour list. How do you adopt it?
    Run the new list over at least one archived quarter to create an overlap window, publish both lines during it, annotate the changeover, and only then retire the old series.
  • How do you detect that a public behaviour list has become contaminated?
    Maintain a private held-out set of comparable items and run it every quarter. When the public line improves and the private line does not, contamination is the leading explanation.
  • What claim should the tracked series never be used to support?
    An absolute safety claim. It is a regression alarm on one fixed instrument; the number would move under a different list, attack, budget or judge.

A public benchmark list leaking into training data is the practice exam being printed in the textbook: scores rise every year and nobody in the room got smarter. The private held-out set is the unpublished paper you keep in the drawer to tell the difference.

saying these in an interview costs you the question

  • Pulling the item list live from upstream each run
  • Swapping the judging procedure mid-series without re-scoring archived quarters
  • Reading a falling public-list rate as proof the defences improved
  • Publishing the tracked percentage without the configuration tuple
  • Splicing a re-baselined instrument into the old line with no overlap window

context