You are asked to set the house standard for attempts per behaviour in red-team results that gate a model release. What do you fix, what do you leave to each team, and what does over-fixing cost you?
answer
- comparable core + free exploration layer
- pin list, minimum n, stopping rule, hit rule, target config
- caption enforced by the tool
- Goodhart on a frozen list
- gate improving while exploration still lands = stale core
basics
~20 sFix what makes numbers comparable across releases: the behaviour list, a minimum attempts per behaviour, the stopping rule, the hit rule, and a caption that states all of them. Leave extra deep runs and new attacks free, reported separately. Over-fixing costs discovery: a frozen budget stops anyone probing where the model actually looks weak.
solid answer
~50 sSplit the standard into a **comparable core** and a **free exploration layer**. The core is a release gate: one frozen behaviour list, a declared minimum attempts per behaviour, one hit rule, one target configuration, and a mandatory caption. Its whole job is that this release's number can be set beside last release's number, so every parameter that moves the rate must be pinned. Change the core deliberately and re-baseline the previous release when you do. The exploration layer is everything else — new attacks, deeper budgets on suspicious categories, novel behaviours — reported as findings rather than as a rate, and never merged into the gate figure. Over-fixing has two costs. It converts red teaming into a compliance run, where nobody spends a query outside the frozen list; and it makes the list a target that gets optimised against, so the gate number improves while real exposure does not. Budget explicitly for the exploration layer, and promote what it finds into the core on a schedule.
go deeper
Understands that everyone should use the same list and the same number of attempts so results can be compared.
Enumerates the parameters that must be pinned and insists the caption carries them.
Sizes the minimum attempt budget from the saturation curve, versions the hit rule, and re-scores history when it changes.
Runs a comparable core plus a funded exploration layer, anticipates Goodhart on a frozen list, and monitors gate-versus-exploration divergence as the staleness signal.
**Why a house standard exists at all.** Its purpose is not statistical rigour; it is that the gate metric means the same thing twice. Any-hit aggregation makes the reported rate a function of the attempt budget, the hit rule, the attack and the target's surrounding configuration. Two teams reporting on the same behaviour list, each choosing those freely, routinely produce numbers that differ by more than any release-to-release change you would actually act on. Without a pinned core, every release review turns into an argument about methodology at the worst possible moment. **Split the standard in two.** The *comparable core* is the release gate. The *exploration layer* is everything else, funded separately and reported as findings rather than as a rate. Merging the second into the first destroys the first. **What the core pins.** - **The behaviour list and its version.** A list is edited, subsetted and re-released; "HarmBench" without a version is not a fixed denominator. - **A minimum attempts per behaviour**, sized from the saturation curve rather than chosen as a round number: take the n past which one more doubling of the budget moves the rate by less than the smallest difference that would block a release. Below that you are measuring your budget; above it you are buying points you will never act on. - **The stopping rule** — whether a behaviour halts at its first hit. It does not change the verdict, but it determines what per-attempt statistics can ever be recovered from the run. - **The hit rule, versioned.** This is the one teams forget to pin, and it is the one that silently rewrites history: swapping a judge model moves every number in the series with nothing else changed. - **The target configuration measured** — bare endpoint, or the shipped stack with its input and output filters. Both are legitimate; they are different systems, so pick one and say which. - **The caption format**, enforced in the reporting tool rather than in a policy document. A convention that lives only in a wiki page is not enforced. **What stays free.** Attack strategies, extra attempts beyond the minimum, novel behaviours, new tools. This is where actual findings come from. Require them to be written up as findings with reproduction steps, never as a competing rate that a reader can mistake for the gate figure. **What the core costs, and why that governs its design.** One gate run is list size times minimum n target generations, plus a judge call per attempt if the hit rule is a hosted model, plus wall clock against a rate-limited endpoint, plus the human hours to triage the candidate hits. If a release cadence is weekly and a gate run takes three days and a five-figure spend, teams will quietly run it at a reduced budget and report the number as though it were the standard one — a standard nobody can afford is worse than a looser one everyone actually runs, because it produces numbers that look comparable and are not. **Where over-fixing bites.** *Goodhart:* a frozen list with a threshold attached becomes the training and prompt-engineering target, so the gate rate improves while real exposure does not move. *Discovery collapse:* if the gate consumes the whole query budget, nobody spends a query outside the frozen list, and red teaming degrades into a compliance run. *Staleness:* attack technique moves faster than a standards revision cycle, so a two-year-old core measures an obsolete threat model with impressive precision. Counter all three deliberately — date the core, keep a rotating held-out portion of behaviours that never enters any training or tuning loop, and fund exploration as its own line item rather than as slack in the gate run. **Where the gate number misleads, even when the standard is followed.** It is a comparability instrument, not an exposure estimate. It says how this release compares to the last one on a fixed list under a fixed budget; it does not say what fraction of real user traffic produces harm, because the list was assembled adversarially and the budget was chosen for affordability. Anyone quoting the gate figure as organisational risk is misusing it, and that misuse is worth pre-empting in the caption. **What you would monitor.** The gate rate against exploration findings: if the gate improves release over release while the exploration layer keeps landing new hits on the same system, the core has gone stale and the improvement is Goodhart, not hardening. Also the measured cost and wall clock of a gate run per release, because a standard drifting past what teams can afford will be silently degraded rather than openly renegotiated. And whenever the hit rule changes, re-score the retained transcripts under the new rule so the series stays continuous, and mark the discontinuity in the report.
- How do you pick the minimum attempts per behaviour rather than guessing a round number?From the saturation curve on a representative attack: choose the n past which a doubling of the budget moves the rate by less than the smallest difference that would block a release.
- You must change the hit rule. What do you owe the historical numbers?A re-scoring of the retained transcripts under the new rule, so the series is continuous, plus a marked discontinuity in the report where the rule changed.
- What single signal tells you the frozen core has gone stale?The gate rate improving release over release while the free exploration layer keeps producing new hits on the same system.
saying these in an interview costs you the question
- Fixing everything, leaving no funded room for new attacks or deeper budgets.
- Standardising the attempt count but leaving the hit rule or target configuration free.
- Never re-baselining history after changing the core.
- Treating the gate figure as the organisation's exposure rather than as a comparability instrument.
- A standard so expensive that teams quietly run it at a reduced budget.