Leadership proposes making a score on a public fixed-attack jailbreak leaderboard, such as a JailbreakBench-style suite, the mandatory release gate for every model or prompt update your product ships. What is the case on both sides, and what would you put in the gate instead?
answer
- tripwire versus gate
- Goodhart once it blocks a launch
- suite attacks a bare model, not your seams
- regenerate attacks each cycle
- publish revision, judge and caveat
basics
~20 sFor: it is cheap, repeatable and comparable across releases, so a regression is visible. Against: a public frozen suite can be optimised against and says nothing about your own application's attacks. Use it as a non-regression tripwire, and gate release on fresh adversarial runs against your deployed stack.
solid answer
~60 s**The case for.** A frozen public suite is the only thing in the safety toolbox that is genuinely comparable release over release and across teams. It is cheap, needs no expert to run, and gives a hard, auditable artefact. As a *tripwire* — this release must not score worse than the last — it is excellent. **The case against.** Three failures make it a bad *gate*. It is a fixed, public target, so once it gates shipping, people optimise against it and the score decouples from real safety — Goodhart in its purest form. It covers a general harm taxonomy, not your product's tools, system prompt or retrieved content. And it ages, so the bar quietly drops over time while the number stays flat. **What I would gate on instead.** A layered gate: the frozen suite as a non-regression check; a freshly generated attack run against the deployed stack, refreshed each cycle so it cannot be memorised; application-surface tests you own; and a human review sign-off for the highest-consequence flows.
go deeper
Says the public score is useful but does not cover the product's own attack surface, so it should not be the only check.
Separates cheap regression signal from safety evidence and proposes adding application-specific tests.
Names the Goodhart dynamic and the ageing bar, and designs the layered gate with a regenerated attack arm.
Adds the organisational safeguards — corpus ownership separated from defence tuning, scheduled re-baselining, and a mandatory caveat on any externally reported number.
**The argument turns entirely on separating two roles: tripwire and gate.** A *tripwire* is a cheap, fixed check that fires when something regresses. A *gate* is the decision procedure that certifies a release fit to ship. A frozen public suite is a first-class tripwire and a poor gate, and most of the disagreement in a room like this comes from people using one word for both jobs. **The honest case for the proposal** A frozen public suite is close to the only artefact in the safety toolbox that is genuinely comparable release over release and across teams. It is cheap — a few hundred to a few thousand target calls, minutes to a couple of hours of wall clock, single-digit to low-tens of dollars per release, and no expert needed once the adapter exists. It produces a hard, auditable number with a date and a revision on it, which is exactly what a compliance function or a board asks for. As a rule of the form *this release must not score worse than the last*, it is excellent and cheap enough to run on every build. **The four reasons it fails as the gate** *Goodhart under organisational pressure.* The moment a fixed public score can block a launch, engineering effort flows towards moving that score, and the cheapest way to move it is to pattern-match its published strings. The metric stops measuring what it measured, and it does so without any individual acting in bad faith. *Wrong scope.* A general jailbreak suite attacks a bare model with single-turn prompts. Your actual risk lives at the seams: the system prompt, the tool-calling surface, retrieved documents, multi-turn state, and the harms specific to your domain. None of that is in the suite, so passing it says nothing about the surfaces that carry your exposure. *Silent drift.* The suite ages against a patched model. "We held the same score for four quarters" can mean the bar fell while the product stood still, and the number reads identically either way. *False precision.* A single percentage invites a numeric threshold, and a threshold invites the release meeting to argue about the third digit rather than about the risk. It also creates a standing incentive to negotiate the threshold rather than investigate a regression. **The layered gate I would write instead** 1. **Non-regression tripwire.** Frozen suite, unmodified shipped judge, pinned revision, denominator recorded. Any worsening beyond the measured rerun spread blocks the release and requires a written explanation — not a threshold renegotiation. Cost: minutes and pocket change per build. 2. **Fresh adversarial run.** Automatically generated attacks against the same harm behaviours, regenerated every cycle against the deployed stack so there is nothing stable to memorise or fit. This is the arm that actually has to pass. Cost: one to two orders of magnitude above the tripwire in calls and wall clock, and it needs rate-limit headroom and internal authorisation for sustained adversarial traffic — budget it explicitly or it will quietly not happen. 3. **Application-surface suite you own.** Attacks against your system prompt, tools, retrieved content and multi-turn flows, versioned in your own repository and maintained like product code. Cost: mostly engineer time, and it is the line item that gets cut first and hurts most. 4. **Human sign-off on the highest-consequence flows,** with a written statement of what was tested and, equally important, what was not. 5. **A dated caveat attached to any frozen number that leaves the building:** which suite revision, which judge, which denominator, and the sentence that it is a regression signal rather than a coverage claim. **Where the numbers mislead, stated for leadership** A passing tripwire is evidence that nothing got worse on a narrow, ageing, public set. It is not evidence that the release is safe, and the difference is not pedantic: those two statements license completely different external claims. A flat score across releases is ambiguous between "we held" and "the bar fell". And a fresh-attack arm that is not actually regenerated each cycle degrades into a second frozen suite within a couple of quarters, at which point you are paying the expensive arm's bill for the cheap arm's evidence. **Organisational safeguards** Keep the team that tunes defences separate from the team that owns and refreshes the attack corpus. Treat any request to "improve the leaderboard number" as a change needing the same review as a product change. Re-baseline the frozen tripwire on a schedule, so its ageing is a decision somebody made and dated rather than a drift nobody noticed.
- Why is a fixed pass threshold on the public score worse than a non-regression rule?A threshold becomes a target to be engineered towards and hides drift; a non-regression rule asks only that this release be no worse than the last and forces an explanation when it is.
- How do you stop the fresh-attack arm from becoming another fixed target?Regenerate it each cycle against the deployed stack and keep its corpus owned by a team separate from the one tuning the defences.
A smoke alarm is worth having on every floor, and it is still not a fire-safety certificate. Making the alarm the sign-off criterion is what gets you a building that is very good at not setting off that particular alarm.
saying these in an interview costs you the question
- Accepting a single public score as the whole release criterion.
- Negotiating the pass threshold instead of investigating a regression.
- No application-surface testing because the public suite passed.
- Letting the same team both tune the defence and own the attack corpus.