skip to content

Fixed-Attack Leaderboards

A suite like JailbreakBench freezes the attacks, the judge and the report shape so two defences compare at all, and the same freeze goes stale as models are patched. Interviewers probe that trade.

on this pageshow

explore

questions

5

A fixed-attack jailbreak benchmark such as JailbreakBench freezes several things at once so that results are comparable. What does it freeze, and why is that freeze what makes two different defences comparable on the same model?

level: juniorimportance: must knowfreq 70%

answer

  1. attacks frozen, judge frozen, report frozen
  2. delta isolates the defence
  3. public prompts get patched first
  4. score answers a narrow question
  5. which revision, which judge, which denominator

basics

~20 s

It freezes the attack prompts, the judge that decides whether a response counts as a jailbreak, and the report format. With all three fixed, two defences see identical inputs scored by identical rules, so the difference in reported success rate reflects the defence and not a different prompt set or a different grader.

solid answer

~50 s

A fixed-attack suite of this class is three frozen artefacts, not one: **the attack set** (a published list of prompts or attack instances, each tied to a behaviour it is trying to elicit), **the decision rule** (the harm judge that labels each response jailbroken or not), and **the report shape** (which numbers are published, over which denominator). The freeze is the whole point. If you change the prompts, you have measured a different threat. If you change the judge, you have changed what "success" means. If you change the denominator, the percentage stops meaning the same thing. Only with all three pinned does the delta between two defences isolate the defence. The cost is equally direct: whatever the freeze excludes is invisible. The number answers "how did this defence do against these attacks, under this judge" and nothing wider.

go deeper

for a junior

Names the three frozen pieces — attacks, judge, report format — and says the freeze is what makes two defences comparable.

for a middle

Explains why swapping any one piece breaks comparability, and states what the number does not cover.

for a senior

Adds the staleness and public-exposure argument, and the five things to check before trusting someone else's posted score.

for a principal

Frames it as instrument versus assessment, and sets policy for where such a suite may appear in a release decision at all.

## A fixed-attack suite is an instrument, not an assessment An **assessment** asks "is this system safe?" and has to consider everything. An **instrument** asks one narrow question and answers it the same way every single time it is used. A fixed-attack jailbreak benchmark — JailbreakBench is the canonical example of the class — is deliberately built as the second thing, and every strength and every weakness it has follows from that one design choice. ## The three artefacts it freezes - **The attack set.** A published, versioned list of attack artefacts, each tied to a **behaviour** — a named harmful capability the attempt is trying to elicit, phrased as a target ("produce instructions for X"). Behaviours are the rows of the experiment; the attack artefacts are the attempts made against each row. A suite of this class typically carries on the order of a hundred behaviours, split across a harm taxonomy, plus a matching set of benign behaviours used to check that a defence is not simply refusing everything. Freezing this means a difference between two runs cannot have come from a different attacker. - **The harm judge.** A fixed decision procedure — a trained classifier, a scripted string rule, or a model prompted in a published, versioned way — that reads the pair (behaviour, model response) and emits one bit: jailbroken or not. This is the component that turns text into a number, and it is emphatically part of the measurement rather than an implementation detail. Two labs scoring the *same* stored responses with different judges will publish different rates. - **The report shape.** Which rate is published, over which denominator, and what is published beside it. Attack-success rate per *attempt* (every attempt is a row), per *behaviour* (a behaviour counts once), and per *behaviour with at least one success out of k attempts* are three different numbers computed from one identical run, and the third rises monotonically as you raise k. Freezing the report shape is what lets two leaderboard entries be lined up on a page at all. ## Why the freeze is what makes two defences comparable Comparison is subtraction, and subtraction is only meaningful when everything except the term you care about is held constant. - Change the prompts and you have measured a different threat. - Change the judge and you have changed the definition of the word "success". - Change the denominator and the percentage has changed units. With all three pinned, the delta between defence A and defence B has exactly one place left to come from: the defence. ## What a run costs Target calls are roughly `behaviours x attack artefacts per behaviour x generations per attempt`. A hundred behaviours with a handful of attack artefacts and one generation each is several hundred to a few thousand calls — on a commodity hosted chat endpoint, single-digit to low-tens of dollars and tens of minutes of wall clock at modest concurrency. A model-backed judge roughly doubles the call count and adds its own bill. None of that is the real cost. The real cost is **engineer time**: writing the target adapter so the suite talks to your *deployed* stack rather than a bare model, pinning the suite revision, and hand-triaging borderline labels. Budget the adapter in days and the run itself in minutes; teams routinely size this backwards and are surprised that the cheap part is the compute. ## Where the number misleads The rate answers "how did this defence do against *these* frozen attempts, under *this* judge, over *this* denominator" — and nothing wider. Four specific misreadings recur. 1. First, quoting it as "our jailbreak rate", which is a claim about all attacks and is not supported. 2. Second, comparing a number computed per-behaviour-with-at-least-one-success at k=10 against someone else's per-attempt number, which is a units error dressed as a comparison. 3. Third, ignoring that the prompts are public, so they are the first strings a vendor blocks by rule or by safety-tuning; the score drifts optimistic over time while looking identical. 4. Fourth, quoting a bare-model figure for a product that adds a system prompt, tools and retrieved content — surfaces the suite never touches. ## What to check before believing a posted number - Which suite and which revision; - whether the shipped judge was used unmodified; - the exact denominator and the attempts per behaviour; - the target configuration, including decoding settings and whether it was the bare model or a deployed stack; - and whether the defence under test was developed while its authors could see those prompts. Absent those five, a posted percentage is a number, not a measurement you can compare yours to.

  • Why is the report shape, not just the prompts, part of what has to be frozen?
    Because the same raw results can be summarised over different denominators — per attempt, per behaviour, or per behaviour with at least one success — and those percentages are not interchangeable.
  • You add ten of your own prompts to the frozen set. Can you still compare against published entries?
    No. You have created a different suite. Report the unmodified suite's number for comparison and your extended run separately.
  • What is the fastest legitimate use of such a suite in a release pipeline?
    As a non-regression tripwire: rerun it unchanged after every model or filter update and investigate any movement, without treating the absolute number as a safety claim.

It is a standardised road-test loop rather than a real commute: identical route, identical stopwatch, so two cars are comparable — and the route tells you nothing about the pothole outside your own office.

saying these in an interview costs you the question

  • Treating the suite's success rate as the model's overall jailbreak rate.
  • Not realising the judge is part of the frozen artefact and can be swapped.
  • Reporting a number without saying which suite revision produced it.
  • Assuming a frozen public suite covers attacks invented after it was published.

context

open as a page

A fixed-attack jailbreak benchmark like JailbreakBench ships a harm judge that labels each model response as jailbroken or not. Why is that judge part of the frozen artefact, and what breaks in your reported numbers if you substitute your own judge?

level: middleimportance: must knowfreq 60%

basics

~20 s

The judge decides what counts as a jailbreak, so it is half the measurement. Swap it and your numbers stop comparing to every published result on the suite: a stricter judge lowers the reported success rate, a looser one raises it. Report your own judge's result separately, never as the leaderboard number.

open as a page

Two candidate input filters were each measured on the same fixed-attack jailbreak benchmark, and one reports a lower attack-success rate. What must you verify before telling the team that filter is the stronger defence to ship?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Check that both ran the same frozen prompt set and the same judge, that the gap is bigger than run-to-run variation, and that neither filter was tuned on those exact prompts. Then look at the per-behaviour breakdown and at what each filter costs in wrongly blocked legitimate traffic.

open as a page

You run a long-published, fixed-attack jailbreak benchmark against a hosted chat endpoint that the vendor has patched many times, and it reports a very low attack-success rate. What are the reasons that number can badly understate your real exposure, and what would you run alongside it?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Its prompts are public and old, so vendors have likely patched exactly those strings and they may sit in training or filter data. A low score then proves those specific attempts fail, not that the model resists new ones. Pair it with freshly generated attacks and attacks aimed at your own application.

open as a page

Leadership proposes making a score on a public fixed-attack jailbreak leaderboard, such as a JailbreakBench-style suite, the mandatory release gate for every model or prompt update your product ships. What is the case on both sides, and what would you put in the gate instead?

level: principalimportance: should knowfreq 30%

basics

~20 s

For: it is cheap, repeatable and comparable across releases, so a regression is visible. Against: a public frozen suite can be optimised against and says nothing about your own application's attacks. Use it as a non-regression tripwire, and gate release on fresh adversarial runs against your deployed stack.

open as a page