skip to content

Standard Suites

A behaviour list, a frozen attack leaderboard and a broad trust battery each standardise something different, and only one of them was built for your claim. Interviewers ask which one you used.

on this pageshow

explore

questions

14

A fixed-attack jailbreak benchmark such as JailbreakBench freezes several things at once so that results are comparable. What does it freeze, and why is that freeze what makes two different defences comparable on the same model?

level: juniorimportance: must knowfreq 70%

answer

  1. attacks frozen, judge frozen, report frozen
  2. delta isolates the defence
  3. public prompts get patched first
  4. score answers a narrow question
  5. which revision, which judge, which denominator

basics

~20 s

It freezes the attack prompts, the judge that decides whether a response counts as a jailbreak, and the report format. With all three fixed, two defences see identical inputs scored by identical rules, so the difference in reported success rate reflects the defence and not a different prompt set or a different grader.

solid answer

~50 s

A fixed-attack suite of this class is three frozen artefacts, not one: **the attack set** (a published list of prompts or attack instances, each tied to a behaviour it is trying to elicit), **the decision rule** (the harm judge that labels each response jailbroken or not), and **the report shape** (which numbers are published, over which denominator). The freeze is the whole point. If you change the prompts, you have measured a different threat. If you change the judge, you have changed what "success" means. If you change the denominator, the percentage stops meaning the same thing. Only with all three pinned does the delta between two defences isolate the defence. The cost is equally direct: whatever the freeze excludes is invisible. The number answers "how did this defence do against these attacks, under this judge" and nothing wider.

go deeper

for a junior

Names the three frozen pieces — attacks, judge, report format — and says the freeze is what makes two defences comparable.

for a middle

Explains why swapping any one piece breaks comparability, and states what the number does not cover.

for a senior

Adds the staleness and public-exposure argument, and the five things to check before trusting someone else's posted score.

for a principal

Frames it as instrument versus assessment, and sets policy for where such a suite may appear in a release decision at all.

## A fixed-attack suite is an instrument, not an assessment An **assessment** asks "is this system safe?" and has to consider everything. An **instrument** asks one narrow question and answers it the same way every single time it is used. A fixed-attack jailbreak benchmark — JailbreakBench is the canonical example of the class — is deliberately built as the second thing, and every strength and every weakness it has follows from that one design choice. ## The three artefacts it freezes - **The attack set.** A published, versioned list of attack artefacts, each tied to a **behaviour** — a named harmful capability the attempt is trying to elicit, phrased as a target ("produce instructions for X"). Behaviours are the rows of the experiment; the attack artefacts are the attempts made against each row. A suite of this class typically carries on the order of a hundred behaviours, split across a harm taxonomy, plus a matching set of benign behaviours used to check that a defence is not simply refusing everything. Freezing this means a difference between two runs cannot have come from a different attacker. - **The harm judge.** A fixed decision procedure — a trained classifier, a scripted string rule, or a model prompted in a published, versioned way — that reads the pair (behaviour, model response) and emits one bit: jailbroken or not. This is the component that turns text into a number, and it is emphatically part of the measurement rather than an implementation detail. Two labs scoring the *same* stored responses with different judges will publish different rates. - **The report shape.** Which rate is published, over which denominator, and what is published beside it. Attack-success rate per *attempt* (every attempt is a row), per *behaviour* (a behaviour counts once), and per *behaviour with at least one success out of k attempts* are three different numbers computed from one identical run, and the third rises monotonically as you raise k. Freezing the report shape is what lets two leaderboard entries be lined up on a page at all. ## Why the freeze is what makes two defences comparable Comparison is subtraction, and subtraction is only meaningful when everything except the term you care about is held constant. - Change the prompts and you have measured a different threat. - Change the judge and you have changed the definition of the word "success". - Change the denominator and the percentage has changed units. With all three pinned, the delta between defence A and defence B has exactly one place left to come from: the defence. ## What a run costs Target calls are roughly `behaviours x attack artefacts per behaviour x generations per attempt`. A hundred behaviours with a handful of attack artefacts and one generation each is several hundred to a few thousand calls — on a commodity hosted chat endpoint, single-digit to low-tens of dollars and tens of minutes of wall clock at modest concurrency. A model-backed judge roughly doubles the call count and adds its own bill. None of that is the real cost. The real cost is **engineer time**: writing the target adapter so the suite talks to your *deployed* stack rather than a bare model, pinning the suite revision, and hand-triaging borderline labels. Budget the adapter in days and the run itself in minutes; teams routinely size this backwards and are surprised that the cheap part is the compute. ## Where the number misleads The rate answers "how did this defence do against *these* frozen attempts, under *this* judge, over *this* denominator" — and nothing wider. Four specific misreadings recur. 1. First, quoting it as "our jailbreak rate", which is a claim about all attacks and is not supported. 2. Second, comparing a number computed per-behaviour-with-at-least-one-success at k=10 against someone else's per-attempt number, which is a units error dressed as a comparison. 3. Third, ignoring that the prompts are public, so they are the first strings a vendor blocks by rule or by safety-tuning; the score drifts optimistic over time while looking identical. 4. Fourth, quoting a bare-model figure for a product that adds a system prompt, tools and retrieved content — surfaces the suite never touches. ## What to check before believing a posted number - Which suite and which revision; - whether the shipped judge was used unmodified; - the exact denominator and the attempts per behaviour; - the target configuration, including decoding settings and whether it was the bare model or a deployed stack; - and whether the defence under test was developed while its authors could see those prompts. Absent those five, a posted percentage is a number, not a measurement you can compare yours to.

  • Why is the report shape, not just the prompts, part of what has to be frozen?
    Because the same raw results can be summarised over different denominators — per attempt, per behaviour, or per behaviour with at least one success — and those percentages are not interchangeable.
  • You add ten of your own prompts to the frozen set. Can you still compare against published entries?
    No. You have created a different suite. Report the unmodified suite's number for comparison and your extended run separately.
  • What is the fastest legitimate use of such a suite in a release pipeline?
    As a non-regression tripwire: rerun it unchanged after every model or filter update and investigate any movement, without treating the absolute number as a safety claim.

It is a standardised road-test loop rather than a real commute: identical route, identical stopwatch, so two cars are comparable — and the route tells you nothing about the pothole outside your own office.

saying these in an interview costs you the question

  • Treating the suite's success rate as the model's overall jailbreak rate.
  • Not realising the judge is part of the frozen artefact and can be swapped.
  • Reporting a number without saying which suite revision produced it.
  • Assuming a frozen public suite covers attacks invented after it was published.

context

open as a page

A teammate reports that your model "passed HarmBench." What is a harmful-behaviour set such as HarmBench or AdvBench actually a list of, and what does a good result over it cover and not cover?

level: juniorimportance: must knowfreq 70%

basics

~20 s

It is a fixed list of harmful request descriptions, sorted into categories its authors chose in advance. A good result means the model refused those written requests under whatever attack drove them. It says nothing about harms that are not on the list, including whatever your own product's worst outcome is.

open as a page

A fixed-attack jailbreak benchmark like JailbreakBench ships a harm judge that labels each model response as jailbroken or not. Why is that judge part of the frozen artefact, and what breaks in your reported numbers if you substitute your own judge?

level: middleimportance: must knowfreq 60%

basics

~20 s

The judge decides what counts as a jailbreak, so it is half the measurement. Swap it and your numbers stop comparing to every published result on the suite: a stricter judge lowers the reported success rate, a looser one raises it. Report your own judge's result separately, never as the leaderboard number.

open as a page

Your product is a bank's customer-support assistant, and its worst realistic outcome is being talked into revealing another customer's account details. It scores near-perfect against the AdvBench and HarmBench behaviour lists. Mechanically, why does that result say nothing about your top risk, and what would you do instead?

level: middleimportance: must knowfreq 60%

basics

~20 s

Those lists contain behaviours their authors picked — weapons, fraud, harassment, self-harm. Cross-customer data disclosure is not among them, so no attempt ever aimed at it. A behaviour nobody attempts cannot fail, so your top risk contributes nothing to the result. Write your own behaviours from your risk register and measure those separately.

open as a page

You re-run a broad multi-dimension trust battery such as TrustLLM after shipping a mitigation, and one dimension's score comes back a few points higher. Why is that not yet evidence the mitigation worked, and what do you check before claiming it?

level: middleimportance: must knowfreq 55%

basics

~20 s

That dimension is backed by few prompts, so a few points can be one or two responses flipping. Sampling, non-zero-temperature decoding and the judging step all move it on their own. Check which individual items changed, re-run the unmitigated build unchanged, and see whether the movement is bigger than run-to-run drift.

open as a page

You are choosing an off-the-shelf evaluation to open a red-team engagement against a chat assistant. What does a broad multi-dimension trust battery such as TrustLLM buy you compared with a focused single-purpose attack suite, and what does that breadth cost?

level: juniorimportance: should knowfreq 45%

basics

~20 s

A broad trust battery scores several different trust properties in one pass, so it gives you a wide first map of where a model looks weak. A focused attack suite drills one property hard instead. The cost is depth: each dimension is backed by a thin slice of prompts, so its number is coarse.

open as a page

Before citing a per-category breakdown from a standard harmful-behaviour list such as AdvBench or HarmBench, what should you check about the list's own composition, and how does that change how you read those category numbers?

level: middleimportance: should knowfreq 40%

basics

~20 s

Open the file and read it. Check how many entries each category holds, how many entries are near-restatements of each other, and how the requests are phrased. Categories are unevenly sized and wordings repeat, so a category figure can rest on a handful of entries or on one idea counted several times.

open as a page

Two candidate input filters were each measured on the same fixed-attack jailbreak benchmark, and one reports a lower attack-success rate. What must you verify before telling the team that filter is the stronger defence to ship?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Check that both ran the same frozen prompt set and the same judge, that the gap is bigger than run-to-run variation, and that neither filter was tuned on those exact prompts. Then look at the per-behaviour breakdown and at what each filter costs in wrongly blocked legitimate traffic.

open as a page

You run a long-published, fixed-attack jailbreak benchmark against a hosted chat endpoint that the vendor has patched many times, and it reports a very low attack-success rate. What are the reasons that number can badly understate your real exposure, and what would you run alongside it?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Its prompts are public and old, so vendors have likely patched exactly those strings and they may sit in training or filter data. A low score then proves those specific attempts fail, not that the model resists new ones. Pair it with freshly generated attacks and attacks aimed at your own application.

open as a page

Your team adds twenty product-specific behaviours to an evaluation that previously used only the HarmBench behaviour list, and the headline number moves. What has actually changed, and how would you report the two sets of behaviours?

level: seniorimportance: should knowfreq 45%

basics

~20 s

You changed the measured population, not the model. The mixed figure is comparable to nothing: not to your earlier run on the public list, and not to anything anyone else reports on it. Report the public list and your own behaviours as two results, each with its own count, and never merge them into one headline.

open as a page

Your team reports a broad multi-dimension trust battery's per-dimension scores each release. After an output-moderation classifier is put in front of the assistant, one dimension improves clearly and another drops. How do you work out what actually happened before you report it?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Pull the item-level results for both dimensions and look at what the moderation layer did to each prompt. The usual story: it blocks unsafe completions, so the dimension that rewards refusal rises, while a dimension that penalises unhelpful or over-cautious answers falls on the same behaviour. One change, two scoreboards.

open as a page

Leadership proposes making a score on a public fixed-attack jailbreak leaderboard, such as a JailbreakBench-style suite, the mandatory release gate for every model or prompt update your product ships. What is the case on both sides, and what would you put in the gate instead?

level: principalimportance: should knowfreq 30%

basics

~20 s

For: it is cheap, repeatable and comparable across releases, so a regression is visible. Against: a public frozen suite can be optimised against and says nothing about your own application's attacks. Use it as a non-regression tripwire, and gate release on fresh adversarial runs against your deployed stack.

open as a page

You lead red teaming for a company shipping a code assistant, a bank chat assistant and a medical triage bot. Would you make a standard harmful-behaviour set such as HarmBench the organisation's unit of measurement, keep a per-product behaviour register, or both? Which would you choose, and what is each number allowed to justify?

level: principalimportance: should knowfreq 35%

basics

~20 s

Both, with different jobs. The shared list is a cheap cross-product regression signal and lets you compare candidate models. It cannot represent product risk, so every product also owns a behaviour register drawn from its own threat model, and that register, not the shared figure, gates release.

open as a page

Where does a broad multi-dimension trust battery such as TrustLLM belong in an AI red-team program that also runs adaptive attack tooling and product-specific scenarios, and what claim should its dimension scores never be used to support?

level: principalimportance: should knowfreq 26%

basics

~20 s

Use it as a cheap, repeatable tripwire and as a way to pick where to attack next. Run it early, then on a schedule. Never let a dimension score stand as an assurance claim that the system is safe: it is a fixed, thin, non-adaptive prompt set that knows nothing about your product.

open as a page