You have one attack-success rate published by HarmBench and one published by JailbreakBench for the same open-weights model. What has to line up before you can put both figures in a single comparison table, and what usually does not?
answer
- two instruments, not two measurements
- list, attempts, attack, target, judge
- confounded difference, no attribution
- re-run both lists in one harness
- report per list, never pool
basics
~20 sBoth rates must count over the same population and be produced the same way: the same behaviour list, the same attempts per behaviour and aggregation rule, the same attack method, the same target configuration, the same procedure for ruling a hit. Independently published suites share none of these, so their headline rates belong on different axes.
solid answer
~60 sEach suite is a complete instrument: a behaviour list, a way of driving attempts against a target, and a procedure that rules whether a response counted. Two suites differ on all three at once, so a difference between their headline rates is confounded and you cannot attribute it to the model. What must match for a legitimate side-by-side: - **Population** — the same items, or at minimum the same harm categories with the same difficulty mix. - **Denominator and aggregation** — behaviours versus attempts, and whether any-of-n counts a behaviour broken. - **Attack** — the same method, same budget of turns or queries. - **Target** — raw weights with the same chat template and decoding settings, or the same deployed endpoint including its system prompt and filters. - **Ruling** — the same judging procedure applied to both sets of transcripts. In practice you cannot get this by quoting two papers. You get it by re-running both behaviour lists yourself through one harness and one judge, and then reporting each list separately rather than pooling them.
go deeper
Should recognise that different suites use different prompt sets and hesitate to compare the percentages directly.
Should enumerate the confounds — list, attempts, aggregation, attack, target, judge — and propose re-running both lists in one harness.
Should also raise target configuration and decoding settings, refuse to pool lists, and note that public lists age and may leak into training data.
Should set an organisational rule for which instrument is canonical, what gets published internally, and how new lists are adopted without breaking existing series.
A published attack-success rate is the output of a complete instrument: an item list, a way of driving attempts at a target, and a procedure that rules whether a response counted. HarmBench and JailbreakBench are two instruments calibrated to nothing in common, so the gap between their headline numbers is not a measurement of anything about the model. ### What differs, and why every difference alone moves the rate - **Population.** Each suite curates its own behaviour list — different size (JailbreakBench's list is on the order of a hundred behaviours; HarmBench's is several hundred), different harm taxonomy, different difficulty mix. A list weighted towards categories a model was heavily tuned on scores low for population reasons. - **Denominator and aggregation.** One suite may report the fraction of *behaviours* broken by any of n samples; another the fraction of *attempts* ruled harmful. These are different statistics even on identical transcripts. - **Attack and budget.** A rate is always the rate of some method at some number of turns or queries. A stronger or simply longer-running attack lifts the number without touching the model. - **Target configuration.** Raw weights with a bare chat template, versus the same weights with a chat wrapper, a system prompt, or a hosted endpoint's own filters. Also decoding: temperature and top-p change what a resampled attempt produces. - **Ruling.** Each suite ships its own harm judge — a classifier or a prompted judge model with its own rubric. Two judges scoring the same transcripts routinely disagree on a noticeable fraction of borderline cases, and that disagreement lands entirely in the numerator. Five free variables, one observation. Nothing can be attributed. ### The move that actually works Hold the instrument fixed and vary only the thing under comparison: 1. Pick one harness and one ruling procedure. 2. Run both behaviour lists through it against the same target configuration, same decoding settings, same attempts per item, same aggregation rule, same attack budget. 3. Report two numbers labelled by list, with the per-category breakdown beside them. Do not average the two into one figure. The lists have different category mixes, so a pooled rate moves with the relative sizes of the two lists — a purely clerical quantity — rather than with the model. ### What that costs This is not free, and the cost is why teams quote papers instead. Two full sweeps of generations plus judging: on the order of `(|list A| + |list B|) x attempts x methods` target calls and a matching number of judge calls, which for a few hundred behaviours at a couple of dozen samples is tens of thousands of calls per model, hours to a day of wall clock against a rate-limited endpoint, and a repeat of the whole thing for every model in the comparison. The larger cost is usually engineer time: the two lists arrive in different schemas with different harm taxonomies and different notions of what a target response is, and normalising them into one harness so that "the same judge" really is the same judge is days of work, not an afternoon. Budget for it once and reuse it; budget for it per-report and it never happens. ### Where the number misleads The seductive reading is that the higher rate is the more rigorous or more honest measurement — that one suite is "harder". It may simply have an easier list scored by a more permissive judge with more attempts per item. There is no shared scale on which one suite over-reports relative to another, so the difference has no sign you can interpret and certainly no magnitude. Two further traps: a leaderboard that ranks models by pooling across suites inherits every one of these confounds at once; and even one suite's numbers from two dates are not automatically one series, because suites revise items and update judges between releases. Public lists also age — once a list is widely mirrored, later models may effectively have seen it, and the rate drifts for reasons unrelated to defences. ### What you would check Before drawing any bar chart: confirm both figures state their aggregation rule and attempts per item; confirm both targeted the same kind of object (bare weights versus a wrapped endpoint); take a couple of hundred transcripts and score them with *both* suites' judges to measure the disagreement rate directly, because that number sets how much of the gap is judging rather than model. If the disagreement is material, the comparison is dead on arrival and the honest deliverable is two labelled series plus the per-category breakdown, which is what a defence can act on anyway.
- Your leadership still wants one number. What do you give them?One number from one instrument, defined explicitly: this list, this attack, this many attempts, this judge, this target configuration, tracked over time. Plus the per-category breakdown, since that is what a defence can act on.
- Is it valid to compare two models on one suite?Much more valid, because the instrument is held fixed. It still requires the same target wrapper, decoding settings and attempts per model, and it only ranks them on that list under that attack.
- Why can pooling two behaviour lists into one rate be worse than reporting them separately?The pooled denominator has no defined population: category mixes and difficulty differ, so the combined rate moves with the relative sizes of the lists rather than with the model.
saying these in an interview costs you the question
- Putting two suites' headline percentages in one bar chart
- Averaging rates from different behaviour lists into a single figure
- Attributing the gap between two published rates to model robustness
- Assuming both suites used the same judging procedure because both call it a jailbreak
- Comparing a raw-weights rate with a deployed-endpoint rate