skip to content

A fixed-attack jailbreak benchmark like JailbreakBench ships a harm judge that labels each model response as jailbroken or not. Why is that judge part of the frozen artefact, and what breaks in your reported numbers if you substitute your own judge?

level: middleimportance: must knowfreq 60%

answer

  1. judge sets the numerator
  2. swap judge = non-comparable entry
  3. strict vs permissive shifts every row
  4. score twice: shipped judge and yours
  5. keep transcripts, re-label offline

basics

~20 s

The judge decides what counts as a jailbreak, so it is half the measurement. Swap it and your numbers stop comparing to every published result on the suite: a stricter judge lowers the reported success rate, a looser one raises it. Report your own judge's result separately, never as the leaderboard number.

solid answer

~50 s

A jailbreak success rate is a ratio whose numerator is decided entirely by the labeller. The attack set fixes the denominator; the judge fixes the numerator. Two teams running the identical frozen prompts against the identical model can publish materially different rates purely because one judge counts a hedged, partially-compliant answer as a hit and the other does not. That is why the suite ships a judge rather than leaving it to you. Substituting your own — even a better one — silently redefines success and makes your entry non-comparable, while still looking like the same metric. The legitimate pattern is to run both: the shipped judge for the comparable number, and your own stricter or product-specific judge for the number you actually act on. Report them as two labelled columns. Also spot-check a sample of judge labels by hand; judge disagreement clusters exactly on the borderline responses that drive small leaderboard deltas.

go deeper

for a junior

Says the judge decides what counts as a success, so changing it changes the score and breaks comparison with published results.

for a middle

Explains numerator versus denominator, that a permissive judge shifts every entry, and proposes scoring twice.

for a senior

Adds judge error profile, partial compliance, drift of a model-backed judge, and stratified hand-labelling of the borderline band.

for a principal

Sets the org rule: transcripts are the durable artefact, the shipped judge is for comparability, and release decisions run on an in-house judge tuned to the product's harms.

**The judge is a measuring device, and it has its own error profile.** A jailbreak attack-success rate is a ratio. The frozen attack set fixes the denominator: how many attempts, against how many behaviours, at what attempt budget. The judge fixes the numerator: which of the resulting responses are counted as hits. That is not a supporting role. It is half of the number, and it is the half that involves a fallible classifier reading ambiguous text. Every harm judge has two error rates. Its **false-positive rate** is how often it labels a refusal, a lecture, or a harmless generic answer as a jailbreak. Its **false-negative rate** is how often it misses a genuinely compliant answer — typically one wrapped in fiction, one written in another language, or one that opens with a refusal paragraph and then complies in the third. Both rates flow straight into the headline. A judge five points more permissive than another shifts *every* entry it scores by roughly five points, which is frequently larger than the gap between the two defences you were trying to distinguish in the first place. **Why the suite freezes it rather than leaving it to you** Without a frozen judge a leaderboard is not a leaderboard. Entries labelled by different graders sit on different scales, and there is no way to tell a genuinely stronger defence from a stricter grader. Freezing the judge is exactly as load-bearing as freezing the prompts, and the standard mistake is to treat the prompt set as "the benchmark" and the judge as a utility you may swap for a better one. **What breaks the moment you substitute** Your entry stops being comparable to every published number on that suite, while continuing to look like the same metric with the same name and the same units. That is the dangerous part: nothing errors, nothing warns, and a slide reading "12% on JailbreakBench" is now unfalsifiable unless the judge is named beside it. A stricter in-house classifier will usually push the rate down and make your defence look better than the leaderboard would; a permissive one flatters the attacker. **Where the shipped judge legitimately falls short** It was calibrated for the suite's own harm taxonomy, not your product's. A response that is unremarkable in general can be disastrous inside your domain — a plausible-sounding dosage, a fabricated policy quote, an internal identifier — and the shipped judge will label it clean. It may also be blind to partial compliance. And if the judge is itself a model behind a hosted API, its behaviour can move under you between runs as the provider updates it, which means your supposedly frozen instrument was never fully frozen. **The pattern that keeps both properties** Score every run twice, from the same stored transcripts: ``` stored responses -+-> shipped judge (pinned) -> the comparable, publishable number \-> your product judge -> the number the release decision uses ``` Report them as two labelled columns, never as one number. Store the raw transcripts, not just the labels: with transcripts, a corrected judge, a new judge, or a scoring bug can be re-applied to every historic run offline. **What it costs** Judging is the cheap half if you keep transcripts, and the expensive half if you do not. Re-labelling a few thousand stored responses with a model-backed judge is a handful of dollars and minutes; re-running the attacks to get those responses back is the full benchmark bill again against a metered endpoint, plus the risk that the endpoint has changed underneath you so the new run is not the old run. Hand-labelling is the genuinely expensive item: a stratified sample of 100–200 responses is several hours of an experienced reviewer, and it needs a second reviewer on a subset if you want to quote human agreement rather than one person's opinion. **Where the number misleads** Two failure modes dominate. First, a judge-driven delta read as a defence-driven delta — you tightened the grader between quarters and the improvement is entirely artefact. Second, a real but unresolvable gap: if defence A and defence B are separated by two points and the judge disagrees with human labels on five points' worth of borderline responses, you have not measured a difference, you have measured the grader's ambiguity band. **What to check** That the shipped judge ran unmodified and its version is recorded; that transcripts are stored, not just labels; that a stratified hand-label sample — clear refusals, clear hits, and heavily weighted to the borderline band — has been done recently; and that judge-human agreement on that band is reported next to any gap you are about to act on.

  • Your judge is clearly better than the shipped one. Do you still run the shipped one?
    Yes — it is the only number comparable to published entries. Run both and label the columns.
  • How do you tell whether a small gap between two defences is judge noise?
    Hand-label a stratified sample weighted to borderline responses, measure judge–human disagreement there, and treat gaps inside that band as unresolved.
  • Why keep full response transcripts rather than just the labels?
    So a judge change or a scoring bug can be re-applied to old runs offline, without paying to re-attack a metered endpoint.

The frozen prompts decide how many shots are taken; the judge decides which ones the referee calls goals. Swapping in your own referee mid-tournament can change the table without a single player improving.

saying these in an interview costs you the question

  • Publishing a suite number produced by a home-grown judge as if it were the standard one.
  • Never hand-checking any judge labels.
  • Discarding raw transcripts and keeping only pass/fail labels.
  • Assuming a model-backed judge behind an API is stable over months.

context