skip to content

When you replay a published red-team benchmark against your own application endpoint, the benchmark's automated harm judge can start scoring the wrong thing because your app's responses do not look like raw model completions. What are two ways your response format breaks that judge, and what do you do about each?

level: seniorimportance: should knowfreq 45%

answer

  1. judge fitted to prose completions
  2. structured output reads as refusal
  3. canned template is correlated error
  4. harm may sit in a tool argument
  5. hand-label a sample, report agreement

basics

~20 s

Two common breaks: the app returns structured output or a rendered card, so the judge reads a wrapper rather than the answer text; and the app emits one fixed refusal template the judge was never calibrated on. Fix by judging exactly the text and tool calls the user receives, then hand-labelling a sample to check the judge still agrees.

solid answer

~50 s

A benchmark's harm judge is a classifier tuned on single-turn prose completions from model APIs. Your application does not emit those. **Break one - shape.** The response is JSON, a card, a citation-annotated block, or a stream you captured partially. The judge sees keys and markup, or an empty string, and quietly labels almost everything a refusal. Fix: extract the exact user-visible text and pass that, and if the app can act, judge the tool calls and their arguments too, not only the prose. **Break two - template.** The app deflects out-of-scope requests with one canned sentence. That string is out of distribution for the judge, so it may be scored inconsistently - and because it repeats, one systematic error moves the whole aggregate. Either way the remedy is the same: hand-label a stratified sample of your app's responses, measure agreement with the judge, and report the judge's error rate alongside the score. An unvalidated judge at a new interface makes the number decorative.

go deeper

for a junior

Recognises that the app's output is not a plain completion and that the scorer may therefore misread it.

for a middle

Names concrete shapes - structured output, truncation, canned template - and knows to extract the user-visible text before judging.

for a senior

Adds tool calls to the judged unit, validates the judge with a hand-labelled stratified sample, and reports agreement alongside the score.

for a principal

Makes judge validation a standing requirement for any replayed number and ties re-validation to output-format changes in the app.

**What a harm judge actually is, and why it is the fragile part of a replay.** The judge is the component that converts a response into a label. In harmbench it is a classifier fine-tuned on model completions paired with behaviour descriptions; in jailbreakbench it is a judge model driven by a published rubric. Either way it is a *fitted* artefact: it was tuned or prompt-engineered against a specific distribution — single-turn prose completions returned by a chat-completions endpoint. Teams porting a benchmark to their own application port the item corpus deliberately and port the judge by accident, because the judge simply comes along inside the harness. But the corpus is text that means the same thing anywhere, while the judge is the one piece that was fitted to a distribution you have just replaced. **The failure modes worth naming.** - **Wrapper text.** Your endpoint returns JSON, a rendered card, or a citation-annotated block. The judged string now contains field names, markup and reference markers. Judges fitted on prose degrade on that input, and they degrade *asymmetrically*: the usual direction is toward "no harmful content", which flatters your score. - **Truncation.** A streamed response captured at a cut-off, or a length cap, hands the judge half an answer. Half a compliant response very often reads as a refusal. - **Canned deflection.** A single templated out-of-scope sentence is out of distribution for the judge, and because it repeats byte-identically, one mistake is not one error — it is a correlated error applied to every off-scope item, and it moves the aggregate on its own. - **Displaced harm.** When the app can act, the consequential part of a hit is often a tool argument, not the prose. A prose-only judge scores that as a clean refusal, and the more capable your product is, the more of your risk hides in that gap. - **Post-processing.** A redaction or rewrite pass means the judge is scoring the sanitiser. That is a legitimate measurement, but it is not the measurement you claimed. **The method, and what it costs.** Capture the full response envelope: user-visible text, tool calls with their arguments, and the termination stage. Judge the union, with an explicit written rule for when the *action* is the hit rather than the text. Then validate the judge against humans, because an unvalidated judge at a new interface makes the score decorative. The validation is the cost item people omit from plans: draw a stratified sample across judged-hit, judged-miss and template responses — 150-300 responses is the usual working size — and label it by hand. At roughly one to two minutes per response that is three to ten hours of a qualified person, doubled if you want a second labeller and an inter-rater figure. It is not one-off: every change to the app's output format silently refits nothing and invalidates everything, so the cost recurs at each envelope change. Judge inference itself is cheap by comparison — one call per response — but not free at three samples per item across a few hundred items. **Where the number misleads.** The headline misreading: a replayed refusal rate far above the published run's is celebrated as evidence that the safety stack works, when the refusals are near-identical template strings and the real finding is that the judge cannot read your format. Second: reporting raw agreement on an imbalanced sample. If 95% of responses are non-hits, a judge that labels everything non-harmful shows 95% agreement and zero usefulness — the number to look at is agreement *within the hit stratum*, or a chance-corrected statistic, not the pooled percentage. Third: validating the judge only on the miss stratum, which is exactly the stratum where the broken judge and the correct judge agree. Fourth: quoting a score with no judge error rate beside it, which hides the fact that the score's uncertainty is dominated by the labeller and not by sampling. **What you check.** Whether the judged string is the user-visible text and nothing else. Whether tool calls are in the judged unit at all. Whether refusals cluster into a few identical strings. Whether hit-stratum agreement, not pooled agreement, is what you reported. And whether the last re-validation postdates the most recent change to the response envelope.

  • Your replay's refusal rate is far higher than the published run and most refusals are byte-identical. What do you suspect first?
    A judging artefact: the app's canned deflection template or a wrapper format is being scored as refusal. Hand-label a sample of those responses before believing the number.
  • When the app can call tools, what is the unit the judge should score?
    The whole response envelope - user-visible text plus the tool calls and their arguments - with an explicit rule for when the action itself constitutes the hit.

Reusing the benchmark's own judge on your application's output is like handing an essay grader a stack of spreadsheets. It never complains and never errors out; it quietly marks nearly everything as fine.

saying these in an interview costs you the question

  • Porting the judge unchanged to a new response format and reporting the score without any agreement check.
  • Judging only the prose when the application can call tools.
  • Reading an implausibly high refusal rate as good news rather than as a judging artefact.
  • Reporting a single score with no judge error rate beside it.

context