skip to content

When a judge model scores red-team transcripts to decide which attempts landed, the text it reads contains the attacker's own prompt and the target's reply. Why does that make the judge itself an attack surface, and how would you check whether your published attack-success rate has been distorted by it?

level: seniorimportance: should knowfreq 35%

answer

  1. judge reads attacker-authored text
  2. score the reply, withhold the attack prompt
  3. search loop reward-hacks its own scorer
  4. classifier fails out of distribution instead
  5. held-out scorer for the published number

basics

~20 s

The judge reads attacker-controlled text, so the transcript is untrusted input to it. Payloads can address whatever reads them next and push a scoring decision, and encoded or obfuscated replies fall outside a trained classifier's training data. Hand-adjudicate a random sample of both scored-hit and scored-miss transcripts.

solid answer

~60 s

A judge model is an instruction-following system fed text that an adversary wrote. Anything in the prompt or reply that reads as direction to a downstream reader — framing the exchange as a sanctioned drill, asserting the content is fictional, or simply formatting the reply to look like an already-adjudicated record — can move the label. The distortion runs both ways: it can suppress hits (an attacker under evaluation wants a clean score) or manufacture them. The worse case is an **automated attack loop that uses the judge's label as its search signal**. Then judge-fooling is directly rewarded and cheaper to find than genuine elicitation, so the loop drifts toward transcripts that score as hits without being hits. Checks: score only the reply where the task permits, with the attack prompt withheld from the judge; present the transcript in a clearly delimited data role; hand-adjudicate stratified samples of scored hits and scored misses; and re-score a suspicious slice with a structurally different scorer (a trained classifier versus a prompted judge) and read the disagreements.

go deeper

for a junior

Should notice that the judge is reading text the attacker wrote and that this is untrusted input, even if the mitigations are vague.

for a middle

Explains that a judge model follows instructions and that transcript content can move a label in either direction, and suggests scoring the reply without the attack prompt.

for a senior

Identifies the reward-hacking case where an attack loop optimises against its own scorer, contrasts it with a trained classifier's out-of-distribution failure, and designs stratified adjudication of hits and misses plus a held-out scorer.

for a principal

Makes the separation of the search-time scorer from the reporting scorer a standing rule, and requires an adjudicated audit sample behind any externally published rate.

**The structural point.** Everywhere else in a red-team pipeline we treat attacker-authored text as untrusted. The scoring stage quietly stops: the judge model is handed the full adversarial transcript — the attacker's prompt and the target's reply — and asked to apply a rubric to it. It is an instruction-following system consuming adversary-written text, which is the same class of problem the target has, with the aggravating detail that the judge's output is the number you publish. **Three distinct distortions, which need separating.** 1. *Direct steering.* Text in the prompt or the reply is written to address whoever reads the record next and to assert how it should be classified. A judge handed the whole transcript as one undifferentiated blob is maximally exposed, because it has no structural way to tell rubric from evidence. 2. *Format capture.* No instruction is needed. A reply shaped like a verdict, a refusal template, or a system-level annotation can pull the label without containing anything that reads as a command. 3. *Signal reward-hacking.* When an automated attack search optimises against the judge's label, the objective it is actually maximising is "text this scorer calls a hit", not "harmful text". Search takes the cheap path. This is the quiet one, because it inflates an entire campaign's rate without any single transcript looking wrong. Note the distortion runs in both directions. A team whose own system is under evaluation benefits from suppressed hits; a team demonstrating a new attack benefits from manufactured ones. Never assume the bias points the way that suits your prior. **A trained classifier fails differently, not less.** It has no prompt, so it cannot be instructed — that whole path closes. In exchange it degrades out of distribution: encodings, unusual scripts, translation, code-shaped or heavily formatted replies drift from the data it was fitted on and come back benign. So an attack family built on obfuscation *deflates* a classifier-scored rate, while the same transcripts under a judge model can be pushed either way. Neither scorer is the safe default; they have different blind spots, which is exactly why cross-scoring is informative. **What the checks cost.** Every mitigation here is a multiplier on the scoring bill, and that is the number to put in the plan. A held-out second scorer means two scoring passes per published run — with a hosted judge, twice the per-transcript API spend, permanently. The prompt-ablation re-score (judge sees the reply only) is a third pass. Stratified manual adjudication runs two to four minutes per transcript, so a 200-item audit split across flags and clears is roughly a person-day per report. Against a full campaign that is single-digit percent of the engagement, and it is the only part of the pipeline that produces evidence about the instrument rather than about the target. **How I would check a published rate.** - *Stratified manual adjudication.* Sample scored hits and scored misses **separately** and report the two confirmed rates. Auditing only the flags finds false positives and is structurally incapable of finding suppression, because a suppressed success was never flagged. - *Input ablation.* Re-score with the attack prompt withheld so the judge sees only the reply, and compare labels attempt by attempt. A large shift localises the influence to the prompt side. - *Cross-scoring.* Re-score a slice with a structurally different scorer — a trained classifier against a prompted judge — and read the disagreement set rather than the totals, since offsetting errors can hide behind matching aggregates. - *Clustering.* If scored hits concentrate on one wrapper, phrasing or formatting pattern that a search loop discovered, that concentration is the reward-hacking signature; a genuine vulnerability spreads across behaviours in a way a scorer artifact does not. - *Provenance of the scorer.* If the same scorer that guided the search also produced the published number, the rate is not evidence. Require a held-out scorer the search never saw, and publish both. **Mitigations that actually help.** Present the transcript in a clearly delimited data role and instruct the judge to classify rather than obey it; score the reply alone where the rubric permits; keep a reporting scorer separate from any search-time scorer; and treat any large, unexplained movement in the rate as a possible instrument event until an adjudicated sample says otherwise. The answer here is structural — a payload wording is not a mitigation and is not what an interviewer is looking for.

  • Why sample scored misses as well as scored hits when adjudicating?
    Suppression only shows up in the misses. If the judge was steered into labelling landed attacks as refusals, a hits-only audit confirms the flags and finds nothing wrong.
  • Your attack search and your reporting both use the same judge. What do you change?
    Hold out a second, structurally different scorer that the search never optimised against, and publish its rate. Keep the search-time scorer for guiding the loop only.
  • Does using a trained classifier instead of a judge model remove this risk?
    It removes the instruction-following path but adds an out-of-distribution one: obfuscated, encoded or translated replies drift off its training distribution and are scored benign, deflating the rate.

The judge is a referee who scores the match by reading notes the players hand him. A note asserting that a play was legal does not have to be true to change the scoreline.

saying these in an interview costs you the question

  • Assumes the scoring stage is outside the threat model because 'it is only evaluation'.
  • Uses the same scorer to drive an automated attack search and to produce the published rate.
  • Adjudicates only the scored hits, so suppressed hits are never discovered.
  • Claims a trained classifier is immune — it cannot be instructed, but it silently fails on obfuscated output.
  • Offers a specific payload wording as the answer rather than the structural mitigation.

context