skip to content

Two teams publish attack-success rates for the same open-weights model on the same behaviour list. One scored replies with the trained harm classifier the benchmark distributes; the other scored them with a hosted chat model prompted with a harm rubric. The rates differ by 14 points. How do you work out whether the model or the scoring choice explains the gap, and what do you require before putting both numbers in one table?

level: seniorimportance: must knowfreq 50%

answer

  1. 2x2: transcripts crossed with scorers
  2. archive raw completions or nothing is re-derivable
  3. disagreement set, not just totals
  4. definitional gap vs scoring error
  5. pin the classifier, pin the rubric prompt

basics

~20 s

Hold the transcripts fixed and vary the scorer. Get both teams' raw completions and run both judges over both sets. If each judge gives similar numbers on either set, the gap is the scoring choice, not the model. Without archived completions the two rates cannot share a table.

solid answer

~60 s

**The experiment is a 2x2.** Two transcript sets (team A's generations, team B's generations) crossed with two scoring rules (the distributed classifier, the rubric-prompted judge). Score every cell. - Rows agree within a judge, columns differ → the judges disagree; the model did nothing. - Columns agree within a transcript set, rows differ → generation differed (decoding settings, system prompt, attack implementation, attempts per behaviour). - Both move → you have two confounds and must fix one before claiming anything. **Prerequisites for the 2x2 to be runnable at all:** raw completions archived per attempt with the prompt, decoding parameters and seeds; the classifier available as a pinned artifact; the rubric judge pinned as an exact prompt plus a model snapshot. If a team published only the aggregate, its number is not re-derivable and does not belong in a shared table. **And note the gap may not be error at all.** The two scorers can encode different definitions of a hit — any actionable detail versus a complete, workable procedure. That is a rubric disagreement, and no amount of accuracy work reconciles it; you have to pick one definition and re-score everything under it.

go deeper

for a junior

Should recognise that a different scoring rule can change the reported rate without the model changing, and that you would need the raw replies to check.

for a middle

Lays out the cross-scoring experiment and names the archival prerequisites — raw completions, decoding settings, a pinned scorer.

for a senior

Also inspects the per-transcript disagreement set, distinguishes scoring error from a definitional difference in what counts as a hit, and reports the judge effect and the generation effect separately rather than reconciling them into one figure.

for a principal

Turns it into policy: no red-team number is publishable unless its transcripts and scorer are archived well enough for a third party to re-derive it, and the table names the scorer in every cell.

**Why the summary statistics cannot settle it.** An attack-success rate is the output of two stages: a generation pipeline (the attack implementation, the target's decoding settings and system prompt, how many attempts per behaviour, how those attempts were reduced to one per-behaviour result) and a scoring rule (here, the benchmark's distributed harm classifier versus a hosted chat model prompted with a harm rubric). The two teams differed in *both*, so 14 points is a sum of two effects with no way to split it from the aggregates. Reasoning about which judge is "better" is not a substitute; the only clean move is to hold one factor fixed and vary the other, and that needs the artifacts, not the numbers. **Run the 2x2.** ``` classifier rubric judge team A outputs a11 a12 team B outputs a21 a22 ``` `|a11 − a12|` and `|a21 − a22|` are the judge effect; `|a11 − a21|` and `|a12 − a22|` are the generation effect. Report both. If the rows move together and the columns do not, the model did nothing and the scoring choice explains the gap. If both move, you have two confounds and must fix one — usually by re-generating one team's configuration yourself — before claiming anything. **What the cross costs.** Two extra scoring passes over both transcript sets. With the distributed classifier that is GPU time: a fine-tuned model in the low tens of billions of parameters, one forward pass per transcript, so a few thousand transcripts is minutes to an hour on one accelerator, and free of per-call charges. With the hosted rubric judge it is one API call per transcript per pass, so the cross doubles an already metered bill and is rate-limit-bound in wall clock. Archiving the completions that make the cross possible is cheap in storage terms — a few kilobytes per attempt, single-digit gigabytes for a large campaign — so the real expense is engineer time and the discipline to have pinned the artifacts before the run, not disk. **Where the number misleads, part one: equal totals are not agreement.** Two scorers can land within a point of each other on the aggregate while disagreeing on a third of individual transcripts, because their false positives and false negatives offset. Any conclusion drawn from matching totals is unfounded. Pull the disagreement set — the transcripts the two rules label differently — sample it, adjudicate by hand, and report per-item agreement, not the totals. **Part two: the gap may not be error at all.** Adjudicate each disagreement into (a) a scoring mistake by one rule, or (b) a genuine definitional difference about what counts as a hit. Only (a) is fixable by picking the better judge. Rubric judges commonly score a refusal-with-partial-explanation as a partial success where a binary classifier does not; they are often more permissive on fictional framing; and "any actionable detail" versus "a complete, workable procedure" are two different metrics wearing the same name. No accuracy work reconciles a definitional gap — you choose one definition and re-score everything under it. **Part three: the boring mechanical causes.** Before crediting anything conceptual, check truncation (one pipeline scored only the first N tokens of long replies), whether the judge saw the attack prompt as well as the reply, format drift — encoded, translated or code-shaped replies sit far from a trained classifier's fitting distribution and come back benign — and whether the two teams reduced multiple attempts per behaviour differently (any-of-n versus per-attempt is a denominator change, not a judge effect, and it can carry the whole 14 points on its own). **What I require before the two numbers share a table.** Per attempt: prompt, full untruncated completion, decoding parameters, seed, attempts per behaviour, and the reduction rule. Per scorer: a pinned artifact, or an exact rubric prompt plus a model snapshot identifier, together with error rates measured on an adjudicated sample of *this* transcript distribution. Given those, publish one number per (transcript set, scorer) cell and name which scorer defines the headline. Lacking them, publish the two rates separately, label each with its scorer and transcript set, and state plainly that they are not comparable. Averaging them, or taking the lower as "conservative", invents a number no experiment produced.

  • The two scorers produce nearly the same overall rate. Does that mean they agree?
    No. They can disagree on many individual transcripts with the errors offsetting. Check per-item agreement on the same transcript set, not just the totals.
  • One team stored only aggregate rates. What can you still salvage?
    Very little — you can re-run their attack configuration to generate fresh transcripts and score those, but that is a new experiment, not a reconciliation of their number. Their published rate stays unverifiable.
  • What would you write in the table caption if you can only publish both numbers as-is?
    The scorer that produced each rate, the transcript set it scored, and an explicit statement that the two cells are not comparable because the scoring rule differs.

saying these in an interview costs you the question

  • Averages the two rates, or picks the lower one as 'more conservative'.
  • Declares the trained classifier authoritative without checking it on this transcript distribution.
  • Compares only the aggregate rates and never looks at per-transcript disagreement.
  • Accepts a table of numbers whose underlying completions were never stored, so no re-scoring is possible.
  • Assumes any judge difference is error, missing that the two rules may define a hit differently.

context