skip to content

Two candidate input filters were each measured on the same fixed-attack jailbreak benchmark, and one reports a lower attack-success rate. What must you verify before telling the team that filter is the stronger defence to ship?

level: seniorimportance: should knowfreq 45%

answer

  1. same revision, same judge, same denominator
  2. rerun: does the ordering survive
  3. variant arm exposes suite-fitting
  4. false blocks on real traffic
  5. per-behaviour, not the average

basics

~20 s

Check that both ran the same frozen prompt set and the same judge, that the gap is bigger than run-to-run variation, and that neither filter was tuned on those exact prompts. Then look at the per-behaviour breakdown and at what each filter costs in wrongly blocked legitimate traffic.

solid answer

~50 s

A leaderboard delta is a claim about one instrument, and four things can invalidate it before it becomes a shipping decision. **Was the instrument identical?** Same suite revision, same unmodified judge, same denominator, same target configuration, same number of attempts per behaviour. Any difference and you are comparing two measurements, not two filters. **Is the gap larger than the run's own variation?** Sampling from the target and, if the judge is model-backed, from the judge, both move the number between reruns. A small gap can be noise. **Was either filter fitted to this suite?** A filter developed while watching this exact prompt set will win on it and lose in production. Confirm with reworded variants of the same behaviours. **What does it cost?** A filter that blocks more attacks by blocking more of everything is not stronger. You need its false-block rate on legitimate traffic, plus a per-behaviour breakdown so a win concentrated in one harm category is visible rather than averaged away.

go deeper

for a junior

Checks that both filters were run on the same suite with the same judge before comparing the two numbers.

for a middle

Adds run-to-run variation and the false-block cost on legitimate traffic as things that can flip the decision.

for a senior

Designs the variant arm that detects suite-fitting and reads per-behaviour results rather than the aggregate.

for a principal

Makes the paired report — attack rate plus false-block rate plus variant drop — the standard artefact for any defence-selection decision.

**A leaderboard delta is a claim about one instrument. Five checks stand between it and a shipping decision.** **Step one: prove the two runs were the same measurement.** Record and compare, for each arm: the suite revision, whether the shipped judge ran unmodified and at what version, the reported denominator, the target endpoint and its decoding settings (temperature and top-p change how often a sampled response drifts into compliance), and the attempts per behaviour. A difference in any of these means you are comparing two measurements, not two filters — and none of it is visible in a summary percentage. This is the single most common reason two internal numbers disagree, and it costs nothing but bookkeeping to rule out. **Step two: distrust a small gap.** The target samples, and a model-backed judge samples too, so the number moves between reruns of an identical configuration. Rerun both arms and check that the ordering survives. A gap that flips between reruns is not a result. If you want a defensible statement rather than a vibe, the run-to-run spread across three reruns per arm is a serviceable proxy for the noise floor, and any gap inside it should be reported as unresolved. **Step three: test for suite-fitting — the decisive arm.** Add a variant arm: the same behaviours in reworded or restructured phrasings the filter's authors never saw, scored by the same judge. What matters is not either raw score but the *drop* from frozen to variant, per filter: ``` filter A: frozen 4% -> variant 31% (fitted to the published strings) filter B: frozen 9% -> variant 12% (generalises) ``` On the leaderboard A wins by five points. In production B is plainly the better filter, because A's advantage was a pattern match against strings an attacker will not use. This inversion is the most valuable thing the comparison can surface and the variant arm is the only thing that surfaces it. Note that a filter can be fitted without anyone cheating: if its rule set was iterated while the public suite was the feedback signal, fitting is the expected outcome. **Step four: price the defence in false blocks.** A filter that blocks more attacks by blocking more of everything is not a stronger defence, and attack-success rate alone cannot tell the two apart — a filter that rejects every request scores a perfect zero. Run both filters over a sampled slice of real, benign production traffic and report the false-block rate beside the attack rate, as a pair. This is the arm with the most product risk attached and the one most often skipped, because it needs a traffic sample, a privacy review of that sample, and someone to define what counts as a wrongly blocked request. **Step five: read the breakdown, not the average.** Per-behaviour or per-category results turn an average into a decision. A filter that halves the aggregate while leaving your highest-consequence category untouched is the wrong choice whatever its headline says, and an aggregate is designed to hide exactly that. **What the whole comparison costs** Four cells (two filters x two arms), each rerun two or three times, is roughly six to twelve times the base benchmark run — still only hours of wall clock and tens of dollars of API spend on a suite of this size, plus the judge's calls. The expensive line items are elsewhere: building the variant arm without letting it drift into a different difficulty, obtaining and clearing the benign traffic sample, and the analyst time to hand-check the borderline labels that drive a close call. Days of a person, not dollars of compute. **Where the number misleads** Three readings recur. A five-point win that is entirely a decoding-setting difference between the arms. A five-point win that sits inside rerun noise and reverses next week. And a five-point win that is real on the frozen prompts and inverts on variants — the case above, and the one that actually ships the worse filter. **What you tell the team** Not "filter A scores better", but: the instrument was identical across arms; the gap survived reruns; the variant arm shows A is fitted to the published phrasings while B generalises; B costs fewer false blocks on real traffic; ship B, and keep the frozen suite as the regression tripwire rather than as the reason.

  • Filter A wins on the frozen suite but collapses on reworded variants. What do you recommend?
    Ship the filter that generalises. Report both arms so the team sees that A's frozen-suite lead came from matching published phrasings.
  • Why report per-behaviour results rather than one aggregate rate?
    Because an average can hide a filter that leaves your highest-consequence harm category entirely unimproved.

saying these in an interview costs you the question

  • Ranking two defences on a single run of one frozen suite.
  • Ignoring the false-block cost on legitimate traffic.
  • Never testing reworded variants, so suite-fitting goes undetected.
  • Comparing arms that used different suite revisions or a modified judge.

context