skip to content

In a red-team report on a moderation guard, why present attack bypass across the guard's whole score range rather than as one percentage at the threshold the team currently runs?

level: middleimportance: must knowfreq 58%

answer

  1. one point hides the shape
  2. flat band vs. cliff edge
  3. scores piled against the line
  4. sweep is free if you kept scores
  5. hold the judge fixed

basics

~20 s

One threshold collapses the whole score range into a flattering number and hides where the guard actually breaks. A curve shows how bypass changes as the threshold moves, so the reader sees whether the guard is genuinely strong or just parked at a lucky setting, and what a small retune would cost.

solid answer

~1 min

The guard emits a score per prompt; the threshold is a line drawn through those scores. Reporting only what crosses the line throws away everything you learned about the distribution. Because you already have the per-prompt scores from the run, the curve is free: sweep the threshold over its range and recompute how many attack prompts land on the allow side at each value. What the curve shows and the single point cannot: - **Where the guard breaks.** A curve that is flat and then falls off a cliff says the guard separates your attacks cleanly; one that slopes gently across the whole range says it barely separates them at all, and the current number is an artefact of where the line sits. - **How fragile the number is.** If a small nudge to the threshold doubles bypass, the reported figure is one config change away from being wrong. - **What the reader is buying.** Every step toward stricter blocking costs benign traffic, so a curve makes the trade visible instead of asserting a single "good" answer. You still quote the number at the operating point in force — that is the reader's reality. You just refuse to let it stand alone.

go deeper

for a junior

Should say that one number depends on where the threshold sits and that showing several thresholds is more honest.

for a middle

Explains that the curve comes free from retained per-prompt scores, and that its shape distinguishes a guard with margin from one whose attacks sit against the line.

for a senior

Holds the success judge fixed across the sweep, calls out that stricter settings cost benign traffic that this corpus cannot measure, and still quotes the number at the operating point actually in force.

for a principal

Makes score retention a standing requirement of the harness so every engagement's guard result is re-derivable at any threshold, and defines how findings are compared across guards and quarters.

## One number, two very different guards A moderation guard scores each prompt; the **block threshold** is a line drawn through those scores; the **bypass rate** counts attack prompts on the allow side of the line whose responses a fixed judge called successful attacks. Reporting only the count at the line throws away the distribution you already paid to collect — and the distribution is where the guard's actual quality lives. Take two guards, each measured at 4% bypass on the same attack corpus at their configured thresholds. Sweep the threshold across its range and they stop looking alike immediately: | block threshold | guard A bypass | guard B bypass | |---|---|---| | strictest | 2% | 1% | | one step stricter | 3% | 2% | | **as configured** | **4%** | **4%** | | one step looser | 5% | 30% | | most permissive | 22% | 41% | Guard A holds near 4% across a wide band: your attacks score far from its line, so the result is a property of the guard. Guard B is at 4% only at this exact setting — its attacks are piled against the line, and the number is a property of where somebody happened to put it. One figure cannot separate those two situations; the sweep separates them at a glance. ## The mechanism The sweep is post-processing over rows you already have, one row per attempt: the prompt, the guard's score (per category if it returns them), the model's response, and the judge's verdict on whether the attack worked. ``` for t in sweep(min_score, max_score): bypass[t] = count(rows where guard_score < t -> allowed AND judge_verdict == attack_succeeded) / n_attacks ``` Two things people get wrong here. First, a bypass is not "the guard allowed it" — it is "the guard allowed it **and** the attack worked". Sweeping the threshold moves only the first term; the judge verdict is a fixed column. Second, the sweep is recoverable only if scores were retained per attempt. A harness that stored the boolean verdict has thrown the axis away, and the curve is not recoverable by any amount of cleverness. ## What it costs If scores were retained, the curve is free: an O(n) pass over a file, seconds of compute, no calls to anything. If they were not, the curve costs a full rerun — for a 1,000-prompt corpus, roughly 1,000 guard calls plus 1,000 target completions plus judge cost, hours of wall clock on a rate-limited endpoint, and a fresh triage pass. The entire price of the curve is therefore paid before the run, in one decision: log the score, not the verdict. That is why score retention belongs in the harness as a standing requirement rather than a per-engagement choice. ## Where the curve misleads - **It argues monotonically for tightening, and it has no benign axis.** Every point on the curve gets better as you tighten, because the corpus is all attacks. The cost of tightening lands on legitimate traffic that this run never measured. A reader shown only this curve will reasonably conclude the threshold should go up; nothing on the chart tells them what that refuses. - **A flat band is ambiguous.** Wide separation is equally consistent with a strong guard and with a mild corpus. Before claiming margin, check the same curve against a harder slice, or against a second guard — if your attacks sit far from *every* guard's line, the finding is about your corpus. - **A moving judge fakes structure.** If the judge is a model re-invoked at each threshold, sampling noise appears in the curve as shape. Judge once per attempt and reuse the column. - **One dial, several knobs.** A guard that scores per category has several thresholds. Sweeping a global cut while per-category rules also apply collapses knobs the reader will assume were held fixed; say which dial you actually moved. ## What to check Validate the sweep against reality before shipping it: pick two thresholds off the curve, actually configure the guard at those values, rerun a slice of the corpus, and confirm the live verdicts match what the sweep predicted. This tests an assumption the recomputation quietly makes — that the deployed threshold is compared against the same score you recorded. Services that layer extra logic on top of the score (keyword short-circuits, per-category overrides, a separate policy engine) break that assumption, and the swept curve will be smoothly wrong. Then quote the number at the operating point actually in force, with the curve beside it: the reader's reality is one point, and your job is to stop that point standing alone.

  • Two guards both report 4% bypass at their configured thresholds. What does the sweep tell you that the point does not?
    Whether 4% holds across a band of thresholds or exists only at that exact setting. The second guard's attacks sit right against the line and a small retune moves the number a long way.
  • What must you hold fixed while sweeping the threshold?
    The judgement of whether an attack actually succeeded. The sweep should only move the guard's allow/block decision; if the success criterion also moves, the curve is uninterpretable.
  • Your harness saved only allow/block booleans. Can you still produce the curve?
    No. The sweep needs the retained per-prompt scores, so you have to rerun the corpus with score capture turned on.

saying these in an interview costs you the question

  • Reporting one bypass percentage and calling the guard strong or weak on that basis.
  • Discarding per-prompt scores and keeping only the allow/block verdict.
  • Re-judging attack success at each threshold, so the sweep mixes two moving parts.
  • Claiming the curve shows the right threshold to run at, which is the app team's call, not the tester's.
  • Implying the curve says anything about refused benign traffic when no benign set was run.

context