Your team quotes a bypass rate for the same input-moderation guard every release, but the attack corpus grows each quarter as new templates are added. Which denominator do you standardise on so two quarters' numbers mean the same thing, and how do you handle the newly added templates?
answer
- frozen versioned baseline slice
- distinct attacks at a fixed budget
- new templates reported separately
- held-out slice, rotate on a schedule
- re-score old transcripts after a decider change
basics
~20 sSplit the corpus. Freeze a versioned baseline slice and report the bypass rate over that fixed set of distinct attacks at a fixed attempt budget - that number alone is compared quarter to quarter. Report newly added templates as a separate first-run number, never blended in. Blending makes the trend move when the corpus changes rather than the guard.
solid answer
~60 s**Standardise on distinct attacks from a frozen, versioned slice, at a fixed attempt budget.** A per-attempt rate over a growing corpus is uninterpretable: it moves whenever the sampling mix or retry policy changes, which is every quarter. **Three numbers, never one.** 1. *Baseline slice* - the frozen set, same templates, same budget, same decider version. This is the trend line, and the only quarter-to-quarter comparison anyone should draw. 2. *New this quarter* - first-contact results for the added templates, reported as counts, since they have no prior to compare against. 3. *Merged corpus* - the current best estimate of overall exposure, reported as a level and explicitly not as a trend. **The trap in freezing.** A fixed slice the guard's owners can see becomes something to tune against, and the trend flattens while real exposure does not. Mitigate by holding part of the baseline back from guard owners and rotating a portion each year, announcing the rotation so the discontinuity in the trend is expected rather than argued about.
go deeper
Should notice that adding templates each quarter makes the numbers incomparable and ask what stayed the same.
Should propose a frozen baseline set and a fixed attempt budget, with new templates reported separately.
Should version the decider and target settings too, keep raw transcripts, and re-score history when the decider changes.
Should treat the denominator as a governance lever: written measurement definition, held-out slice, scheduled rotation with an overlap period, and clear limits on what the number may be used to claim.
### This is a governance problem wearing a statistics costume The moment a bypass rate becomes a release gate or a quarterly slide, the denominator becomes a lever. Anyone who adds templates, changes the retry budget, upgrades the judge that labels a hit, or nudges the target's decoding settings can move the headline without touching the guard - and will usually do so for a perfectly innocent reason. A lead's job is not to win an argument about this quarter's number but to remove the lever, by fixing the measurement definition in writing and versioning everything that feeds it. ### Standardise on distinct attacks from a frozen slice at a fixed budget A per-attempt rate over a growing corpus is uninterpretable as a trend: its value depends on the mix of templates and on how many times each was retried, and both change every quarter. The comparable statistic is the fraction of a **frozen, versioned set of distinct attack templates** that bypassed at least once **within a fixed attempt budget N**, scored by a **named decider version**. Publish three numbers and never fewer: 1. **Baseline slice** - the frozen set, same templates, same N, same decider. This is the only series anyone may draw a trend through. 2. **New this quarter** - first-contact results for newly added templates, reported as counts. They have no prior, so they are a level, not a trend point. 3. **Merged corpus** - the current best estimate of overall exposure against everything you own, reported explicitly as a level. Blending 2 into 1 is the specific mistake: the line then moves when the corpus changes, and a corpus event gets read as a security event in both directions - a quarter of easy new templates looks like a regression, a quarter of hard ones looks like a win. ### What must be versioned, and what that costs Version the template set (with a stable id and technique-family label per template), the attempt budget N, the decider and its version, and the target's sampling settings. Change one at a time. When a decider upgrade lands, **re-score the previous quarters' stored transcripts with the new decider and republish the series**, so any step change is attributable to the relabelling rather than argued about. That requires keeping raw transcripts - request, guard verdict, model reply, decider output - which is the real bill this discipline incurs: on the order of a few hundred megabytes to a few gigabytes per quarter for a corpus of hundreds of templates at a double-digit budget, plus the re-scoring calls, which is one judge call per stored transcript for every prior quarter you republish. It is worth arguing for, because without stored transcripts every methodology change orphans the history and the series restarts. ### Countering overfit to the frozen set A fixed set that the guard's owners can see becomes something to optimise against, and there is no clean line between fixing a real hole and pattern-matching the test. The trend flattens while real exposure does not. Split the baseline: one slice the owners may iterate against freely, one held-out slice reported only in aggregate and never handed over template-by-template. Rotate a fraction of the held-out set on an announced schedule, and during the rotation run old and new slices for one overlap period so the offset between them is measured rather than guessed. That is the ordinary discipline of any metric redefinition, and it is what turns an unexplained jump into a published constant. ### Where the number misleads The two live traps: a **flat trend on a visible frozen set** reads as a stable guard when it may mean the set has been learned; and a **jump after a corpus or decider change** reads as a security event when it is a measurement event. Both are fixed by the same habit - show corpus version, N, decider version and set size in the same view as the rate, so a reader cannot see the line without seeing what defines it. And the boundary matters: a guard's bypass rate on an attack-only corpus is an instrument reading about one control. It is not production exposure, not the engagement's attack-success rate, and not a ship decision; those are separate reports with separate owners, and conflating them is how a sound measurement gets used badly. ### What to check each quarter That the baseline slice is byte-identical to last quarter's; that N and the decider version are unchanged, or that history was re-scored if they are not; that new templates are reported outside the trend; that the held-out slice is still held out; and that the counts, not just the percentages, are in the report so anyone can re-derive the number.
- You must rotate part of the frozen baseline. How do you avoid an unexplained jump in the trend?Run old and new slices in parallel for one reporting period, measure the offset between them, and publish it alongside the changed series.
- The decider that labels a hit is upgraded. What do you owe the historical series?Re-score the stored transcripts of prior runs with the new decider and republish the series, so any step change is attributable to the labelling change rather than the guard.
- Engineering asks for one headline number instead of three. What do you give them?The baseline-slice distinct-attack rate at the fixed budget, with counts and the corpus version. The other two stay in the report body so nobody quotes a level as a trend.
saying these in an interview costs you the question
- Blending newly added templates into the trend line and calling a corpus change a security improvement or regression.
- Letting the guard's owners tune against the whole baseline set with no held-out portion.
- Upgrading the decider mid-series without re-scoring history.
- Reporting a per-attempt rate over a corpus whose composition changes every quarter.
- Presenting the guard's bypass rate as though it were the product's production exposure or a ship decision.