skip to content

Your guardrail test corpus is half attack items and half benign items, but production traffic is overwhelmingly benign. How do you choose the corpus mix, and how do you report the results so the numbers still say something about production?

level: principalimportance: should knowfreq 33%

answer

  1. oversample the rare class deliberately
  2. per-class rates transfer, mixed ones do not
  3. project onto real daily volume
  4. state the volume assumption
  5. slice by category, language, surface

basics

~20 s

Keep the balanced mix for measurement — you need enough attack items to estimate the block rate at all — but never quote a corpus-level precision as if it were production's. Report the two rates separately, then project the daily wrongly blocked volume using production's real benign volume and its base rate.

solid answer

~50 s

The corpus mix is a sampling decision, not a model of production. Balanced sampling is the efficient choice: attacks are rare in traffic, so a production-proportional corpus would need enormous volume to contain enough of them to estimate anything, and most of your labelling budget would buy easy benign items. What does not survive the reweighting is any number that mixes the two classes. The share of attack items caught and the share of benign items wrongly blocked are each estimated within their own class, so both transfer. Anything of the form "what fraction of blocks were real attacks" depends on how many of each class you fed in, and on a balanced corpus that number is meaningless for production. So report the two per-class rates plus the counts behind them, then translate: multiply the wrong-block rate by real benign volume to get blocked legitimate requests per day, and the miss rate by the expected attack volume. Those two operational quantities are what a decision maker can weigh.

go deeper

for a junior

Knows the corpus mix is chosen, not sampled from production, and that both rates must be reported.

for a middle

Explains that per-class rates transfer while mixed-class figures depend on the ratio you supplied.

for a senior

Projects the rates onto real volumes with stated assumptions and reports per-slice rather than one aggregate.

for a principal

Frames the threshold as an explicit business trade owned by the product, publishes the curves and assumptions that let it be made, and keeps provenance and contamination in the same report so the projection can be trusted.

**Why balanced, and what balancing costs you.** Attack traffic against a live endpoint is a thin slice of volume — often well under one percent. Sample a corpus in proportion and nearly the whole hand-labelling budget buys ordinary requests, while the attack half stays too small to distinguish a good block rate from a bad one. Oversampling the rare class is the standard and correct answer. The price is that the corpus no longer resembles production in composition, and that mismatch has to be handled in the reporting instead of quietly ignored. **Which quantities survive the reweighting, and which do not.** A rate estimated *within* one class does not depend on how many of the other class you supplied. The share of attack items blocked is computed over attack items only; the share of benign items wrongly blocked is computed over benign items only. Both transfer. Any quantity that mixes the classes — above all "of everything the guard blocked, what fraction were genuinely attacks" — is a function of the mix you chose, and it moves violently with the base rate. Work it through with a guard that blocks 80% of attacks and 3% of benign items: | corpus | attacks | benign | true blocks | wrong blocks | share of blocks that were attacks | |---|---|---|---|---|---| | 50/50 test corpus | 500 | 500 | 400 | 15 | ~96% | | one production day | 50 | 10,000 | 40 | 300 | ~12% | The guard did not change between the rows. Only the denominator under its positives did. Quoting the 96% to a decision maker is the classic error in this leaf, and it is how a threshold nobody would have accepted gets shipped. **How to report so the numbers say something about production.** Per slice, carry: attack-item count and share blocked; benign-item count and share wrongly blocked; the provenance of each half; and the guard version and threshold. Then add a projection section that applies the two within-class rates to real volumes — at 10,000 benign requests and 50 attack attempts a day, a 3% wrong-block rate is about 300 blocked legitimate requests daily and a 20% miss rate is about 10 attacks through. State the assumed attack volume in the open, because it is the softest number in the report and the one a reader should be free to argue with. **The transfer condition people skip.** Within-class rates transfer *only if that class's composition resembles production's*. The benign half is usually built deliberately heavy on near-miss items, precisely so it discriminates — which means its wrong-block rate is a near-miss rate, not a traffic rate, and projecting it directly onto total benign volume overstates the cost, sometimes by an order of magnitude. The fix is to stratify: keep easy-benign and near-miss-benign as separate strata with separate rates, estimate each stratum's real share of traffic, and reweight. If you cannot estimate the shares, say so and report the stratum rates unreweighted rather than presenting a blended figure that belongs to no population. **What it costs.** The corpus size sets both bills. Labelling is the large one — hundreds of items, minutes each, with a double-labelled sample on top — and slicing multiplies it, because a per-language or per-category rate needs enough items *per cell* to be worth quoting; forty items in a slice gives a confidence interval wide enough to swallow most decisions. Inference is the small one for a hosted guard (cents to a few dollars, minutes of wall clock for a full sweep) and a real one for a self-hosted guard on GPU or an in-situ run that pays application-model tokens per item. Decide the slice list before labelling, since it determines how many items you must buy. **Where the numbers mislead.** Beyond the precision trap: a single blended accuracy over a mix you composed is an average over your own sampling decision and means nothing outside it. A rate quoted without its item count invites a percentage resting on three items. Two guard candidates with equal aggregate scores can differ wildly per language or per category, and the aggregate hides it. And a flattering catch rate measured on a contaminated attack half poisons the projection downstream, so provenance belongs in the same report as the volumes. **What to check.** Recompute the projection yourself from counts, not from percentages, and see whether the arithmetic reproduces the headline. Confirm both rates came from one run at one threshold on one pinned guard version. Ask what the benign half is made of and whether its strata were reweighted. Put a confidence interval, or at minimum the item count, on every rate. Finally, make sure the threshold decision is framed as an explicit trade — blocked legitimate users against prevented incidents — and put it in front of whoever owns the product, rather than letting it be settled implicitly by whoever picked the vendor's default.

  • Two guardrail candidates have identical per-class rates on your corpus. What else decides between them?
    Per-slice behaviour and cost: which categories, languages and surfaces each fails on, latency and price per call, and how each degrades on paraphrased items whose provenance you control.
  • A stakeholder asks for a single number. What do you give them?
    A projected operational quantity with its assumption attached — for example, expected wrongly blocked legitimate requests per day at the proposed threshold — rather than a corpus-level accuracy.

The same test looks brilliant in a lab where half the samples are infected and nearly worthless in a clinic where one in a thousand is. The test never changed; only the denominator underneath its positives did.

saying these in an interview costs you the question

  • Quoting a corpus-level share-of-blocks-that-were-attacks as a production figure.
  • Sampling the corpus proportionally and ending up with too few attack items to measure anything.
  • Reporting one aggregate score with no per-class counts.
  • Choosing the operating threshold alone, without the product owner seeing the trade.
  • Projecting onto production volume without stating the assumed attack volume.

context