A guardrail assessment has a fixed number of paid calls to a hosted moderation endpoint. How do you decide what share of that query budget goes to benign hard negatives instead of attack probes, and how do you defend the split to a stakeholder who wants everything spent on attacks?
answer
- size from the smallest rate worth acting on
- stratify benign across near-miss slices
- attack half: distinct techniques, not variants
- reserve capacity for a re-measure
- catch-rate-only is unfalsifiable one way
basics
~20 sSpend enough on benign traffic to detect an over-block rate the product would actually care about, then give the rest to attacks. Size the benign half from the smallest rate worth acting on, not from what is left over. Defend it by noting a catch-rate-only report cannot tell a good guard from a closed door.
solid answer
~60 sWork backwards from the decision, not from the corpus. The benign half exists to answer one question: at this configuration, does the guard refuse legitimate requests often enough to matter? So size it from the smallest over-block rate the product would act on — if a 1-in-100 refusal on clinician traffic would trigger a fix, you need enough benign items in that slice to see it at all, and enough to distinguish it from zero. That reasoning also tells you the shape: a modest number of items spread across many near-miss categories beats a large undifferentiated pile, because the finding you are hunting is concentration. The attack half has diminishing returns of a different kind — after the first few dozen distinct techniques, extra calls mostly buy variants of things you already found. Defending the split is a framing argument: the deliverable is a tuning decision, and a report with only one of the two costs can only ever argue in one direction. You are not spending budget on benign traffic instead of finding bypasses; you are spending it so the bypass number can be safely acted on.
go deeper
Should recognise that some of the budget has to go to benign traffic or the result is one-sided.
Should reason about diminishing returns on repeated attack variants and reserve some calls for benign items.
Should size the benign half from the smallest actionable over-block rate, stratify it across near-miss slices, and hold capacity back for a re-measure after any guard change.
Should own the stakeholder conversation — framing the split as what makes the bypass number safely actionable — and write the split, its rationale and the smallest detectable rate into the methodology.
This is a portfolio decision under a hard constraint, and the honest version of it is stated before the engagement starts rather than discovered when the budget runs out. ### Size the benign half from what "zero" would mean The arithmetic that settles the argument is the rule of three: if a slice of n benign items returns zero over-blocks, the upper 95% bound on the true rate for that slice is about 3/n. A hundred items returning zero is consistent with a true over-block rate as high as 3%. So if the product would act on a 1% refusal rate for clinician traffic, a 100-item clinical slice returning zero has failed to rule out three times the rate you care about; you need roughly 300 items in that slice before "zero" means "below 1%". Two consequences fall straight out. First, the arithmetic applies **per slice**, not to the pile, so stratification is forced — an undifferentiated benign heap cannot bound anything for the slice where the concentration lives. Second, full coverage is usually unaffordable: ten slices at 300 items is 3,000 benign items, and labelling those is more than a week of scarce reviewer time. So narrow the claim rather than fake the coverage. Fund the three or four slices where over-blocking would actually change what the product does at 250–300 items each, sample the rest thinly, and report the bound the thin samples support. "Clinical, security-research and non-English slices measured at 300 each; three further slices sampled at 60, bounding over-blocking below roughly 5%" is an honest sentence. "Over-blocking: 0%" from 60 items is not. ### Size the attack half from marginal information Distinct techniques and distinct surfaces dominate the value of the attack half. The tenth variant of an already-confirmed bypass adds report volume, not decision value — the decision it would inform was already made when the first one landed. Once a technique is confirmed to work, further calls proving it works again are the cheapest thing in the plan to cut. ### Reserve for the re-measure Any guard change during the engagement invalidates cross-configuration pairing. Hold back on the order of 20% of the budget so both halves can be re-run against the final configuration. Teams that spend everything on configuration A end up quoting two numbers from two different guards, which is worse than quoting one. ### The non-query budget Calls are usually the cheap input. Human labelling, the privacy review that lets sampled production traffic into the corpus, and the calibration of any response scorer are the scarce ones, and they have lead times measured in days rather than minutes. Book them at kickoff; leftover capacity never materialises. ### The stakeholder conversation The person who wants everything spent on attacks is optimising for finding something, which is a legitimate instinct rather than an obstacle. Two arguments land. The mechanics: an all-attack corpus is maximised by a guard that blocks everything, so a catch-rate-only report cannot distinguish a well-tuned guard from a closed door. The number is not merely incomplete — it is unfalsifiable in one direction, because no change that increases refusals can lower it. The consequence: this report will be used to tune the guard. If bypasses are the only measured cost, tightening looks free right up to the point where it reaches users, and by then the tightening is attributed to your engagement. Then offer the concrete trade instead of the principle — a bounded benign sample, named in calls, in exchange for cutting redundant variants of confirmed bypasses. That turns the conversation from "is benign testing worth doing" into "which calls", which is a conversation you can win. ### How the number misleads The most dangerous number this leaf can produce is a zero over-block rate from an under-powered benign sample. It reads as "no cost". It means "below what this sample could see", and the rule of three says exactly what that ceiling is, so report the ceiling. Second, quoting the benign share as a fraction of the corpus — "30% of our items were benign" — tells a reader nothing about detectability. Report items per slice and the minimum detectable rate per slice instead. Third, if the stakeholder wins and no benign testing is funded, the catch rate must ship with its one-sided caveat in the summary, not in a footnote. Running the attacks and quoting the rate unqualified is worse than not measuring at all, because an unqualified number gets acted on. ### What you write down Put the split, its rationale, the per-slice item counts and the minimum detectable over-block rate into the methodology before the first call is spent. It converts a judgment call into something a reader can check, it protects the run from being re-scoped mid-flight, and it makes the reserved capacity visible so nobody quietly spends it.
- The guard is retuned halfway through the engagement. What does that do to your budget plan?It splits the run in two. You need reserved capacity to re-measure both halves against the final configuration, or you end up quoting numbers from two different guards.
- Where do you cut first when the budget is short?Repeat variants of bypasses already confirmed. They add report volume without changing any decision, whereas cutting the benign half changes what the report can conclude.
- The stakeholder still refuses to fund benign testing. What do you do?Run what they fund, but state plainly in the report that no over-block cost was measured and that the catch rate must not be used to justify tightening the guard.
Seeing no over-blocks in a hundred benign requests is like seeing no rain in a single week and calling the region arid: the sample is perfectly consistent with a rate three times the one you would have acted on.
saying these in an interview costs you the question
- Treating benign testing as optional leftover capacity
- One large undifferentiated benign pile with no stratification
- Spending the whole budget on the first guard configuration with nothing held back for a re-measure
- Padding the attack half with repeat variants of a bypass already confirmed
- Conceding an all-attacks split to the stakeholder and reporting the catch rate unqualified anyway