You test a moderation guard that returns a per-prompt score, and most of your attack prompts that were blocked scored only barely above the block threshold. What does that clustering tell you, and what goes in the report?
answer
- blocked, but with no margin
- share within a narrow band above the line
- config drift, guard drift, paraphrase drift
- whole family at the line = one technique
- sensitivity table, not one number
basics
~20 sIt means the guard barely separates those attacks from allowed traffic: the blocks are luck of the line, not margin. A small loosening, a guard update, or a slightly reworded attack flips many of them through. Report the margin distribution and how bypass moves for small threshold changes, not just the current number.
solid answer
~60 sA block with no margin is a block you cannot rely on. **What it tells you.** The guard's score distribution for your attacks overlaps the threshold rather than sitting well above it. The current bypass number is therefore a property of exactly where the line sits, and three ordinary events undo it: the owner loosens the setting, the guard is updated and its scores shift, or an attacker paraphrases a prompt down a few points. **What goes in the report.** Not one number but three things: - the **margin distribution** — what fraction of blocked attack prompts sit within a narrow band above the threshold; - a **sensitivity line** — bypass at the operating point and at small steps either side, so the reader sees how far the number travels; - **which attack families** are clustered there, since a whole family sitting at the line is a single technique the guard nearly misses, not scattered noise. **What not to claim.** Tightening the threshold would recover those blocks, but the cost lands on benign traffic that this all-attack corpus never measured. Say the trade exists and that it needs its own measurement; do not price it from this run.
go deeper
Should notice that scraping past the threshold is weaker evidence of safety than blocking with a wide margin.
Quantifies the margin as a share within a band and shows bypass at a step either side of the operating point.
Names config drift, guard drift and paraphrase drift as the routes from near-miss to bypass, checks whether the cluster is one attack family, and refuses to price a tighter threshold from an all-attack corpus.
Turns margin reporting into a standing metric so guard fragility is tracked over time and a threshold change triggers re-measurement rather than an argument.
## What the clustering means Keeping raw scores instead of verdicts buys you exactly this diagnostic. The headline figure tells you how many attacks got through today; the **margin distribution** — how far the blocked prompts sat from the block threshold — tells you how much to believe that figure tomorrow. When most blocked attack prompts scored only just above the line, the guard's score distribution for your attacks *overlaps* the threshold rather than sitting clear of it. The guard is not confidently distinguishing those attacks; the line happens to be drawn just under them. A block with no margin is a block you cannot rely on. Summarise it as one quantity: of the attack prompts the guard blocked, what share scored within a narrow band above the threshold? A small share means confident separation and a stable bypass rate. A large share means a knife-edge result. ## The three routes from near-miss to bypass Name all three in the report, because all three are routine rather than hypothetical. 1. **Config drift.** Whoever owns the dial loosens it — usually in response to over-blocking complaints from the product side — and a band of your blocked attacks becomes allowed in one edit. 2. **Guard drift.** The guard is retrained, or the hosted model behind it is swapped, and the score distribution shifts wholesale. The same threshold value now means something different, and nobody is notified. 3. **Attacker drift.** A rewording, a formatting change, or an innocuous-looking wrapper shifts a prompt a few points down the score. Where a whole attack family sits at the line, that is close to free for an adversary — which is precisely why the finding is worth writing up as a technique-level gap rather than a statistic. ## What it costs to measure Quantifying the margin costs nothing extra if the run retained per-prompt scores: it is a histogram over a column you already have. The expensive part is the follow-up — confirming the near-misses actually flip. That means running a rewording pass over the clustered family and re-measuring, so budget roughly *k* variants × family size × (one guard call + one completion + one judgement) and note that this pass has to be sampled, not exhaustive, on a metered endpoint. Report the fragility you measured; the report's job is the finding, not a catalogue of the wordings that worked. ## Where the number misleads - **The band width is a free parameter.** Choose it after looking at the histogram and you can manufacture any margin share you like. Fix it in advance — as a stated fraction of the score range, or in whatever units the guard documents — and print the width next to the share. - **Scores are not probabilities.** A guard's 0.62 is not "2% more harmful" than its 0.60, and two guards' score scales are unrelated. Margin shares are therefore comparable *within* one guard over time and **not** across guards; a table putting two vendors' margin shares side by side is comparing rulers with unmarked units. - **Quantisation fakes clustering.** Some guards round scores to two decimals or return bucketed severity bands. On those, everything looks piled at the line because the resolution is coarse, not because the guard is weak. Check the number of distinct score values before drawing a conclusion. - **A blocked prompt is not a safe prompt.** Clustering is often read as good news — "we blocked them all" — which inverts the finding. The correct reading is that the blocks were decided by the line's position, not by the guard's discrimination. - **Tightening cannot be priced from this run.** An all-attack corpus has no benign side, so it cannot say what a stricter setting refuses. Flag the trade and say it needs its own measurement; do not recommend a specific threshold from these data. ## What to check Look at whether the clustered prompts share a technique or a harm category — a whole family at the line is one concrete finding an engineer can act on, while scattered near-misses are a general fragility statement. Check the guard's per-category scores on those prompts: often one category fires strongly and the one that actually matters barely moves, which localises the gap. Run the same prompts against a second guard; if they sit at *its* line too, the honest reading is that your corpus is mild rather than that this guard is weak. Finally, check the benign set's margin distribution as well — if legitimate traffic is also piled just under the line, the guard has no separation anywhere and this family is a symptom rather than the finding. ## How to report it Not one number, but three: bypass at the operating point, bypass one step either side of it, and the margin share with its band width stated. The sensitivity steps show how far the figure travels; the margin share explains why; and both are readable by someone who will never open a score histogram. That form also pre-empts the conversation six months later in which the percentage is quoted as though it were a constant.
- How would you summarise the margin in one number for an executive summary?The share of blocked attack prompts scoring within a narrow band above the threshold — a high share means the bypass figure is one small change away from being much worse.
- The clustered prompts all belong to one attack family. What changes in your recommendation?It becomes a specific finding about one technique the guard nearly misses, which can be addressed directly, rather than a general call to tighten the threshold for everything.
- Why not just recommend lowering the threshold to catch them?Because the cost lands on benign traffic that an all-attack corpus never measured. Flag the trade and say it needs its own measurement before anyone moves the knob.
Blocks with no margin are like a student who passes every exam by a single mark: the transcript says pass every time, but you would not bet on next term's.
saying these in an interview costs you the question
- Reading the clustering as good news because the prompts were blocked.
- Reporting only the bypass number when the margin distribution shows it is knife-edge.
- Recommending a specific tighter threshold with no benign measurement to price the cost.
- Not checking whether the clustered prompts are one attack family.
- Assuming scores keep the same meaning after the guard is updated.