skip to content

You are writing up over-blocking results for a moderation guard in front of a chat product. What counts as an over-block event, and why is a single aggregate false-positive percentage not enough for the product team to act on?

level: seniorimportance: should knowfreq 45%

answer

  1. refusal, hedge, truncation, silent rewrite
  2. score the response, not the flag
  3. disaggregate by benign category
  4. same guard configuration for both rates
  5. appendix of verbatim refusals

basics

~20 s

Count an over-block whenever a benign request fails to get the answer it deserved: a hard refusal, a hedged non-answer, or a silently rewritten reply. One overall percentage hides which kinds of user are hit, so break it out by benign category and hand over the actual blocked requests as examples.

solid answer

~50 s

**Event definition first.** Blocking is rarely binary in a deployed product. A request can be refused outright, answered with a generic safety notice that contains no information, truncated, or quietly rewritten into something blander. To the user all four are the product failing; if your harness only counts hard refusals, you will under-report the cost. Decide the rule up front, write it into the report, and apply it consistently — a scorer that reads the final response is usually needed, not just the guard's own verdict flag. **Then disaggregate.** A 2% aggregate can be 0% for most traffic and 30% for one category — clinicians, security researchers, one non-English locale, one age-appropriate creative use case. That concentration is the actionable finding, and averaging destroys it. Report per benign category, and per source where you have it. **Then make it concrete.** Percentages get argued about; five verbatim benign requests the guard refused, with the responses it gave, get fixed. Include them.

go deeper

for a junior

Should recognise that a refusal on a legitimate request is a cost worth reporting, and that examples help.

for a middle

Should broaden the event definition past hard refusals to hedges and rewrites, and ask for a per-category breakdown.

for a senior

Should define the event before the run, score responses rather than the guard's flag, disaggregate, hold the guard configuration constant across both rates, and attach verbatim examples.

for a principal

Should own how the finding lands: naming concentrated over-blocking on a profession or language as a fairness issue, and ensuring the cost of caution reaches the decision-maker who only asked about bypasses.

The report is the deliverable, and over-blocking is the half of it with no natural advocate in the room. The stakeholder came to hear about bypasses, so the cost of caution reaches the decision only if the write-up carries it deliberately. ### Where a block can actually happen A deployed product is a stack, and the guard's own verdict flag sees one layer of it. An input classifier can refuse before the model runs. An output classifier can gut the answer after it. Bedrock Guardrails' `ApplyGuardrail` call reports an `action` of `GUARDRAIL_INTERVENED` and substitutes a configured blocked-message string, so the intervention is visible in the API response while the user sees only the substitute. A NeMo Guardrails dialogue flow can decide to refuse inside Colang, and the refusal returns looking like an ordinary assistant turn. A cautious system prompt produces the same user-visible outcome with no flag anywhere at all. So a harness that counts only what the guard's boolean reports will under-count over-blocking, and will misattribute the instances it does count. Attribute the layer where you can — the fix for an over-eager input classifier is not the fix for a system prompt that hedges. ### Defining the event Write the rule down before the run: a benign item is an over-block if a reasonable user would not consider their request served. That covers the hard refusal, the content-free safety notice, the truncated answer and the silently softened rewrite. Applying that rule means scoring the *response*, not reading a flag — which means a scorer, and a scorer is a classifier with its own error rate landing directly inside your headline number. Calibrate it: hand-label around fifty responses, measure the scorer's agreement with those labels, and report that agreement beside the over-block rate. A judge that counts every hedge as a refusal inflates the rate; a judge that only recognises "I can't help with that" deflates it. Cost it honestly too — an LLM judge over 500 benign responses is 500 more calls plus the human hour to calibrate, on top of the calls that produced the responses. ### How the number misleads **The aggregate is a weighted average whose weights you invented.** A single false-positive percentage over the benign corpus is the mean of per-slice rates weighted by how many items you happened to put in each slice. You chose that mix when you built the corpus, and it is not production's category prevalence. Quote per-slice rates as the primary result; if you want one headline number, say whether it is corpus-weighted or reweighted to production traffic, and show the weights. **Averaging destroys the finding.** 2% overall can be 0% across most traffic and 30% inside one professional slice or one non-English locale. The concentration is the actionable engineering result, and — when it lands on a profession or a language — the fairness result the product owner needs named rather than diluted. **Cross-configuration pairing.** Comparing an over-block rate measured under one guard configuration against a catch rate measured under another is the most common way these write-ups mislead, and it happens by accident whenever the guard is retuned mid-engagement. Stamp a configuration identifier on every run and refuse to pair numbers that do not share it. **Silent drops.** Items that errored, timed out or exceeded a length limit and were skipped leave the numerator, and frequently leave the denominator too. Both directions bias. Count them explicitly and put the count in the methodology. ### What to check before publishing Reconcile four counts: benign items sent, decisions recorded, responses scored, items dropped. They must add up, and the drop count is part of the result rather than an embarrassment to hide. Confirm one configuration identifier across the whole benign run and across the attack run you pair with it. Confirm the scorer's agreement on the hand-labelled sample. Re-derive the aggregate under a second, production-like set of slice weights and see how far it moves; if it moves a lot, the aggregate was never a fact about the guard. ### Making it land Attach the examples. Five verbatim benign requests the product refused, each with the response it gave and the slice it came from, convert a percentage into an obvious bug that someone will fix. Keep the same examples across reruns so the team can see which ones got resolved. And put the concentration, not the average, in the executive summary — the average is the number that lets the finding be ignored.

  • A benign request passes the input classifier but the final answer is generic safety boilerplate. Whose problem is it, and how do you report it?
    Still an over-block from the user's point of view. Report it as one, and attribute the layer — output classifier or cautious system prompt — where you can, since the fix differs from an input-classifier change.
  • Over-blocking is 1% overall but 25% on one professional category. What goes in the summary?
    The 25%. The concentration is the finding; the average is the number that would let it be ignored.

saying these in an interview costs you the question

  • Counting only hard refusals and ignoring content-free safety boilerplate
  • Reporting one aggregate percentage with no breakdown
  • Comparing an over-block rate and a catch rate measured under different guard configurations
  • Dropping benign items that errored during the run without noting the bias
  • Leaving over-blocking out of the executive summary because the stakeholder asked about bypasses

context