Your organisation's harm policy names categories that the deployed guard model's fixed hazard taxonomy does not cover, and the guard emits categories your policy never mentions. How do you scope and report a guardrail test engagement so that the pass rate you hand leadership is not read as evidence that the policy is enforced?
answer
- policy category to guard category mapping table
- two denominators, never one number
- untested is not passing
- gaps get a control, an owner, a date
- remap when the guard model changes
basics
~20 sMap every policy category to a guard category first and mark the unmapped ones unmeasurable at this layer. Report two numbers, never one: prompts blocked out of prompts sent, and how many policy categories this instrument can express at all. Unmapped categories go in the report as untested, with an owner, not as passing.
solid answer
~1 minThe trap is a single headline percentage whose denominator is the guard's category list while the reader's denominator is the policy. **Do the mapping before the engagement.** A three-column table — policy category, corresponding guard category if any, how it will be measured — is the scoping artefact. It usually reveals that a meaningful share of the policy is not testable against this layer at all, which is itself the most valuable finding and often changes the engagement's shape before a prompt is sent. **Report on two axes.** Measured performance (blocked over sent, per category) says how well the layer does its job on what it can see. Policy expressibility (policy categories the instrument can represent, over policy categories in force) says how much of the policy the layer could ever cover. Presenting only the first is how a green number gets attached to an untested policy. **Name owners for the gaps.** Each unmapped policy category needs a proposed control — a second classifier, a rules layer, retrieval-side restrictions, human review — and a named owner, or the gap will reappear identically next quarter. **Re-do the mapping when the guard model changes**, since categories and behaviour belong to a specific model, not to a product name.
go deeper
Recognises that the guard's categories and the company policy are different lists and should not be conflated.
Builds the mapping table and reports untested categories separately from failed ones.
Scopes the engagement from the policy, reports per-category counts with both denominators, and pairs the block rate with an over-block rate.
Owns the framing so the number cannot be misread downstream, converts each gap into a proposed control with an owner and cost or a dated acceptance, and decides what becomes continuous monitoring versus a periodic engagement.
This is a scoping and reporting problem wearing a testing problem's clothes, and it is where guardrail engagements most often mislead the people who commission them. ## The two lists, and why they never match Your organisation's harm policy is a list of things the business has decided it will not do or allow: regulated-advice boundaries, contractual and brand constraints, sector-specific abuse patterns, plus the universal harms. The guard model's hazard taxonomy is a different list, fixed at that model's training time and published by its authors, built to be broadly useful across customers rather than correct for you. The overlap is partial in both directions. The guard emits categories your policy never mentions, and — the direction that matters — your policy names categories the guard has no way to express. The instrument cannot produce a signal for those, so no threshold, no prompt engineering and no amount of test volume will ever measure them at this layer. ## Scope from the policy, not from the tool The scoping artefact is a three-column table built before a single prompt is sent: policy category, corresponding guard category if any, and how it will be measured. Record the mapping quality per row, because there are three states and not two: - **Clean map** — a guard category corresponds to the policy category; measure it directly. - **Partial map** — the guard has an adjacent category that catches part of the intent. These are the rows that quietly turn into overclaims later, because they produce a real number that covers less than the reader assumes. Say what part is covered. - **No map** — unmeasurable at this layer. This row is the most valuable output of the whole exercise, and it usually arrives before any testing money is spent. ## What it costs The mapping is engineer and policy-owner hours, typically a day or two of joint work, and it routinely changes the engagement's shape before a prompt is sent — discovering that five of twelve policy categories are unmeasurable here is worth more than a week of probing the other seven. The measured half then costs what any model-guard run costs: roughly three inferences per case with two-sided screening, multiplied by the per-category sample size you need for the rate to mean anything. A category tested with six prompts does not support a percentage; decide the minimum per-category sample up front and report counts, not just rates, so a reader can see which cells are thin. ## Where the number misleads — the denominator swap A single headline percentage is computed over the categories the instrument can express, and read as if it were computed over the policy. That one substitution converts every unmeasured category into an implied pass. Write the number so the swap is impossible: > "The guard blocked 86% of the 420 prompts sent, across the 7 of 12 policy categories it can express. The remaining 5 policy categories are not testable at this layer and are untested." A bare "86% blocked" is not a summary of that sentence; it is a different claim. Two related traps. **Untested and passing look identical in a table of green cells** — give untested its own visual state and assume the slide will travel without your caveats attached. And **a block rate quoted without an over-block rate** invites a leadership decision to tighten the guard, with no visibility into the product incident that follows; the two numbers belong in the same row. ## Turn gaps into a plan, with names on it Each unmapped policy category gets a proposed control — a second or bespoke classifier, a rules layer, retrieval-side restriction, human review — with a rough cost and a named owner. Some gaps are correctly accepted; the honest outcome is a recorded, dated risk acceptance, not silence, because silence converts an accepted risk into an assumed-covered one within a quarter. ## What you would check before publishing That both denominators appear wherever the number appears, including in the executive summary. That per-category counts are visible under the headline and thin cells are marked. That the benign over-block rate sits beside the block rate. That every gap row has a control, an owner and a date, or a signed acceptance. That the mapping is pinned to the specific guard model and serving version, and is re-done — with controls re-run — before last quarter's numbers are ever quoted again after a model swap. And that the words "pass" and "compliant" appear nowhere against a category the instrument cannot express.
- Leadership wants one number for the dashboard. What do you give them?Two, tightly coupled: share of policy categories this layer can measure at all, and the block rate within them, with the over-block rate beside it. If forced to one, give the policy-coverage share, since it bounds the meaning of everything else.
- A policy category has no control anywhere and no budget. What goes in the report?A recorded, dated risk acceptance naming the owner, plus the cheapest control you would build if funded. Silence turns an accepted risk into an assumed-covered one.
It is the difference between a school reporting that 86% of the students it tested passed and a school reporting that 86% of its students passed. The classifier chose which subjects were examined, and the five it never sat show up on the transcript as blank, not as pass.
saying these in an interview costs you the question
- One headline block rate with no stated denominator.
- Green cells for policy categories the instrument cannot express.
- Scoping the engagement from the tool's category list instead of the policy.
- Reporting block rate without the over-blocking rate, inviting a change that breaks the product.
- Quoting last quarter's numbers after the guard model was swapped.