You have to report guardrail coverage for an application protected by a rail stack that combines deterministic rules, similarity-matched routing and a judge-model self-check. The three layers share no denominator. How do you report coverage without inventing one number that misleads?
answer
- three units, three rows, no blend
- rules exercised is not safety
- routing space has no denominator
- name the uncovered set explicitly
- validity window: which models, which config
basics
~20 sReport three numbers with their own denominators and refuse the single one. Rules exercised out of rules configured is countable. Routing is a continuous space with no enumerable denominator, so report a fall-through rate over paraphrase families. The judge layer is a pass rate at a fixed repeat count against a named judge model.
solid answer
~60 sThe three layers are measured in different units and a combined percentage silently picks a unit for you. **Deterministic rules** have a real denominator: the configured rule set. Rules exercised over rules configured is honest, and it is also the least interesting number, because it says nothing about payloads nobody wrote a rule for. **Similarity routing** has no enumerable denominator at all — the space of phrasings is not finite. The reportable quantity is a rate over deliberately constructed paraphrase families, plus the fact that the unmatched path exists and what it does. **The judge layer** is measured as a pass rate at a fixed repeat count, valid for one judge model. So the report has three rows with three units, an explicit statement of what is not in any denominator (harms nobody encoded), and a validity window tied to the configuration identity: rule set, example set and cutoff, embedding model, judge model. The value you add as a lead is refusing the one number, because a single figure will be read as 'we are that safe' by people who will never read the rows.
go deeper
Recognises the layers are counted differently and that mixing them into one percentage is not meaningful.
Gives each layer its own denominator and states which one is enumerable and which is not.
Adds the named uncovered set and the validity window tied to configuration and model identity, and re-measures when any of those change.
Refuses the single number to the audience that wants it, offers a ranked ungoverned-surface headline instead, and argues the structural fix on the unbounded layer.
### There is no single right answer here, but there is a wrong one The wrong one is a blended percentage. Aggregation requires a shared unit, and these three layers do not have one. **Deterministic rules** have a genuine denominator: the flows and rules actually written in the deployment's `config.yml` and Colang files. *Rules exercised over rules configured* is countable and honest. It is also the least interesting figure in the report, because it says nothing whatsoever about payloads for which nobody wrote a rule. **Similarity routing** has no enumerable denominator at all. The set of phrasings that could express an intent is unbounded; the framework decides by embedding the turn, taking the nearest canonical form from the `define user` blocks, and applying `embeddings_only_similarity_threshold`. What you can report is a fall-through rate over deliberately constructed paraphrase families, plus the existence and behaviour of the unmatched path. **The judge-backed self-check** is a rate: k of n replays approved, for one payload family, against one named judge model and one prompt template. A count over a finite authored set, a sample from an unbounded space, and a repeat-rate. Averaging them picks a unit silently, and the unit it picks is whichever one has the most items. ### The failure mode that proves the point A blended coverage figure can be moved by an edit that changes nothing about the application's exposure. Add ten flows to `config.yml` that no test touches: the exercised count is unchanged, the denominator grows, and "coverage" falls. Delete an unused rule and coverage rises. **Any metric a no-op edit can move will eventually be moved by one**, usually in the week before a review, and usually by someone acting in good faith. That is the argument to make out loud, because it does not depend on anyone agreeing with your taste in reporting. ### What it costs to produce each number Worth saying in the report, because the three are not comparable work. The rules-exercised figure is nearly free — it falls out of a run you were doing anyway, from the trace's activated-rails list. The judge rate costs 2n model calls per payload family across its life (n to measure, n to re-verify after a fix), which at n=30 with judge and downstream generation is hundreds of calls and an hour of wall clock per family. The routing rate is the most expensive of all and most of the cost is human: authoring paraphrase families that are genuinely meaning-preserving, then re-running them every time the example set, threshold or embedding model moves. A report that quotes all three as if they cost the same invites a request to "just refresh the numbers weekly" that nobody has budgeted. ### The structure that survives contact with an audience - **Per layer**: the number, its denominator stated in words, and how it was measured. - **A named uncovered set**: harms with no rule and no canonical example, and the unmatched routing path. This row deliberately has no number — that is what stops a reader summing the others and calling the remainder small. - **A validity window**: rule-set version, example set and similarity threshold, embedding model, judge model and prompt template. Change any and the numbers expire. - **One qualitative verdict in a sentence**, so the audience that wants a headline gets one you wrote rather than one they derived. ### Where the number misleads, and the organisational pressure Someone will ask for the single figure anyway, usually for a slide, and "82% guardrail coverage" will be read as "we are 82% safe" by people who will never see the rows. The productive counter is to offer a different headline: the ranked list of what is ungoverned. It is decision-relevant and it does not degrade into a safety score. If you must give a number, give the one with a real denominator — rules exercised — and label it plainly an *exercise* metric. Two further points separate a strong answer. First, **the layers run in series, so coverage is not additive in the direction people assume**: a payload class's exposure is set by whichever layer it can walk past, not by the mean of the three, and averaging lets a well-tested input rail mask the routing hole that is actually being used. Second, the routing layer's coverage argument is **structurally unbounded** — you cannot enumerate a continuous space, so no number of added examples ever closes it. That is the real argument for making the unmatched path deny by default: it is the only change that converts an unbounded coverage claim into a bounded one. ### What you would check Before publishing any of the three figures, confirm the denominators are what you think: read the flows actually loaded from the run log rather than from the file you edited, confirm which threshold and embedding model were live, and confirm the judge identity from the run rather than the config. Then re-read your own summary sentence and ask whether a reader who sees only that sentence would be misled. If they would, the sentence is the deliverable to fix, not the table.
- Leadership insists on one number for a slide. What do you offer?A ranked list of what is ungoverned, plus at most the rules-exercised figure labelled as an exercise metric. A headline you wrote beats one they derive from your rows.
- Why can a coverage percentage fall without safety changing?Because someone added rules to the configured set. The denominator grew while nothing about the application's exposure moved.
Blending the three is like reporting one security score for a building by averaging the fraction of doors locked, the fraction of the perimeter fenced, and how often the night guard stays awake. The average looks reassuring precisely when one of the three is zero.
saying these in an interview costs you the question
- Produces one blended guardrail-coverage percentage
- Presents rules-exercised as a safety figure
- Omits the unmatched routing path from the report because it has no number
- Reports coverage without naming the judge model and embedding model it was measured against
- Treats layers in series as if their coverage averaged