Three payloads get through an application protected by a rail stack: one matched no pattern rule, one was approved by the rail that asks a judge model, and one matched no canonical example so the turn reached the application model on the unmatched path. Why file three separate defects rather than one, and what does the fix look like for each?
answer
- three owners, three definitions of done
- gap vs score vs ungoverned path
- ungoverned usually ranks highest
- verification recipe per ticket
- state what invalidates the fix
basics
~20 sEach has a different owner and a different fix. The unmatched pattern is a rule gap: widen or add a rule and it closes deterministically. The judge approval is a score problem: no edit guarantees that payload stays blocked, only the rate moves. The unmatched routing is worse still: that turn was never governed at all.
solid answer
~60 sA single 'guardrail bypassed' ticket is unfixable because the three causes live in different parts of the configuration and have different definitions of done. **Rule gap.** Enumerable and deterministic. Add or widen the rule, replay once, done. The residual risk is the payloads nobody enumerated, which is a coverage argument, not this ticket. **Judge approval.** A soft boundary. The owner can retune the judge's instruction, raise its strictness, or swap the judge model, and then the honest verification is a rate at a fixed N, not a single clean replay. This ticket can only be closed to a number. **Unmatched routing.** The turn never entered a governed path at all, so none of the checks attached to any intent applied to it. Adding one canonical example closes your phrasing and re-partitions the neighbourhood; the structural fix is making the unmatched path itself safe by default. The report should name the layer, the evidence (trace or differential run), the verification method, and what would invalidate the result — a new example, a moved cutoff, a new judge or embedding model.
go deeper
Recognises the three results came from different parts of the configuration and should not be one ticket.
Names the fix for each: add or widen a rule, retune or swap the judge and re-measure a rate, add an example or make the unmatched path safe by default.
Attaches per-ticket evidence and a verification recipe, ranks the ungoverned path highest despite its quiet symptom, and states what would invalidate each fix.
Makes the layer-and-verification split a house reporting standard, and pushes the architectural fix — deny by default on the unmatched path — rather than accumulating per-phrasing examples forever.
### The instrument's unit is a hit; the report's unit is a defect someone can close Three hits, one ticket, and the information the owner needs is gone. The three results differ along every axis a tracker cares about. | | Pattern rule missed | Judge rail approved | No canonical form matched | |---|---|---|---| | Config artefact to change | the rule or action inside the flow | the `self_check_input` template in `prompts.yml`, or the judge under `models:` | a `define user` example, the similarity threshold, or the default path | | Who owns it | the rule-set author | whoever owns judge configuration | whoever owns routing and the unmatched path | | What "closed" means | one replay of the exact payload | a pass rate at the same n as the original measurement | a family retest, or a structural default-deny | | What invalidates the fix later | a new payload variant | a judge-model or template change | a new example, a moved threshold, a new embedding model | Three owners, three definitions of done, three expiry conditions. That is three tickets in any tracker that has ever been used in anger. ### The mechanism behind each fix The **rule gap** is enumerable and deterministic: widen or add the pattern, replay once, done. The residual risk — payloads nobody enumerated — is a coverage argument for a different document, not an open state on this ticket. The **judge approval** has no closed state, only a moved number. The owner can tighten the self-check instruction, raise strictness, or swap the judge model, and each of those changes the *distribution* of verdicts near the boundary. Verification is a re-measurement at the original n, comparing rates. The **unmatched routing** is the structural one. Because the turn cleared no canonical form's threshold, nothing attached to any intent ran for it; flows under `rails.input.flows` and `rails.output.flows` are not intent-attached and still executed, so state that scope precisely rather than saying "ungoverned" flatly. The tempting fix — add your exact phrasing as a new `define user` example — closes one point and re-partitions the neighbourhood around it, which can pull unrelated turns onto the wrong path. The real fix is making the unmatched path safe by default. ### What it costs The costs are as unequal as the fixes, and that asymmetry is worth putting in the ticket so nobody plans them as one unit of work. Re-testing the rule gap is one request. Re-testing the judge ticket is 2n requests minimum — n before, n after — which at n=30 with a judge call plus a downstream generation per replay is a couple of hundred model calls and an hour of wall clock per verification round. Re-testing the routing ticket is the whole paraphrase family times its replay count, plus the authoring time to keep that family meaning-preserving, plus a re-run whenever the example set or embedding model moves. Triage time is the hidden line item: attributing each of the three hits to a layer, via trace or differential run, is usually more engineer-hours than finding them was. ### Where the number misleads Someone will ask for a headline: *"three bypasses in a 200-payload run, so a 1.5% bypass rate."* Two things are wrong with it. The denominator is **your own probe set** — it measures what you chose to send, and it moves if you send more of what already worked. And the numerator pools three incommensurable events: a binary rule miss that is certain, a soft-boundary rate that is a sample, and a skipped routing path that is not a rate at all. Averaging a certainty with a sample yields a number that is wrong in both directions at once. A combined severity does the same damage more quietly. And do not describe the ungoverned turn as *"the model answered a harmful question"*: that phrasing routes the ticket to the model team, when the finding is that the governing layer never ran. ### What you would check Each ticket carries the evidence that establishes its layer — a trace excerpt showing which rails executed, or the differential run that flipped the verdict — because without it the ticket is an assertion. Each carries a verification recipe the owner can actually execute: the exact payload for the rule, the payload plus n and the judge identity for the self-check, the paraphrase family plus example set, threshold and embedding model for the routing. Rank the ungoverned path highest despite its quiet symptom; candidates get this backwards because it produced the least dramatic output, but a hole in one rule is smaller than a hole through the routing layer. And say the measurability line out loud: deterministic layers fail as gaps you can name and close, model-backed layers fail as scores you can only move. A guardrail report written in a single voice is hiding one of them.
- Which of the three would you rank highest, and why is that counterintuitive?The unmatched routing. It looked like an ordinary answered turn, but it means the governing layer never ran, so every check attached to every intent was skipped.
- The owner asks for one severity number covering all three. What do you give them?Three severities plus a short narrative. Averaging a deterministic gap with a soft-boundary rate produces a number that is wrong on both counts.
- How does the verification step differ between the rule-gap ticket and the judge ticket?The rule gap is verified by one replay of the exact payload. The judge ticket is verified by the same payload replayed at the same N as the original measurement, comparing rates.
saying these in an interview costs you the question
- Files one combined guardrail-bypass ticket with a single severity
- Ranks the quiet unmatched-path finding lowest because it produced no dramatic output
- Gives the judge-layer ticket a binary closed state
- Omits from each ticket the evidence that establishes which layer failed
- Describes the ungoverned turn as a model problem rather than a routing problem