When you promote a case into a guardrail regression suite, what should be stored with it so a run months later can decide pass or fail on its own — and why is storing the response text captured on promotion day a poor expectation?
answer
- pin the decision, not the prose
- verdict + category + reason
- text match dies on a model swap
- score assertions need a margin
- provenance is what lets you delete
basics
~20 sStore the input, the decision you expect (blocked or allowed), the policy category it should fire on, and a one-line reason the case exists. Pinning the exact response text fails because generation is not stable across model swaps or sampling, so the case goes red on harmless wording changes and gets deleted or ignored.
solid answer
~50 sPin the **decision**, not the prose. A case needs: the input, the expected verdict (block or allow), the category or rule that should fire, the reason it was promoted with a link to the original finding, and the identity of what it was verified against. Exact-text expectations break for a reason specific to this stack: the response is generated, the guard's confidence is a float, and both move when the model behind the guard changes or sampling is non-zero. A text-match case then reports red for wording that is behaviourally identical. Teams respond by re-running until green, or by deleting the case — either way the case stops protecting anything. The same argument caps how tight the expectation may be. A verdict is stable; a confidence to three decimals is not. If the case is about a numeric score, express it as a direction and a margin, and expect to revisit that margin when the guard is upgraded.
go deeper
Knows the case stores an input and an expected outcome, and that comparing full generated text is brittle.
Names the fields — verdict, category, provenance, what it was verified against — and explains why generation and confidence scores are unstable assertions.
Handles the awkward cases: right verdict for the wrong reason, threshold-adjacent cases, and re-verification after a guard upgrade instead of treating rescaled scores as regressions.
Treats the record as the contract that makes the suite maintainable — provenance so cases can be retired, identity so reds can be attributed, and assertion tightness as a deliberate trust-versus-noise choice.
### What "the expectation" is, and why its tightness is the whole design A regression case is a stored input plus an assertion that a harness evaluates months later against whatever the stack returns. That assertion is the entire design decision. Too loose and the case passes while the behaviour underneath it rots. Too tight and it reports red for changes that mean nothing — and that is the worse failure, because a case that cries wolf gets muted, re-run until green, or deleted, and then it protects nothing at all. Two things move independently beneath every case. The **guard** — a hosted moderation endpoint such as OpenAI Moderation or Azure AI Content Safety, or a classifier you run yourself such as Llama Guard, ShieldGemma, Granite Guardian or Prompt Guard — is upgraded, retrained and re-thresholded, which rescales its confidence outputs. Separately, the **model generating the assistant's answer** behind that guard is swapped, and even at temperature 0 its wording is not guaranteed byte-stable across serving stacks, quantisations or minor version bumps. An expectation that touches generated prose or a raw float is standing on both of those moving surfaces at once. ### The record that survives ``` input: exact prompt/request, plus the channel it enters through expected_verdict: block | allow expected_category: the policy label that should fire (the guard's own taxonomy key) provenance: why this was promoted; link to the triaged finding or incident verified_against: guard build/endpoint + config/threshold + serving-model id + date ``` - **Input** must include the entry point when the case is about a specific channel; the same string screened on a different path is a different case. - **Expected verdict** is the assertion, and it is the only field the run passes or fails on. It is stable because it is the behaviour a user experiences. - **Expected category** is checked and reported, not asserted as a hard failure. A case that blocks for the wrong reason is a weaker pass than a case that blocks for the right one, and it often precedes real drift. - **Provenance** is what makes deletion possible. Without it nobody can say what a case is the only witness for, so nothing is ever safe to retire and the suite grows monotonically. - **verified_against** is what lets a later red be attributed to *the stack changed* rather than *the guard regressed*. ### Two tempting shortcuts to reject Hashing the moderation endpoint's full JSON body gives a beautifully exact assertion that breaks on any added field, key reordering or score movement — all behaviourally irrelevant. Capturing the model's refusal text verbatim turns the case into a paraphrase detector: an equally-correct refusal, worded differently, fails it. Both feel rigorous and both convert a routine upgrade into a wall of reds. ### What brittleness costs The cost is paid in human triage, and it compounds. Assume 1,000 cases and 5% of them assert on text or a raw score. A guard upgrade rescales confidences and rewords refusals, so roughly 50 cases go red at once. At five to fifteen minutes each to open the raw response, decide "cosmetic", and re-baseline, that is a full engineer-day per upgrade — for zero findings. Teams do the rational thing after the second occurrence: they stop reading the reds. From that point the suite's real detection rate is zero regardless of what it costs to run, which is a far more expensive outcome than the inference bill it was optimising. ### Where the number misleads Two readings go wrong. First, a **green from the wrong reason**: the verdict matched, so the harness reports pass, while the block came from a different policy category, a truncated response, or an unrelated layer such as an auth failure returning a canned refusal. Checking the category alongside the verdict is what surfaces that. Second, a **score assertion after a rescale**: if a case pins the guard's confidence with a direction and margin ("score above the block threshold by at least 0.1"), an upgrade that changes the scale makes that red an artefact, not a regression — and, in the other direction, a rescale can silently widen the margin so a case that used to sit near the boundary now passes with room to spare and no longer tests anything. Threshold-adjacent cases are inherently fragile; after a guard upgrade the honest move is to re-verify them by hand rather than to read their reds as findings or their greens as safety. ### What to check on an inherited suite Search the case definitions for assertions on response text, response hashes or raw floats, and count them. Count cases with no provenance line — those cannot be retired and cannot be justified. Compare each case's `verified_against` guard build to the one currently deployed; every case verified against a build that no longer exists is an expectation nobody has confirmed since the last change.
- The guard blocks the case, but under a different policy category than the one recorded. Pass or fail?Report it distinctly from a plain pass. The user-visible outcome is right, but the reason changed, which often precedes a real drift; investigate rather than silently green-lighting it.
- Why record what the expectation was verified against?So a later red can be attributed. Without the guard build, configuration and serving-model identity, you cannot tell a genuine regression from the stack having been swapped underneath the suite.
saying these in an interview costs you the question
- Asserting on exact response text or a hash of the moderation endpoint's full JSON.
- Pinning a confidence value to several decimals and calling any movement a regression.
- No provenance line, so nobody can justify retiring a case and the suite only grows.
- Recording nothing about the guard build or serving model, leaving every red unattributable.
- Silently re-baselining an expectation whenever it goes red, which converts the suite into a recorder of current behaviour.