skip to content

A guardrail regression suite has failed the same three cases for months. This week it passes all three, with no change to the suite. What do you check before recording it as fixed, and what should each run have been recording to let you answer that?

level: seniorimportance: must knowfreq 52%

answer

  1. unexplained green = investigate
  2. diff run identities, not pass counts
  3. harness error counted as pass
  4. hosted guard changes with no version
  5. stable canary cases as drift signal

basics

~20 s

Treat a green that nobody caused as suspect. Check whether the stack changed underneath the suite: guard version or endpoint, its policy configuration, and the model serving responses. Also check the harness — a verdict check that errors and counts as pass looks identical. Every run should record what it tested against, not just the results.

solid answer

~50 s

An unexplained green is an event to investigate, not a win. Three explanations compete: 1. **The guard genuinely improved** — a version bump or policy change closed the gap. 2. **The stack changed underneath you** — the endpoint now serves a different model, or the guard is a hosted service that updated without anything you control changing. The three cases may no longer exercise what they were written against. 3. **The harness broke** — the object deciding whether a response counts as a hit is erroring, timing out or returning a default, and the run counts that as pass. You can only distinguish these if each run records the identity of what it tested: guard build/endpoint, its configuration and thresholds, the serving model identifier, the suite revision, plus per-case raw responses and any harness errors. Without that, the diff between two runs is only pass counts, which is exactly the information that cannot answer the question. Re-verify by hand on the cases that flipped before updating any expectation.

go deeper

for a junior

Says to re-run the cases and look at the actual responses rather than trusting the pass count.

for a middle

Names the competing explanations — real fix, changed model or guard, broken harness — and knows the run needs to record what it tested against.

for a senior

Gives a triage order, separates harness errors from case failures, and knows a hosted guard can change with no version signal so drift must be detected behaviourally.

for a principal

Designs for attributability up front: identity captured per run, drift alerts routed to an owner rather than a build gate, and a standing rule that an unexplained green is investigated like a red.

### Why an unexplained green is an incident, not a win The premise of this suite is that the thing under test moves without asking you. A hosted guard — OpenAI Moderation, Azure AI Content Safety, a vendor firewall — is upgraded on the vendor's schedule. A self-hosted classifier such as Llama Guard or ShieldGemma is re-pinned to a new weights tag by a dependency bump. The model serving the assistant's answers behind the guard is swapped for a cheaper or newer one by a routing change you never saw. Any of those can flip a long-failing case to green without a single line of your code changing. So a pass has to be *attributable*, and attribution is a property of what the run recorded at the time — it cannot be reconstructed afterwards from a pass count. Three explanations compete for the flip, and they demand opposite actions: | explanation | what the evidence looks like | what you do | |---|---|---| | the guard genuinely improved | guard build/config changed; raw responses show a correct block for the right reason | close the finding, keep the cases, note the fix against their provenance | | the stack changed underneath you | serving-model or guard identifier differs from the last red run | re-verify: the cases were written against something that no longer exists | | the harness broke | non-zero error count, empty or truncated raw responses, suspiciously uniform outcomes | fix the harness; the run is void, not green | ### What every run must record Alongside the results: the guard identity (build or endpoint, plus configuration and thresholds), the serving-model identifier, the suite revision, the raw request and response body per case, per-case latency, and harness errors tracked separately from case failures. The point is that two runs can be diffed on *inputs and identities*, not only on outcomes. A run artefact that stores "1,997 passed, 3 failed" has thrown away the only information that can answer this question. Storage is cheap by comparison — a few kilobytes of raw response per case, so a 2,000-case run is single-digit megabytes, and thirty days of nightly runs is a few hundred megabytes with response bodies redacted or hashed where policy requires. ### Triage order 1. **Diff the recorded identities** between the last red run and this green one. A changed guard build or serving-model id explains the flip immediately and reframes it: those cases were verified against a system that is gone. 2. **Check harness health.** A scoring step that throws and is caught into a default outcome produces a clean sweep of greens. Any run with a non-zero error count should not be published as a pass at all. 3. **Re-run the three cases by hand and read the raw bodies.** A block for a new reason, a truncated response, an empty completion, or a canned refusal produced by an expired credential all read as "blocked" to an automated verdict check. 4. **Only then record it as fixed**, and write the change that fixed it against each case's provenance so the case is justifiable later. ### Where the number misleads The pass count is the most misleading artefact a run produces, because every failure mode above *raises* it. A harness whose verdict step defaults to "blocked" scores 100% on an attack-heavy suite. An expired API key that makes the application return a fixed refusal scores 100%. A swapped endpoint that no longer routes through the guard at all may also score 100% if the model happens to refuse on its own. In each case the number moves in the direction people celebrate. This is why the symmetric heuristic matters and is worth stating explicitly: a suite that goes red *across the board* is far more often a swapped endpoint, an expired credential or a harness fault than a simultaneous regression in dozens of behaviours — and the same recorded identity answers both directions. ### The hosted-guard blind spot A hosted guard may expose no version identifier at all, so the identity diff has nothing to compare. The compensating control is behavioural: carry a small set of stable, deliberately unchanging cases — a handful of clear blocks and clear allows — and record their verdicts and, where available, their raw scores on every run. Movement in that set is evidence the *service* changed even though nothing in your configuration did. Treat that as a drift signal, not a build failure: it should page the owner of the suite to re-verify expectations, not turn a deploy red, because the guard vendor changing is not your application regressing. ### What to check before recording "fixed" Concretely: the identity diff between the two runs; the run's error count; the raw response bodies for the three cases read by a human; whether the block came under the expected policy category; and whether the drift canaries moved in the same window. Only when all five agree is "the guard fixed it" the cheapest explanation.

  • The entire suite goes green overnight, not just three cases. First hypothesis?
    The harness, not the guard. A broken verdict step, a swapped endpoint or an auth failure that returns a canned refusal will pass everything. Check error counts and raw responses before believing it.
  • The hosted guard exposes no version identifier. How do you detect that it changed?
    Behaviourally: keep a small stable set of cases, record their verdicts and scores every run, and alert on movement in that set. It signals drift in the service rather than a regression in your application.
  • Once you confirm the guard really did fix the three cases, what do you do with them?
    Keep them, with a note of the change that fixed them. They now guard against the fix being reverted or lost in the next upgrade, which is the whole point of a permanent slot.

A harness that catches its scoring exception into a default of "blocked" is an exam that marks every unanswered question correct: the highest score in the room belongs to the candidate who never picked up a pen. That is why the error count has to be reported next to the pass count, not folded into it.

saying these in an interview costs you the question

  • Closing the finding on the strength of a green run with no check of what changed.
  • Runs that record only pass and fail counts, making any flip unattributable.
  • Counting a harness or network error as a passing case.
  • Assuming a hosted guard is static because you did not change its configuration.
  • Re-baselining the three expectations to match the new behaviour before anyone has read the raw responses.

context