Why can a green policy test suite still miss what the real enforcer sends?
answer
- two documents, one rule
- defaults fill in what you thought optional
- the input arrives already rewritten
- capture rather than hand-write
- a rule that never fires is a finding
basics
~20 sBecause the fixtures were hand-written from documentation while the enforcer sends a defaulted, already-rewritten document. The tests prove the rule works on a shape nobody produces. Validate captured real output against the declared contract in CI.
solid answer
~50 sHand-written fixtures encode the author's mental model of the document; the gate receives whatever the pipeline actually emits, and those diverge predictably. The producer fills in defaults, so a field the author believed optional is never absent in production. Earlier steps rewrite the document - overlays patch it, a chart injects a sidecar, a normalising step reformats it - so the gate sees a post-rewrite object the fixture never modelled. Values get normalised too: a quantity written as a number arrives as a string. Each difference turns a passing test into evidence about nothing. The remedy is to stop trusting fixtures as specifications: capture the document the gate actually received, base the fixture on it, and validate captured real output against the declared input contract on every CI run, so producer drift fails a build instead of silently disabling a rule.
go deeper
Understand that the document used in tests may not look like the one the gate receives, and that defaults and earlier rewrites are the usual reasons.
Explain the specific drift mechanisms — defaulting, rewriting, encoding, envelope — and why each turns a passing test into evidence about nothing.
Show the operational habit: capture the real input, keep a contract-conformance check in CI, and treat a rule that has never fired as a finding to investigate rather than a success.
Own the question of who is accountable when a producer upgrade silently disables enforcement, and what standing evidence exists that gates were still matching across the period an auditor asks about.
## Two documents, one rule There are always two documents in play. The one in the fixture directory, written by the rule's author from field documentation and a half-remembered example. And the one the enforcement point hands the engine, produced by a chain of steps the author did not write. A green suite says the rule behaves correctly on the first. It says nothing about the second, and the gap between them is where enforcement quietly disappears. ## How the two drift **Defaulting.** Producers fill in fields. A field the author treated as optional — and therefore wrote a deny-on-absence branch for — arrives with a default on every real document. The branch is dead code in production and fully covered in the suite. **Earlier rewrites.** The document is rarely delivered as authored. Overlays and patches are applied, a chart injects a sidecar container, a templating helper rewrites names, a normalising step reorders or reformats. If any step rewrites the document before the gate, the fixture must be taken from *after* that step; otherwise the rule is tested against an input that only ever existed in the author's head. This is the same ordering discipline as at admission, where mutating admission runs before validating admission, so a validating rule only ever sees post-mutation objects. **Encoding.** Numbers arrive as strings, booleans as the words `true` and `false`, a single item where the fixture had a list, an empty list where the fixture had a missing key. Comparisons that pass in the fixture stop matching. **Envelope.** The suite feeds one manifest; the gate feeds every object a render produced, in one stream. A rule that assumed a single object either evaluates the wrong thing or evaluates nothing. ## The signature of drift It is not a failing test. It is a rule with a perfect suite and no production denials — or, occasionally, the mirror image: a rule that suddenly denies everything after a producer upgrade. Both mean the same thing. Nobody was comparing the fixture's shape to the producer's output. ## Making drift a build failure 1. **Capture the real document.** Take the input at the enforcement point — the exact bytes the gate handed the engine — and use it as the basis of the fixture rather than typing one from documentation. The captured document already contains the defaults, the injected sidecar and the encodings. 2. **Contract-conformance test.** Keep the declared input contract executable and validate a sample of real producer output against it on every CI run. When the producer starts emitting a different shape, that check fails first, with a diff, instead of a rule silently ceasing to match. 3. **Pin and review the producer.** Record which version of the rendering chain the captured fixtures came from, and re-capture when it moves. A producer upgrade is a policy-relevant change even though no rule was touched. 4. **Watch the fire rate.** A rule that has produced zero violations since it shipped deserves a deliberate check: is the estate genuinely clean, or is the rule looking at a shape nobody sends? A single known-bad document pushed through the real path answers it in minutes. ## Why authors resist this Hand-written fixtures are small, readable and diffable; captured ones are long and noisy. That is a real cost, and the usual compromise is to keep captured documents as the source of truth for shape and trim them to the fields under test, mechanically, with the trimming itself reviewed. What is not an acceptable compromise is inventing the shape, because the invented shape is precisely the thing the suite cannot catch being wrong. ## The honest framing at interview A green suite is evidence that the rule computes what the author intended over the documents the author imagined. Confidence about the gate needs a second thing: evidence that those documents look like what the enforcer sends. Naming that as a separate check — and saying who runs it and when — is what separates someone who has operated a gate from someone who has only written rules.
- A rule denies when a field is absent, yet it has never fired in production. What is your first hypothesis?That the producer defaults the field, so it is never absent in a real document. The branch is exercised only by a hand-written fixture that omitted it. Confirm by capturing a real input at the enforcement point and looking for the field; if it is always present, the rule needs to decide on the defaulted value instead of on absence.
- How do you keep captured fixtures readable without inventing shape?Trim captured documents mechanically down to the objects and fields the rule reads, keeping the encodings and the envelope exactly as captured, and review the trimming step itself. Shape comes from the producer; brevity comes from a deterministic reduction. What you never do is retype the document from documentation, because that reintroduces the guess the capture removed.
- Which check would you run before trusting a newly written rule in a blocking gate?Push a known-bad document through the real enforcement path, not the test runner, and confirm the denial and its message. That single end-to-end pass catches envelope mismatches, defaulting surprises and a mis-registered gate at once — the class of problem a unit suite structurally cannot see.
It is unit-testing a parser against examples you typed yourself while production feeds it output from a machine you never inspected.
saying these in an interview costs you the question
- Treats a green suite as evidence the gate works
- Hand-writes fixtures from field documentation
- Ignores steps that rewrite the document before the gate
- Never questions a rule with zero violations
- Assumes optional fields stay absent in real output