skip to content

Before a policy engine major upgrade, how do you prove your rules still deny what they denied before?

level: seniorimportance: should knowfreq 44%

answer

  1. the estate's objects, not the fixtures
  2. run both pairs, diff the outcome
  3. compare decisions, ignore message churn
  4. two flip directions, opposite meanings
  5. deny-to-allow blocks the window

basics

~20 s

Replay a corpus of the estate's real stored objects through both the old and the new engine-and-rules pair, then diff the decisions. Deny-to-allow flips are enforcement regressions; allow-to-deny flips are the blast radius you are about to inflict.

solid answer

~50 s

Rule unit tests only prove the rules behave on inputs their authors imagined, which is exactly the set an upgrade does not surprise. So I build a corpus from what the estate actually contains — every stored object of the governed kinds, deduplicated by the subtree the rules read, plus retained real denial inputs and every must-deny fixture — and evaluate it twice: once on the current engine with the current rules, once on the new pair. Then I diff the *decision*, not the message text, because reason strings churn cosmetically across majors and will drown the signal. Two flip directions matter and mean opposite things. Deny to allow is an enforcement regression and blocks the upgrade until explained. Allow to deny is the blast radius: those denials will land on real teams on day one, so they need either a fix window or a staged rollout. A clean replay is necessary, not sufficient — it says nothing about object shapes the corpus does not contain.

go deeper

for a junior

Know that rule tests use inputs their authors invented, and that an upgrade is best checked against real objects from the running estate. Understand that the comparison is between the old and the new engine on the same inputs.

for a middle

Explain how the corpus is assembled and deduplicated, and why the diff is on the decision rather than the message. Be able to say what a deny-to-allow flip means versus an allow-to-deny flip.

for a senior

Show that you have run this: sourcing denial inputs, converging the corpus on distinct shapes, triaging both flip directions, and stating plainly what replay does not cover. Expect a scenario asking whether findings block the window.

for a principal

Own the standing investment — replay as a scheduled job rather than an upgrade ritual, who pays for the retained inputs, and how the result becomes the artefact you hand someone asking whether the control kept operating across the change.

## Why the rule suite is not the evidence A rule library's own tests are written by the people who wrote the rules, against the inputs those people had in mind. An engine major upgrade breaks rules in ways nobody had in mind — a changed input document, a differently assembled result, a shift in how a matching construct behaves at the edges. The fixtures move with the rules, so they keep passing. **A green suite after an engine upgrade is a statement about your imagination, not about your estate.** The evidence an upgrade actually needs is behavioural: for the inputs that really occur, does the new pair decide what the old pair decided? ## Building the corpus The corpus is what the estate holds, not what the tests hold. Sources, in order of value: - **Stored objects of the governed kinds**, pulled from the live environments. This is the bulk, and it is the only source that carries the shapes you did not anticipate — including objects created years ago under versions nobody writes any more. - **Retained denial inputs.** If the enforcement point records the input of each refusal (scrubbed of secrets, and with a retention window you have agreed), those are gold: they are the only naturally occurring examples of the deny path. - **Every must-deny fixture you own**, so the corpus exercises denials even where the estate is fully compliant. The estate is large and mostly repetitive, so deduplicate — but on the right key. Hashing the whole object keeps thousands of near-identical replicas; hashing only the kind and version throws away real variation. Deduplicating on the kind, the API version, and a normalised projection of the subtrees the rules actually read keeps one representative per distinct decision-relevant shape and usually collapses an estate by an order of magnitude or two. One caution about the deny path: a corpus drawn from live objects is by definition the set of things that were *admitted*. If a rule has been enforcing for a year, the estate contains almost no violations of it, and the replay will exercise that rule's allow branch only. That is why the retained denial inputs and the deliberately bad fixtures belong in the corpus — without them you are proving that the new pair still allows, which is the half you care less about. ## Diffing decisions, not messages Run the corpus through old-engine-plus-old-rules and new-engine-plus-new-rules and compare, per input, the decision: allowed, denied, or errored. Normalise or ignore the human-readable reason text — majors reword it, reorder violation lists, and change how they render paths, and a message-level diff turns into thousands of findings you will learn to skip. Four buckets come out, and they mean different things: | Flip | Meaning | Response | |---|---|---| | deny to allow | Enforcement regression. A rule has gone quiet on inputs it used to catch. | Blocks the upgrade until each case is explained. This is the bug the whole exercise exists to find. | | allow to deny | Either a rule now catches something it always should have, or the new engine is stricter in a way you did not intend. | Triage per case; this is your day-one blast radius on real teams. | | error, either side | A rule that no longer evaluates, or one that now does. | Fix before the window; an error's handling differs by enforcement point and you do not want to discover which. | | no change | The expected majority. | Report the count — it is the denominator that makes the other numbers meaningful. | The allow-to-deny bucket is the one that gets an upgrade rolled back at 9am, and it is often not a defect at all: the new pair is correct and the estate was never compliant. That is still a decision to make deliberately — announce it, give teams a window, or ship the engine in a non-blocking posture first — rather than discovering it as an incident. ## What replay does not prove Be explicit about the limits when you present the result: - It covers **shapes the corpus contains**. The first object created in a new API version after the upgrade is, by construction, not in it. - It covers **the rules as evaluated**, not the deployment: bundle loading, scoping, and where the enforcement point sits are separate risks with separate checks. - It is a **point-in-time** result. If the rule library or the engine build moves again before the window, rerun it; a replay from three weeks ago is a document, not evidence. So pair it with the checks that cover the other blind spots: a must-deny canary per rule written in each accepted shape, and denial-rate monitoring for the days after the window. Replay is the strongest single piece of pre-upgrade evidence available, and it is worth keeping as a scheduled job rather than an upgrade ritual — run monthly, it also catches the platform-side changes that nobody scheduled with you.

  • Most of your estate is compliant, so the replay barely exercises the deny path. How do you fix that?
    Add inputs that were actually refused — if the enforcement point retains denial inputs, scrubbed and time-boxed, those are the real deny-path corpus. Top up with every must-deny fixture in the rule library, and with mutated copies of real objects: take a representative object per shape, break the one field each rule cares about, and add it. That gives each rule at least one denial to prove on top of the estate's allow cases.
  • The replay shows forty existing workloads that the new pair denies and the old pair allowed. Do you block the upgrade?
    Not on the count alone — first separate the causes. If the new pair is correct and those forty were always violations the old rules missed, the upgrade is fine and the forty are a remediation campaign with a deadline. If the new engine is stricter in a way the policy never intended, that is a rule bug to fix before shipping. Either way you do not surprise forty teams: announce, or ship the engine non-blocking first and turn enforcement on afterwards.
  • How large a corpus is enough?
    Size is the wrong metric; distinct decision-relevant shapes is the right one. Deduplicate on kind, API version and a normalised projection of the fields the rules read, then check the curve — when adding another environment's objects stops adding new shapes, you have converged. A few thousand distinct shapes usually covers a large estate, and it runs in minutes, which matters because you want to rerun it on every rule-library change, not just at upgrades.

saying these in an interview costs you the question

  • Treats the rule library's own unit tests as upgrade evidence
  • Diffs reason text instead of the decision
  • Replays only compliant objects, never exercising a deny
  • Reads a clean replay as proof for future object shapes
  • Ignores allow-to-deny flips until teams are blocked

context