skip to content

Your policy gate denied an unchanged plan at 09:00 and allowed it at 10:00 — what happened, and what do you change?

level: seniorimportance: should knowfreq 44%

answer

  1. the plan did not move; something else did
  2. diff the context, not the change
  3. which revision of the list decided?
  4. edits to facts need review too
  5. explicable beats fresh for a guardrail

basics

~20 s

The plan did not move; the facts around it did. Someone edited the approved list, or a live lookup answered differently. Pin the fact set to an identifiable revision and name it in every verdict.

solid answer

~50 s

A byte-identical change with two verdicts means the decision depended on something outside the change, and that something moved. Usually the approved-regions or approved-types list was edited between the runs, or the engine looked the fact up live and the source answered differently. Diagnose it by diffing the *facts*, not the plan: which revision of the fact set was loaded at 09:00 versus 10:00, and did any rules change too. The nondeterminism itself is not always wrong, an allow-list is supposed to be editable, but it must be explicable. So: give the fact set a published identity, make the engine report which revision decided alongside the verdict, and route edits to the fact set through the same review as rule edits. Then a differing re-run reads as "approved-regions moved from revision 46 to 47" rather than as an engine nobody trusts.

go deeper

for a junior

Recognise that a policy verdict depends on more than the file you submitted, so an unchanged change can be judged differently when the lists it is compared against are edited.

for a middle

Be ready to enumerate what could have moved between the runs, the fact set, the rules, a live lookup, a time-dependent rule, and to say which piece of evidence separates them.

for a senior

Demonstrate the diagnosis by comparing decision context rather than the change, and propose concrete fixes: versioned fact sets, the deciding revision named in the verdict, and fact edits reviewed like rule edits.

for a principal

Own the stance that a guardrail's product is the argument, not just the verdict, and decide organisationally when a re-evaluation may legitimately reverse an earlier pass and who may edit the facts that make it so.

## What the symptom actually tells you The same bytes went in and different verdicts came out. That rules out the change and points at everything else the decision consumed. In practice there are four candidates, and they are worth separating because the fixes differ. **The fact set moved.** Someone added a region to the approved list, or removed an instance type, between the two runs. This is the common case and it is the one the leaf is really about: the rule compared the plan against a set that is maintained by other people on their own schedule, and that set changed underneath. **The rules moved.** A new rule version was published in the interval. Diagnostically identical to the first case and equally invisible if nothing records which version decided. **A live lookup answered differently.** If the engine calls out during evaluation, the verdict tracks a remote system's state at that instant. Nothing local changed at all, and there may be no local record of what came back. **Something time-dependent is in the rule.** A rule that consults the current time, or an expiry, will flip at a boundary with no fact and no rule having changed. Rarer, and usually deliberate, but it presents identically. ## Diagnosing it Stop diffing the plan; you already know it is identical. Diff the *decision context*: 1. Which revision of the fact set was loaded by the engine at each run? If you cannot answer that, you have found the real defect already, and it is bigger than this incident. 2. Which version of the rules was loaded at each run? 3. Did any rule make an outbound call, and if so, was the response recorded anywhere? 4. Do the two verdicts name different rules, or the same rule reaching different conclusions? A different rule firing points at a rule publish; the same rule flipping points at a fact. The fastest confirmation is usually to find the edit: an allow-list committed at 09:40 that adds the region the 09:00 run rejected closes the case in one step. ## Why it matters more than one confused engineer A gate's authority rests on being explicable. The blocked engineer at 09:00 was told no; their colleague at 10:00 was told yes for the same plan. If nobody can say why, three things follow quickly. Engineers learn that retrying is a valid strategy for getting past the gate, which it now is. Anyone reviewing the record cannot tell whether a change that passed was actually compliant when it passed. And the next argument about a denial has no facts in it, only assertions about what the engine "usually" does. Note what is *not* wrong here. An allow-list is supposed to be editable, and a change legitimately becoming allowed when a region is approved is the system working. The defect is not that the answer changed. The defect is that the change of answer was invisible, unattributable and indistinguishable from a bug. ## The fixes, in order of value **Give the fact set an identity and report it.** Every published revision of the approved lists gets a version, and every verdict carries the versions that produced it. A denial reading "eu-south-2 not in approved-regions revision 46" is self-diagnosing: the engineer checks revision 47, sees their region, and asks for a republish instead of opening a ticket about a flaky gate. **Make fact edits as reviewed as rule edits.** If an allow-list can be edited by one person with no review while the rules require approval, the fact set is the soft underbelly of the whole gate. Anyone who can widen the list can render any rule that consults it toothless, and the change of behaviour will be attributed to the rule, not to them. **Publish facts as discrete revisions rather than continuously.** Continuous mutation means a change evaluated at 09:00:01 and 09:00:02 can differ. Discrete publishes, even frequent ones, at least give the answer a boundary you can name and a version you can quote. **Prefer a stale fact you can identify to a fresh one you cannot.** This is the judgment call the incident forces. A live lookup gives you the most current answer and the least explicable verdict; a replicated snapshot gives you an answer that may be hours old and an exact statement of what decided. For guardrails, explicability usually wins, because the gate's product is not just the verdict but the argument behind it. **Decide, deliberately, when a re-run is allowed to change the answer.** For a merge gate, re-evaluating at merge time against current facts is often right: policy tightened, and the change should be re-judged. For an approval already granted, silently re-deciding erases the meaning of the approval. Write down which you are doing, because both are defensible and doing them accidentally is not. ## The one-line summary for the postmortem The change was identical; the context was not. Any decision that consults facts outside the change is only reproducible to the extent that those facts are versioned and reported, and a gate that cannot name what decided cannot defend what it decided.

  • How do you tell a fact-set change from a rule change when both runs look the same?
    Compare which rule produced each verdict. A different rule name firing points at a rule publish in the interval; the same rule reaching opposite conclusions points at the data it consulted. If neither version is recorded with the verdict, that missing identity is the first thing to fix, because every future incident of this shape is equally unanswerable.
  • Is nondeterminism here always a defect?
    No. An allow-list is meant to be editable, so a change becoming permitted after a region is approved is the system working correctly. The defect is that the flip was invisible and unattributable. Make the answer explicable and the same behaviour stops being an incident and becomes a routine, readable outcome.
  • Should a re-run at merge time be allowed to reverse an earlier pass?
    It depends what the earlier pass meant. Re-judging against current facts is right for a merge gate, since policy may have tightened and stale approval is exactly what you are guarding against. Where the earlier pass was a granted approval, silently reversing it destroys its meaning. Pick one deliberately and state it, rather than inheriting whichever the tooling happens to do.
  • Why route allow-list edits through the same review as rule edits?
    Because widening the list defeats the rule just as effectively as deleting it, and with less visibility. If one person can add a region unreviewed, the fact set becomes the soft path around every rule that consults it, and the resulting behaviour change gets blamed on the rule rather than traced to the edit.

saying these in an interview costs you the question

  • Assumes the engine is flaky rather than checking the facts
  • Re-diffs the plan after being told it is identical
  • Fixes it by caching verdicts so answers stop changing
  • Treats an editable allow-list as inherently a bug
  • Leaves the fact revision out of the verdict message

context