How can a policy engine get a fact, such as an approved-region list, that the change never carries?
answer
- only three shapes exist
- ahead of time, alongside, or mid-decision
- each buys a different freshness
- live lookup inherits another system's uptime
- most facts move slowly
basics
~10 sThree routes: replicate the fact into the engine ahead of time, have the caller pass it in with the change, or look it up live during evaluation. Each buys different freshness and fails differently.
solid answer
~50 sThere are only three shapes. **Replicate ahead of time**: the approved-regions list is published to the engine with the rules, so the decision is fast, self-contained and reviewable, but only as current as the last publish. **Pass it in**: whatever assembles the request collects the fact and includes it in the same document as the change, so the fact that decided is captured next to what it decided about, and freshness becomes the caller's problem rather than the engine's. **Look it up live**: the engine calls out mid-evaluation, giving the freshest possible answer at the cost of latency, an extra dependency that can take the gate down with it, and a decision whose inputs nobody recorded. I default to replicating slow-moving facts such as region and instance-type allow-lists, pass in facts that belong to this specific request, and reserve live lookups for facts that genuinely cannot tolerate being minutes old.
go deeper
Know that a fact missing from the change has to be supplied somehow, and be able to name the obvious two: it was loaded into the engine beforehand, or it was sent along with the change.
Be ready to lay out all three routes and what each buys, and to say concretely how stale a nightly-published allow-list can be and what that means for a region added this morning.
Show judgment: sort the facts by how fast they move and how much a stale answer costs, defend replication as the default, and explain what a live lookup does to the gate's availability and to explaining a verdict afterwards.
Own the platform-wide stance on whether decisions may depend on remote systems at all, and the operational commitment that follows, since every allowed lookup is another service whose outage becomes a deployment freeze.
## The problem A rule says instances may only be created with an approved type, in an approved region. The proposed plan supplies the type and the zone. It cannot supply the approved lists, because those are organisational facts, not properties of the change. So the fact has to reach the engine somehow, and there are exactly three shapes for that. Choosing between them is a policy-author decision with consequences that outlive the rule. ## Route 1: replicate the fact ahead of time The approved-types and approved-regions lists are maintained somewhere, and a copy is published to the engine, arriving with or alongside the rules that use it. - **Freshness**: as current as the last publish. If the list is republished nightly and someone adds a region at 10:00, changes evaluated before the next publish are judged against yesterday's list. - **Failure mode**: the engine keeps deciding even if the source system is down, because it is not consulted at decision time. The risk is not unavailability but silent drift, a copy that has stopped being refreshed while everyone assumes it is current. - **Reviewability**: the strongest of the three. The fact is a published artifact with an identity, so "which list decided?" has an answer, and changes to the list can go through review the same way rule changes do. - **Fits**: slow-moving facts of modest size. Allow-lists, ownership maps, environment classifications. ## Route 2: pass it in with the change Whatever assembles the request gathers the fact and puts it in the same document as the change, so the engine receives one self-contained input. - **Freshness**: exactly as fresh as the moment the caller collected it, which is usually seconds before the decision. - **Failure mode**: the caller's problem. If the fact cannot be collected, the caller decides whether to proceed, retry, or refuse, and it does so where a human can see it rather than deep inside rule evaluation. - **Reviewability**: good, because the fact that decided travels in the same document as the thing it decided about, so what the engine saw is a single object. - **Cost**: every caller now has to know which facts the rules need. Add a rule that needs a new fact and you are changing the caller too, which is a coupling worth being deliberate about. - **Fits**: facts specific to this request, and facts too large or too volatile to replicate wholesale. ## Route 3: look it up during evaluation The engine reaches out to another system while deciding. - **Freshness**: the best available. The answer reflects the source at the instant of the decision. - **Failure mode**: the worst. A gate that calls another system inherits that system's availability and latency on every decision, and "the approvals service is slow" becomes "deployments are blocked", or worse, "deployments are allowed", depending on how the enforcer handles an error. - **Reviewability**: the weakest. The verdict now depends on something outside the recorded input, so two runs over the identical change can legitimately disagree and nothing kept says why. - **Fits**: genuinely time-sensitive facts, such as a break-glass exemption that must take effect immediately, and even then it should be a narrow, bounded call rather than the general way facts arrive. ## Choosing Sort your facts by how fast they move and how badly a stale answer hurts. | Fact | Moves | Stale answer costs | Route | | --- | --- | --- | --- | | Approved regions | Rarely | A new region rejected for a few hours | Replicate | | Approved instance types | Occasionally | A new type rejected until republish | Replicate | | This request's requester and branch | Per request | Wrong decision entirely | Pass in | | An emergency exemption just granted | Minutes | The fix for a live incident stays blocked | Look up, narrowly | Two things fall out of the table. First, most facts a gate needs move slowly, so most facts should be replicated, and the instinct to reach for a live lookup is usually the expensive answer to a cheap problem. Second, the question is never "is stale data acceptable?" in the abstract. It is "how stale, and who notices?" Every route is stale by some amount, including the live lookup, which is stale by exactly the network round trip and by however long the source system took to learn the fact itself. ## Make the staleness visible Whichever route you pick, the fact set that decided should be identifiable in the outcome: which published revision of the approved-regions list, or which snapshot the caller collected. A denial reading "eu-south-2 is not in approved-regions revision 47" tells the blocked engineer that their region was added after revision 47 and that they need a republish, not an argument. The same denial without that identity tells them only that the robot said no, which is how gates acquire the reputation that eventually gets them switched off.
- Why is replicating the allow-list usually the right default rather than a compromise?Because most facts a gate needs move slowly and are small: approved regions, approved types, ownership maps. Replication makes the decision fast, self-contained and survivable when the source system is down, and it gives the fact set an identity you can point at when someone asks which list denied them. Freshness measured in hours is genuinely fine for a list that changes twice a year.
- What is the honest cost of having the caller pass facts in with the change?Coupling. Every caller must know which facts the current rules need, so adding a rule that needs a new fact becomes a change to every enforcement point, not just to the rules. That is tolerable when there are few callers and painful across dozens, which is often the deciding factor rather than freshness.
- Your gate must honour an exemption granted five minutes ago. Does that force a live lookup?Not necessarily. It forces the exemption fact to reach the engine within minutes, which a frequent republish of just that small fact set also achieves. Prefer shortening the replication interval for the one volatile fact over making every decision depend on a remote call, and if you do call out, bound it tightly and decide in advance what an error means.
saying these in an interview costs you the question
- Reaches for a live lookup as the default route
- Claims replicated data has no staleness because it is versioned
- Ignores that a live lookup adds an availability dependency
- Treats freshness as the only axis of the choice
- Cannot name which fact set produced a given verdict