skip to content

Policy Engine Fundamentals

Every policy engine splits deciding from enforcing, places that decision somewhere in delivery, and must still answer when the engine is unreachable. Interviewers open here before naming tools.

on this pageshow

explore

questions

page 1 of 2

In a policy decision, what separates the change under evaluation from the surrounding facts?

level: juniorimportance: must knowfreq 72%

answer

  1. two documents, not one
  2. the proposal versus the standard
  3. the allow-list is not in the plan
  4. could the author edit it true?
  5. claims come from the gated party

basics

~20 s

The change under evaluation is the document being decided on, such as a proposed plan or manifest. The surrounding facts are everything else the rule needs, like the list of approved instance types, which the change never carries.

solid answer

~50 s

A policy engine evaluates a document it is handed, and that document has two halves. The **change under evaluation** is what the author proposes: a plan to create an instance, a manifest, a job definition. The **surrounding facts** are what the rule must compare it against, and they are not properties of the change at all: which instance types are approved, which regions the company may run in, who owns which namespace. A rule like "only approved instance types, only approved regions" reads `m5.24xlarge` and `eu-central-1` straight out of the plan, but the approved lists appear nowhere in it and never will. The useful test is: could the change's author make this predicate true by editing their own file? If yes, it is a claim from the party being gated, not a fact. The change is untrusted input; the fact set must arrive by a path the author does not control.

code

json · 15 lines
json
{
  "resource_changes": [
    {
      "address": "aws_instance.batch",
      "type": "aws_instance",
      "change": {
        "actions": ["create"],
        "after": {
          "instance_type": "m5.24xlarge",
          "availability_zone": "eu-central-1a"
        }
      }
    }
  ]
}

go deeper

for a junior

Be ready to state plainly that a decision compares two things: the change being proposed and the facts it is judged against. Know that the second half is supplied separately and is not hiding inside the plan or manifest.

for a middle

Explain why a value the change's own author controls is a claim rather than a fact, and give an example of a fact the change cannot carry at all, such as current account spend or an approval recorded elsewhere.

for a senior

Show that you decide this split deliberately for each rule, and that you have thought about what the gate does when the fact set is missing or empty rather than assuming it will always be there.

for a principal

Own the question of who maintains each fact set and by what review path, since the trust in every fact-driven rule reduces to whether the gated party can influence the standard they are measured by.

## Two halves, not one document A policy engine does not "look at your infrastructure". It evaluates a document that something hands it, and that document splits into two conceptually different halves. Confusing them is the first mistake a new policy author makes, and it survives into surprisingly mature gates. The **change under evaluation** is the thing being decided on: a proposed plan for infrastructure, a manifest about to be admitted to a cluster, a pipeline job definition, an image's metadata. It is authored by the person the gate exists to constrain, and it describes what they want to happen. It is self-contained and self-describing. The **surrounding facts** are everything else the rule needs to turn that description into a verdict. Which instance types finance has approved. Which regions the company is permitted to operate in. Which cost-centre codes are valid this quarter. Which team owns which environment. None of that lives in the change, and none of it can, because it is not a property of the change. It is a property of the organisation at the moment the change is judged. ## The worked case Take a rule that says: an instance may only be created with an approved type, in an approved region. The plan carries an instance type and an availability zone. Extracting those two values is trivial. The rule still cannot decide anything, because the interesting half of the comparison, the *set* of approved types and the *set* of approved regions, is nowhere in the plan. The change says what is wanted. The fact set says what is permitted. A decision is the comparison. An engine given only the first half can answer only trivial questions such as "is an instance type set at all?". ## The test that separates them Ask one question of every value your rule reads: > Could the author of this change make my predicate true by editing their own file? If the answer is yes, that value is part of the change, and it is a **claim**, not a fact. A manifest carrying a label reading `approved: true`. A plan with a tag declaring `environment = "dev"` so the production rules skip it. A build that states its own risk tier. Each of these is an assertion written by exactly the party the gate constrains. A rule that trusts one has not been defeated by a clever attacker; it was written to ask the gated party for permission. This is why the split is a security property rather than a modelling nicety. The change is untrusted input. The fact set is trusted, and it earns that trust from *where it comes from*: a list a different group maintains, reviewed on a different path, delivered to the engine by something the change's author cannot edit. That does not make every value in the change useless. Values that describe what will actually be created, the instance type that will be provisioned, the image that will run, are exactly what you want to gate on, because the platform will honour them. What you must not do is let the change also supply the *standard* it is judged against. ## Facts a change can never carry Some facts are not merely absent from the change; it is structurally impossible for the change to carry them. How many instances the account already runs. What has been spent this month. Whether a human approved this in a ticketing system. Whether the same rule denied this yesterday. A change is a description of one proposed state, so anything about *the rest of the world*, or about *history*, has to arrive from somewhere else, or the rule cannot be written at all. Spotting early which of your intended rules need such a fact is what tells you whether a gate is cheap or expensive to build, long before anyone writes a line of rule. Broadly there are three ways to get such a fact in front of the engine: replicate it into the engine ahead of time so it is loaded alongside the rules, have the caller collect it and pass it in with the change, or have the engine look it up during evaluation. They buy very different freshness and very different failure behaviour. ## Where this goes wrong in practice - Writing a rule against whichever field the change happens to carry today, without ever deciding whether that field is a claim or a fact. - Assuming the engine can see the world. Most engines see only what they were given. - Degrading quietly when the fact set does not arrive. Many rules, deprived of the list they compare against, produce no opinion rather than a denial, and a gate with no opinion lets the change through. - Treating "the engine saw it" and "the engine was told it" as the same statement. Only the second is ever true. ## What good looks like Before writing the rule, name the two halves out loud: this is the change, these are the facts, and this is who maintains each. Nearly every hard question later in a gate's life, staleness, reproducibility, who may edit the allow-list, is a consequence of that one split being drawn deliberately rather than by accident.

  • The plan already names its region, so why is that not enough to decide whether the region is allowed?
    Because the plan states what the author wants, not what the organisation permits. The region name is one operand; the approved-regions set is the other, and it lives outside the change. Without it the rule can only check that a region is present at all, which denies nothing anybody cares about.
  • A manifest carries a label saying it was approved. Can a rule act on that?
    Only as a signal, never as the fact. The author of the manifest writes that label, so it is a claim by the party being gated. If approval matters, the record of it has to come from wherever approvals actually live and be supplied to the engine on a path the author does not control.
  • What happens to a rule when the fact set fails to arrive?
    Usually nothing visible, which is the danger. A comparison against a missing or empty set often yields no denial rather than a denial, so the gate silently turns into a pass-through. Decide explicitly what absent facts mean and make the engine or its caller fail loudly instead of quietly allowing.

A passport control officer reads your passport, but the watch-list is not printed inside it. The traveller supplies the claim; the state supplies the standard.

saying these in an interview costs you the question

  • Thinks the approved-types list is somewhere inside the plan
  • Assumes the engine can query the cloud account by itself
  • Trusts a label the change's own author wrote
  • Cannot say where a fact outside the change comes from
  • Says a missing fact set makes rules deny by default

context

open as a page

What do PDP, PEP and PIP each do in a policy-as-code gate?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A policy decision point evaluates the rule and returns an answer. A policy enforcement point sits in the change's path, gathers the input, asks, and acts on the reply. A policy information point supplies facts the change itself does not carry.

open as a page

A CI job enforces backup retention with a shell script that exits 1. What does restating it as a policy rule buy?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A rule file states the condition as data, so the engine can report which rule failed on which resource, and the same rule can be listed, reviewed and reused elsewhere. An exit code carries one bit: pass or fail.

open as a page

A policy gate blocks your deploy with only 'denied by policy' — what should that denial have carried?

level: juniorimportance: must knowfreq 66%

basics

~20 s

A usable denial names the rule that fired, the object and the exact field it failed on, what the rule requires, and one concrete fix. It should list every violation found, not only the first.

open as a page

A deploy tool calls a policy decision service and gets no answer - what do fail-open and fail-closed mean here?

level: juniorimportance: must knowfreq 66%

basics

~20 s

Fail open means the deploy proceeds when the decision service cannot answer; fail closed means it is rejected. Open protects delivery and lets unassessed changes through; closed protects the guardrail and can halt every deploy in the organisation.

open as a page

A build-time policy check verifies a container image's base image is on the approved list — what can it not establish?

level: juniorimportance: must knowfreq 61%

basics

~20 s

It establishes one fact: the recorded base image matches a list. It cannot establish whether an unapproved base was justified, whether the approved one was the right choice, or whether anything layered on top of it is safe.

open as a page

A policy rule can run in audit, warn or enforce mode - what does each do, and who does it reach?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Audit, warn and enforce are the same rule with different consequences: audit records the violation in a report for later review, warn returns a message to the requester but allows the change, enforce rejects the change outright.

open as a page

Where can a policy check run between a developer's editor and a running container, and what does each point see?

level: juniorimportance: must knowfreq 68%

basics

~20 s

A change passes several decision points: editor and pre-commit, pull request, build, artifact promotion, admission, and runtime. Each is handed a different document — raw source early, the fully rendered object at admission, the live workload last — so each can decide different rules.

open as a page

Why can a cloud provider's own built-in guardrail not be bypassed the way a policy engine you deploy can?

level: juniorimportance: must knowfreq 66%

basics

~20 s

A provider-native guardrail is evaluated inside the provider's own control-plane API while it authorizes the request. Every path reaches it — console, CLI, SDK, pipeline — and there is no component of yours to delete, skip or route around.

open as a page

What is the difference between a preventive and a detective policy control?

level: juniorimportance: must knowfreq 76%

basics

~20 s

A preventive control evaluates a change before it takes effect and can refuse it, so the non-conforming state never exists. A detective control inspects state that already exists and can only report what it finds.

open as a page

Why can a shape-matching policy rule not express 'every multi-replica workload needs a PodDisruptionBudget'?

level: middleimportance: must knowfreq 61%

basics

~20 s

A match document describes the shape of one resource. This requirement is a relation between two — for each multi-replica Deployment, some PodDisruptionBudget must select the same pods — and quantifying across sibling documents needs a logic language.

open as a page

In policy as code, how does a declarative match document differ from a rule written as a single expression?

level: juniorimportance: should knowfreq 52%

basics

~20 s

A match document states the shape a resource must have, and the engine compares the resource against that shape. An expression rule is a boolean over the resource's fields that the engine evaluates. Shape comparison versus computed verdict.

open as a page

How can a policy engine get a fact, such as an approved-region list, that the change never carries?

level: middleimportance: should knowfreq 58%

basics

~10 s

Three routes: replicate the fact into the engine ahead of time, have the caller pass it in with the change, or look it up live during evaluation. Each buys different freshness and fails differently.

open as a page

Why is a policy decision point's deny inert without the enforcer?

level: middleimportance: should knowfreq 47%

basics

~20 s

A decision point returns a message; it has no hands on the resource. The enforcer sits in the change's path and owes it three things: gather the input, apply the verdict including any correction, and never read a missing answer as approval.

open as a page

A backup-retention rule must run on a pre-merge plan and on a nightly scan of live databases. What makes one rule serve both?

level: middleimportance: should knowfreq 46%

basics

~20 s

Both surfaces must hand the rule the same fact in the same agreed shape, so each enforcement point normalises its input before evaluating. The rule must also state what it does when the plan leaves the value unknown until apply.

open as a page

What fields should a policy violation record carry beyond the human-readable message?

level: middleimportance: should knowfreq 48%

basics

~20 s

A violation record needs discrete fields other systems can read: a stable rule identifier, severity, the outcome, the resource identity, the offending field path, and a remediation hint. Prose cannot be filtered, aggregated or placed on a line.

open as a page

What middle designs sit between fail-open and fail-closed when a policy decision service is unavailable?

level: middleimportance: should knowfreq 48%

basics

~20 s

Split rules by what they protect: a small named subset stays fail-closed, everything else degrades to a static allow. Or answer from a last-known-good rule set with an expiry and a loud signal that the fallback is in use.

open as a page

An auditor asks what your approved-base-image gate actually establishes — what controls do you name beside it?

level: middleimportance: should knowfreq 44%

basics

~20 s

Name the gate as a preventive check on one recorded fact, the human review that judges whether a deviation was necessary, and runtime detection that watches what the image does once it runs. Say what each decides.

open as a page

A rule in warn mode fires on every namespace apply, yet violations keep landing - who is reading the warning?

level: middleimportance: should knowfreq 52%

basics

~20 s

A warning rides back on the response to the request that triggered it, so it reaches whichever client made the call. When that client is a CI job or a reconciler, the message lands in a machine's log.

open as a page

Why does a "no secrets in plain environment variables" standard need rules at two decision points?

level: middleimportance: should knowfreq 46%

basics

~20 s

Because the two rules read different documents. A source-text rule catches a secret typed literally into a file; only a check over the workload's resolved environment catches one that a reference, a pipeline or an injector puts there at start-up. Neither sees the other's case.

open as a page

What does a provider-native cloud control cost you compared with the same rule in your own policy engine?

level: middleimportance: should knowfreq 51%

basics

~20 s

You give up the pre-change verdict, the wording of the denial, the policy language itself, portability to another provider, and the ability to unit-test the rule locally in CI. You buy coverage and unbypassability with all of that.

open as a page

A pre-change policy gate has been green all quarter, yet a sweep of live resources finds violations — why?

level: middleimportance: should knowfreq 58%

basics

~10 s

A pre-change gate only judges changes routed through it. Resources created by hand in a console, by a third-party integration, or by a platform component never reach the gate and so never fail it.

open as a page

Your policy gate denied an unchanged plan at 09:00 and allowed it at 10:00 — what happened, and what do you change?

level: seniorimportance: should knowfreq 44%

basics

~20 s

The plan did not move; the facts around it did. Someone edited the approved list, or a live lookup answered differently. Pin the fact set to an identifiable revision and name it in every verdict.

open as a page

How do you review a policy rule you cannot follow well enough to confirm it matches the written standard?

level: seniorimportance: should knowfreq 45%

basics

~20 s

You do not approve it. Approval attests that the rule says what the standard says, and you cannot attest to logic you cannot follow. Ask for named intermediate steps and a case the rule must not flag.

open as a page

Your TLS 1.2 rule passes its tests but listeners still deploy at TLS 1.0 — where do you look?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Split the failure across the three roles. The rule tests clean, so suspect enforcement first — a deploy path with no guard, a verdict received and ignored, a swallowed error — then the input and the facts fed in.

open as a page

You are converting a 400-line shell check suite to policy rules. Which checks should stay scripts, and why?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Keep the procedural parts — paging a listing API, joining two systems, retrying, waiting — because they gather facts rather than decide anything. A pure comparison, such as retention of at least 35 days, is what becomes a rule.

open as a page

Your approved-base-image rule has grown a dozen special-case clauses and still blocks legitimate builds — what do you change?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Stop encoding judgement in the predicate. Split the rule: keep a hard block for the part that is mechanically decidable, and demote the rest to a signal that routes the build to a human review instead of failing it.

open as a page

Admission blocked a deploy for a rule an editor check could have caught days earlier — what do you change?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Add an earlier home for the same standard rather than weakening or moving the late one. The problem is feedback latency, not the rule: the developer paid three days for a verdict the editor could have given in a second. Keep the late check as the backstop.

open as a page

Encryption at rest is off on a data store created six hours ago that already holds data — what now?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Establish the exposure first: what was written, who could reach it, for how long. On most managed data stores encryption at rest is fixed at creation, so remediation is a migration, not a setting change.

open as a page

At 02:00 your policy decision service is down and every deploy is blocked - do you flip the default to allow?

level: principalimportance: should knowfreq 38%

basics

~20 s

Only as a narrowly scoped, time-boxed change that expires by itself, and never for the small set of rules whose whole purpose is stopping unassessed privilege changes. Restoring service or serving last-known-good rules is the better move.

open as a page

showing 1–30 of 39