skip to content

A policy engine such as Open Policy Agent or HashiCorp Sentinel is normally fed a machine-readable description of the change an infrastructure tool intends to make, rather than the source files the engineer wrote. Why do teams evaluate that representation instead of the source, and what is it still unable to tell the engine?

level: middleimportance: should knowfreq 38%

answer

  1. source is a template, outcome is data
  2. the deciding value often is not in the file
  3. loops expanded, references resolved, defaults merged
  4. plain JSON, so a generic engine can read it
  5. some attributes are unknown until apply

basics

~20 s

Source files are a template, not an outcome — variables, loops and module inputs are only resolved when the tool computes the change. The engine is given that resolved description so it judges the concrete resources and values that will exist, not the text that produces them.

solid answer

~50 s

The source is one level of indirection away from the truth. The same file can produce a compliant result in one environment and a violation in another, because the deciding value arrives through a variable, a module input, or an expression evaluated at run time. A rule that reads the source would have to interpret the language to know, and is trivially evaded by moving the offending value behind an indirection. The computed change is already resolved: loops expanded, defaults merged in, references dereferenced, and each entry carries the action and the resulting attribute values. That is a plain data structure, which is why engines like Open Policy Agent can evaluate it generically as JSON without knowing anything about the tool that produced it. What it cannot show is anything the tool does not yet know — values the provider computes during apply appear as unknown — plus resources this tool does not manage, and any intent behind the change.

code

json · 13 lines
json
{
  "resource_changes": [
    {
      "address": "module.web.aws_s3_bucket.assets",
      "type": "aws_s3_bucket",
      "change": {
        "actions": ["create"],
        "after": { "bucket": "acme-assets-prod", "force_destroy": false },
        "after_unknown": { "arn": true, "id": true }
      }
    }
  ]
}

go deeper

for a junior

Know that the check runs on what the tool has decided to do, not on the text of the file, and that this is why it still catches a value that arrived through a variable.

for a middle

Explain the resolution step: variables substituted, loops expanded, modules flattened, defaults merged — leaving plain structured data that a general-purpose engine can evaluate without understanding the tool.

for a senior

Be the person who raises unknown-at-apply values unprompted, and who says what they do about a rule that depends on one rather than pretending the rule is enforced.

for a principal

Take a position on engine choice and its blast radius: one general engine and rule language across infrastructure, admission and service authorisation, versus a per-tool native policy layer that is simpler today and fragments your rule estate over five years.

## Two candidate inputs When you decide to enforce a rule automatically, there are two obvious things to point the rule at: - the **source files** the engineer edited, or - the **computed change** the tool produces when it works out what it is going to do. Almost every serious policy-as-code setup chooses the second, and understanding why is the mechanism half of this topic. ## Why the source is the wrong input Infrastructure source is a *template*. The value that decides whether a rule passes very often is not in the file at all: - it arrives as an input variable, set differently per environment; - it is produced by an expression — a conditional, a lookup, a string built from other values; - it comes from a shared module whose source you are not even reading; - it is inherited from a default the tool merges in without anyone writing it down. So a rule reading source would have to reimplement the configuration language's evaluator to know what the file *means*, and it would still be wrong whenever a value came from outside the file. Worse, it is evaded by accident: an engineer refactors a literal into a variable, and the rule silently stops matching. The pattern generalises — a check on text has to guess; a check on the resolved outcome does not. ## What the resolved change looks like By the time the tool has computed the change, it has done the interpretation for you. Every element of the proposed change set is enumerated concretely: which resource, of which type, what action is being taken on it (create, update in place, replace, destroy), and the attribute values it will have afterwards. Loops have been expanded into individual entries; module boundaries have been flattened into addresses; references have been resolved to values wherever the tool already knows them. Crucially, that is now *data* — nested objects and arrays, with no language semantics left in it. That is the property the engine shapes exploit: - **Open Policy Agent** is a general-purpose engine that evaluates rules written in its Rego language against arbitrary JSON input. It has no idea what infrastructure is; you hand it the change document and it applies your rules to it. The same engine is therefore used to police completely unrelated systems. - **Sentinel** is HashiCorp's own policy language and framework, embedded in their platform, which exposes the plan, the configuration and the recorded state to a policy as structured objects. - **Provider-native policy services** work on the same principle from the other side: the rule is evaluated against a structured description of a resource rather than against the file that declared it. The common shape is: *tool produces structured description of intent → engine evaluates rules against that description → verdict, with a message naming the offending element.* ## What the resolved change still cannot tell you **Unknown values.** Some attributes are assigned by the provider only during apply — generated names, allocated network addresses, identifiers of resources being created in the same run. The change description marks these as not yet known. A rule that depends on one cannot return an honest verdict; you must either force the value to be set explicitly, fail closed, or move that rule to a check that runs against reality afterwards. **Everything this tool does not manage.** The document describes one tool's proposed change. Resources created by another pipeline, by hand, or by an autoscaler are simply absent — as are resources this tool manages but is not touching in this run, unless the engine is also given the recorded state. **Intent.** A rule can see that a bucket policy permits anonymous reads. It cannot see that this bucket is the public marketing site and that permission is the entire point. This is why exemptions exist, and why a rule that cannot distinguish a legitimate case from an illegitimate one accumulates them. **Behaviour.** Configuration is not runtime. That logging is configured is visible; that logs actually arrive is not. ## The interview point The answer to give is: **evaluate the outcome, not the text that produces it** — because the text is a template whose meaning depends on inputs, and because the resolved outcome is plain data that a general-purpose engine can reason about. Then name the limits, because that is what separates someone who has run this from someone who has read about it.

  • Why does Open Policy Agent work for infrastructure policy despite knowing nothing about infrastructure?
    Because its input is just JSON. OPA evaluates Rego rules against an arbitrary document, so any system that can describe its proposed change as structured data can be policed by it. That decoupling is the point: one engine, one rule language and one set of testing habits cover infrastructure changes, admission decisions and API authorisation alike, instead of a different bespoke rule dialect per tool.
  • An engineer moves a hardcoded value into a variable and a policy that used to catch it stops firing. What went wrong?
    The rule was reading the source text rather than the resolved change. Text-matching rules are brittle against exactly this refactor, and the failure is silent — the check still reports green. Evaluating the computed change avoids it entirely, because by then the variable has been substituted and the engine sees the value that will actually exist.
  • Can a policy engine evaluate rules that span several resources in one change, such as "any public-facing load balancer must have logging enabled"?
    Yes, and that is a real advantage of evaluating the change set as one document. The engine has all the entries at once and can correlate them by address or by reference. The caveat is that a resource involved in the rule but not being modified may not appear in the diff, so cross-resource rules usually need the recorded state supplied alongside it.

saying these in an interview costs you the question

  • Writing policies that grep the source files for forbidden strings
  • Assuming every attribute value is known at plan time
  • Believing the engine sees resources the tool does not manage
  • Confusing the engine (generic, evaluates data) with the rules you write for it

context