skip to content

Policy as Code

Policy-as-code moves 'no public storage buckets' off a wiki page and into a check that fails the pipeline. Interviewers ask it to see whether I can enforce guardrails without becoming a manual approval bottleneck.

on this pageshow

questions

6

Policy checks in an infrastructure-as-code pipeline usually run against the proposed plan or diff rather than against the resources that are already running. Why is evaluating the proposed change the decisive design choice, and what does each of the two placements catch that the other misses?

level: middleimportance: must knowfreq 62%

answer

  1. before it exists versus after it exists
  2. preventive on the diff, detective on the estate
  3. the exposure window is the risk
  4. a diff has direction; a snapshot does not
  5. its authority stops at the pipeline

basics

~20 s

Evaluating the proposed change is preventive: the violation is judged before it exists, so the bad resource is never created. Scanning what is already running is detective — it finds violations only after they exist and after the exposure window has opened.

solid answer

~50 s

A policy that reads the plan sees a change that has not happened yet, so failing it means the misconfiguration is never created at all — no exposure window, no cleanup, no incident write-up. It also sees the whole proposed change set at once, including deletions, so it can reason about things a snapshot of live infrastructure cannot: that a database is about to be replaced, or that a rule is being relaxed. The cost is that its authority stops at the pipeline. It never sees a console edit, a change made by another team's tooling, or a violation that predates the rule, and it cannot judge attributes whose values are only known after apply. So the two placements are complementary rather than alternatives: preventive checks on the diff keep new violations out, and a periodic detective sweep of the live estate tells you what is actually there. Anyone who claims one replaces the other has only half a control.

go deeper

for a junior

Know the two words and which is which: preventive stops a bad change before it is applied, detective finds a bad resource that already exists. Be able to say why stopping it earlier is cheaper.

for a middle

Explain the mechanics: the check reads the computed diff, so it can fail a change that has not happened, and it sees deletions and relaxations that a snapshot of live resources cannot express.

for a senior

Demonstrate the blind spots without prompting — bypass paths, pre-existing violations, attributes unknown until apply — and describe the ratchet where preventive checks hold the line while a detective sweep works down the backlog.

for a principal

Own the coverage argument. Be able to state what fraction of production actually flows through the gated path, what the plan for the ungated remainder is, and how you would defend the claim "this control is effective" to someone who does not take the pipeline on trust.

## Two places a rule can live The same rule — say, *no storage bucket may be publicly readable* — can be enforced in two very different positions: 1. **Against the proposed change.** The IaC tool computes what it is about to do, and the policy engine evaluates that description before anything is applied. This is a **preventive** control. 2. **Against what is running.** Something periodically reads the live resources from the provider's APIs and reports which ones violate the rule. This is a **detective** control. Both are legitimate. The reason the plan-time position is the one policy-as-code programmes are built around is that it is the only one that can stop the violation from ever existing. ## What the preventive placement buys **No exposure window.** A detective control, however fast, is still a race: the bucket is public between the moment it is created and the moment someone reacts. If the rule is about data exposure, that window is the entire risk. A plan-time failure closes it to zero. **Cheap remediation.** Failing a change costs an engineer a commit. Fixing a running resource costs a change request, a maintenance window, and sometimes data movement — and, if the resource has been in use, an incident review of what happened while it was wrong. **Feedback where the mistake was made.** The failure lands on the engineer who wrote the line, in the context of the change, minutes after they wrote it. A finding on a dashboard three days later lands on whoever owns the dashboard. **The whole change set is visible at once.** This is the subtle advantage. A snapshot of live infrastructure can only see what *is*. The proposed change also contains what is about to *stop being*: a rule can fail a change that deletes a production database, that removes a logging configuration, or that widens an existing network rule — none of which is expressible as a property of the current estate. A diff carries direction; a snapshot does not. **The change is evaluated as a unit.** A live scan sees each resource independently. The plan can express a relationship across the set: this network exposure is being added *and* the accompanying restriction is not. ## What the preventive placement cannot see This is the half candidates forget, and it is where the interview usually goes. - **Anything outside the pipeline.** Console edits at 3am, another team's tooling, an autoscaler, an incident responder with admin rights — none of it passes your check. The plan-time control has authority over exactly one path to production. - **Pre-existing violations.** A new rule applies to new changes. Resources created before it existed sit there, untouched, until something happens to modify them. Many teams discover only during an audit that their coverage is "every change since March" rather than "the estate". - **Values not known until apply.** Some attributes are computed by the provider — generated names, allocated addresses, resolved identifiers. In the plan they are marked unknown, and a rule that depends on one cannot decide. The honest options are to fail closed (safe but noisy), to require the value be set explicitly in code, or to move that particular rule to the detective side. - **Runtime reality.** Configuration is not behaviour. A plan can prove that logging is enabled; only observation shows that logs are arriving. - **The state of the world after apply.** A plan is a proposal. An apply can fail halfway, leaving something the policy never approved. In practice, teams close this by making the applied change the same artifact that was evaluated, rather than re-deriving it. ## How the two are combined The common shape is a ratchet: preventive checks on every change stop the number of violations from growing, and a periodic detective sweep of the live estate measures the backlog and shrinks it. The preventive side gets the rules that are cheap to evaluate and unambiguous; the detective side gets the rules that need real-world values, plus the job of catching everything that never went through the pipeline at all. A good answer names both, says which risk each retires, and does not pretend that a green pipeline is a statement about the estate.

  • A rule needs an attribute whose value the provider only assigns during apply, so the plan shows it as unknown. What do you do?
    Three honest options. Fail closed and require the engineer to set the value explicitly in code, which turns an undecidable rule into a decidable one. Accept the unknown at plan time and cover that rule with a detective check after apply. Or narrow the rule to a property that *is* known, if one carries the same risk. What you must not do is silently pass unknowns and describe the rule as enforced.
  • Why can a plan-time check fail a change that a scan of the live estate could never flag?
    Because the plan contains destructive and subtractive actions. Deleting a production database, removing an audit-log configuration, or widening an existing firewall rule are all properties of the *transition*, not of the resulting state. A snapshot taken afterwards sees only the outcome, and often a perfectly compliant-looking one.
  • Your policy check passes on the plan, but the apply half fails and leaves resources behind. Has the control done its job?
    Only partly. The check certified a proposal, not an outcome. Teams close the gap by applying the exact artifact that was evaluated rather than recomputing the change, and by keeping a detective sweep that reads the live estate — which is also what catches partially applied changes that no one noticed.

saying these in an interview costs you the question

  • Claiming a green policy pipeline proves the running estate is compliant
  • Treating live scanning as redundant once plan-time checks exist
  • Forgetting that resources created before the rule existed were never evaluated
  • Assuming the check applies to console changes and other teams' tooling
  • Passing resources whose relevant attribute is unknown at plan time and calling the rule enforced

context

open as a page

You are introducing automated policy checks to teams that already ship infrastructure changes daily. What enforcement levels sit between "record the finding" and "block the apply", and how would you sequence the rollout so the guardrails survive contact with delivery pressure?

level: seniorimportance: must knowfreq 45%

basics

~20 s

Three levels are standard: advisory reports only, soft-mandatory fails but a named owner can override with the override recorded, and hard-mandatory cannot be overridden at all. Roll out advisory first to find false positives and the real violation rate, then promote rules one at a time.

open as a page

In infrastructure-as-code work, what does "policy as code" mean, and what does expressing a rule such as "no storage bucket may be publicly readable" as an executable check give you that the same rule written in a standards document does not?

level: juniorimportance: should knowfreq 48%

basics

~20 s

Policy as code expresses organisational rules as machine-executable checks that run automatically on every proposed infrastructure change. Unlike a written standard, the rule is unambiguous, never forgotten by a reviewer, version-controlled like the code it governs, and leaves a pass/fail record.

open as a page

A policy engine such as Open Policy Agent or HashiCorp Sentinel is normally fed a machine-readable description of the change an infrastructure tool intends to make, rather than the source files the engineer wrote. Why do teams evaluate that representation instead of the source, and what is it still unable to tell the engine?

level: middleimportance: should knowfreq 38%

basics

~20 s

Source files are a template, not an outcome — variables, loops and module inputs are only resolved when the tool computes the change. The engine is given that resolved description so it judges the concrete resources and values that will exist, not the text that produces them.

open as a page

Any set of infrastructure guardrails eventually meets a change that legitimately has to break one of the rules. How would you design the exemption path so that exceptions remain possible without the guardrail degrading into a formality?

level: principalimportance: should knowfreq 32%

basics

~20 s

Make exemptions explicit, narrowly scoped to one rule and one resource, time-bounded with an expiry that re-fails, attributed to a named requester and an independent approver, and visible in aggregate. Blanket skips and permanent unowned exceptions are what hollow out a guardrail.

open as a page

An auditor asks you to demonstrate that no publicly readable storage bucket was created in your production environment over the last twelve months. What can a policy-as-code pipeline offer as evidence, and what does a year of green policy runs genuinely not prove?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

A policy pipeline evidences that every change passing through it was evaluated against named, versioned rules, with verdicts, overrides and approvers recorded per commit. It proves nothing about changes made outside that path, or about rules you never wrote.

open as a page