skip to content

In infrastructure-as-code work, what does "policy as code" mean, and what does expressing a rule such as "no storage bucket may be publicly readable" as an executable check give you that the same rule written in a standards document does not?

level: juniorimportance: should knowfreq 48%

answer

  1. rules that execute, not rules that are read
  2. a wiki page cannot fail a build
  3. ambiguity has to be resolved to be coded
  4. sampling versus every single change
  5. versioned, reviewable, leaves a record

basics

~20 s

Policy as code expresses organisational rules as machine-executable checks that run automatically on every proposed infrastructure change. Unlike a written standard, the rule is unambiguous, never forgotten by a reviewer, version-controlled like the code it governs, and leaves a pass/fail record.

solid answer

~50 s

Policy as code means the rule stops being a sentence humans are supposed to remember and becomes a program that returns a verdict on a specific proposed change. Three things change. First, ambiguity is forced out: a document can say "no public buckets" and leave *public* undefined, while an executable rule has to name exactly which attribute combinations fail, and that definition is itself reviewable. Second, coverage goes from sampled to total — a reviewer catches what they happen to notice on the changes they happen to read, whereas the rule evaluates every change, every time, including the Friday-evening one. Third, the rule inherits the software lifecycle: it is versioned, reviewed, tested, rolled out gradually, and it emits a record you can point an auditor at. It does not replace the cloud provider's own permission controls; it is an earlier, cheaper layer in front of them.

go deeper

for a junior

Be able to define the term plainly: the rule is a program that runs on every change, not a paragraph someone is supposed to remember. Give one concrete example rule and say what happens when it fails.

for a middle

Be ready to explain why an executable rule is different in kind, not just in convenience: it forces a precise definition, it evaluates every change rather than a sample, and it is versioned and testable like any other code.

for a senior

Show you know where the check has no authority — anything applied outside the pipeline — and that you pair it with provider-side enforcement and periodic checking of the live estate rather than trusting one layer.

for a principal

Own the question of who writes the rules, how a team disputes one, and how the policy set is kept honest over years. A rule that half the estate quietly works around is a design problem in the policy programme, not a discipline problem in the teams.

## The idea in one line An organisation always has rules about its infrastructure: production data must be encrypted, nothing may be exposed to the public internet without sign-off, every resource must carry an owner tag, no instance type above a certain size without approval. **Policy as code** is the practice of writing each of those rules as an executable check — a small program that is given a description of a proposed change and returns *pass* or *fail*, usually with a message naming the offending resource. The alternative it replaces is not "nothing". It is a wiki page, an onboarding slide, a checklist in a pull-request template, and a senior engineer who remembers. Those work exactly as well as human attention scales, which is to say they work until the team grows, the reviewer is on holiday, or the change is urgent. ## What actually changes when a rule becomes a program **Ambiguity is forced out.** "No publicly readable buckets" is a comfortable sentence precisely because it is vague. Does a bucket fronted by a CDN count? What about one with a permissive policy but a block-public-access setting on top? A prose rule can survive that vagueness for years; an executable rule cannot even be written until someone decides. The decision then lives in a file that can be read, argued with, and changed deliberately. **Coverage becomes total rather than sampled.** Human review is sampling: a reviewer scans a diff and notices what stands out. An automated rule is population testing — it sees every resource in every change. This is the difference that matters most in practice, and it is why the same rule catches things reviewers had waved through for months when it is first switched on. **The rule gets a lifecycle.** Because the policy is a file in version control, you can review a change to the rule, test it against known-good and known-bad examples, roll it out to one team before all of them, and answer "when did this start applying and who approved it". A wiki page has none of that; it has a last-edited timestamp and an anonymous edit. **It produces evidence.** Every evaluation leaves a record tied to a specific change: which rules ran, which version of them, what the verdict was, who approved any exception. That record is worth far more to an auditor than an attestation that a standard exists. ## What policy as code is not It is *not* the cloud provider's own permission system. A provider-side permission decides whether a caller may perform an action at all, at the moment of the API call, regardless of which tool made it. A policy-as-code check runs inside your delivery path and can be bypassed by anyone acting outside that path — clicking in a console, or running the tool from a laptop. The two are complementary: policy as code gives fast, specific, early feedback with a good error message; provider-side permissions give an enforcement boundary that cannot be walked around. A serious platform uses both, and treats a policy check as a way to stop mistakes rather than as a way to stop a determined insider. It is also not a linter for code style, and not a substitute for periodically checking what is actually running. It judges what a change *proposes*, which is a different question from what the estate *contains*. ## The common failure modes - **Rules written only from incidents.** The policy set becomes a scar tissue of last year's outages and never expresses the standard as a whole. - **False positives.** A rule that fails legitimate changes destroys trust faster than any amount of coverage buys it, and teams start looking for ways around the check. - **No owner.** A policy set that nobody maintains ages into a set of rules people mechanically work around. - **Unactionable messages.** "Policy violation: rule 47" wastes an engineer's afternoon. A good failure names the resource, the rule, the reason, and the fix. The useful mental model is that a policy check is a very cheap, very consistent reviewer who reads every line of every change, knows exactly one thing, and never gets tired of saying it.

  • If the cloud provider can already deny the action at the API, why bother with a policy check in the pipeline at all?
    Because they answer different questions at different costs. A provider-side denial happens at apply time, on a partially built change, with an error message about a request rather than about your code. A pipeline check fails before anything exists, points at the exact line, explains the rule, and can encode organisational rules the provider has no concept of — naming conventions, cost ceilings, required tags. Use both: the pipeline for feedback, the provider boundary for enforcement.
  • Who should own the policy set — the platform team, the security team, or the product teams it applies to?
    Security typically owns the intent, the platform team owns the implementation and the rollout, and product teams must have a real path to challenge a rule. The arrangement that fails is one where the group writing the rules never has to pass them; they lose the feedback loop that tells them a rule is wrong rather than a team being sloppy.

It is the difference between a sign on the lab door saying "no food inside" and a door that will not open while you are holding a sandwich. The sign relies on everyone reading and remembering; the door does not.

saying these in an interview costs you the question

  • Calling it policy as code when the rules still live in a document
  • Assuming the check enforces anything against someone bypassing the pipeline
  • Treating the policy set as write-only, with no owner or tests
  • Believing a passing check means the running estate is compliant

context