skip to content

It is 02:00 and every apply in HCP Terraform is failing at the policy check — how do you diagnose it?

level: seniorimportance: nice to knowfreq 33%

answer

  1. scope tells you where to look
  2. one change cannot break every workspace
  3. denial versus failed evaluation
  4. you cannot run it on a laptop unaided
  5. download the run's mock data

basics

~20 s

Read the policy check output first: a failure across every workspace at once points at the policy set or its inputs, not at anyone's change. Then reproduce offline by downloading a failed run's mock data and evaluating the policy locally.

solid answer

~50 s

Start from the blast radius. One team's change can violate a rule; it cannot break every workspace at once, so the cause is on the policy side — a policy set version just published, a parameter changed, a scope widened, or a plan whose shape moved with a Terraform or provider upgrade so an attribute the rule indexes is gone. Read the policy check output and separate a real denial, which names concrete offending values, from an evaluation that collapsed into undefined. Then get a fast loop: this rule only executes inside the platform, so download a failed run's mock data — its plan, state and configuration — and evaluate the policy locally with the Sentinel CLI until you can read the trace. Treat those mocks as sensitive; they carry the run's real values. Revert the policy set to stop the bleeding, then fix forward.

go deeper

for a junior

Know that the policy check is a distinct stage of a run and that its output names which policy failed. Being able to read that output and report it accurately is the expectation here.

for a middle

Be ready to explain the difference between a rule denying a change and a rule failing to evaluate, and to name what inputs a local reproduction would need.

for a senior

Demonstrate the triage: reason from blast radius, classify the failure, reproduce offline against downloaded mock data, mitigate by reverting the policy set, then fix forward and verify against several runs.

for a principal

Own the operational consequence: the gate is a production dependency, so how policy changes reach every workspace, and how quickly they can be rolled back, is a design decision you are accountable for.

## The gate is the outage A policy check that blocks the wrong change is an inconvenience. A policy check that blocks *every* change is an outage, and at two in the morning it is your outage rather than the developers'. ### Step one: read the blast radius, not the rule The single most informative fact is scope. If runs are failing across unrelated workspaces owned by unrelated teams, the proposition "they all wrote a violating change tonight" is not credible. Something on the policy side moved. The candidates, roughly in order of likelihood: - **A policy set version was published.** Someone merged a rule change hours ago and it reached the platform on its own schedule. - **A policy parameter or its value changed** — a list of approved regions, a threshold — so a rule that was satisfiable is now not. - **The policy set's scope widened**, attaching rules to workspaces that were never expected to pass them. - **The plan's shape changed.** A Terraform or provider upgrade renamed, restructured or stopped emitting an attribute the rule indexes. The rule did not change; its input did. This is the one that surprises people, because the policy repository's history is clean. ### Step two: is it a denial or a failure to evaluate? Read the policy check output on a failing run and classify it. A genuine denial names concrete values: this address, this region, this tag. An evaluation that fell apart looks different — the trace shows a rule that came out undefined, or an error, with no offending value to point at. The two demand opposite responses. A genuine denial across the fleet means the rule's intent is now wrong or too broad. An undefined result usually means the rule is reading something that is not there in the current plan shape, which is a bug in the rule rather than a judgment about anyone's infrastructure. ### Step three: get a loop outside the platform This is the part that makes the leaf what it is. A rule written against the managed run's inputs cannot simply be run on your laptop against a plan file, because it reads documents the platform supplies — prior state, run metadata, the cost estimate — and it is evaluated by the platform's policy runtime. The supported escape hatch is mock data. From an affected run, download the generated mock files reproducing that run's plan, state and configuration data. Point a local policy test configuration at them and evaluate the policy with the Sentinel CLI. You now have the trace, sub-second iteration, and a way to confirm a fix without publishing a policy set and waiting for someone's run. Two cautions come with it. Mocks contain the run's real values — resource names, addresses, attribute values, possibly more — so they are sensitive material, and downloading them is a permissioned action rather than something to paste into a shared channel. And a mock is a snapshot of *one* run: a fix verified against a single workspace's data can still fail on another whose plan contains an attribute yours did not. Pull mocks from two or three affected workspaces before you declare it fixed. ### Step four: stop the bleeding, then fix forward Diagnosis and mitigation are separate. If a policy set version broke the fleet, reverting to the previous version restores service immediately and buys you daylight to fix the rule properly; narrowing the policy set's scope is the same move at smaller granularity. Neither requires you to reason about whether tonight's blocked changes are safe. What you should resist is editing the rule live against production runs. Each iteration costs a full plan cycle for somebody, the feedback is a policy check result rather than a trace, and it is how a two-hour incident becomes a six-hour one. ### Afterwards The durable output of the night is the mock data from the run that broke, kept alongside the rule, so the same plan shape is exercised the next time the rule changes. The second durable output is knowing which of the four causes above it was, because "the plan shape changed under us" implies watching provider upgrades, and "a policy set shipped straight to every workspace" implies something about how policy changes reach production.

  • What in the policy check output tells you the rule errored rather than legitimately denied?
    A real denial cites concrete values — the address, the attribute, what it found versus what it required. An evaluation failure shows a rule that resolved to undefined, or an error, with nothing specific to point at. If the output cannot tell you which resource offended, suspect the rule before you suspect the change.
  • Why can't you just run the rule locally against a plan JSON file?
    Because the rule reads inputs the platform supplies alongside the plan — prior state, run and workspace metadata, the cost estimate — and it is evaluated by the platform's own policy runtime. A plan file on its own is one of several documents the rule expects, which is why the mock download exists.
  • The policy set genuinely did not change. What else could have?
    The plan's shape, most often after a Terraform or provider upgrade changed or removed an attribute the rule indexes. Also a policy parameter's value, a workspace newly brought into the policy set's scope, or a workspace whose data now exercises a path in the rule that no earlier run reached.
  • How do you confirm a fix before publishing it?
    Evaluate the corrected policy locally against mock data pulled from several affected runs, not one, and against a run that was passing before, so you can see that the fix does not quietly stop the rule from matching anything. A rule that now denies nothing is not a fixed rule.

saying these in an interview costs you the question

  • Blames the developer's change without checking scope
  • Edits the rule live against production runs
  • Assumes a failing policy means the plan is unsafe
  • Cannot say how to evaluate the rule outside the platform
  • Verifies the fix against a single workspace's data
  • Pastes downloaded mock data into a shared channel

context