skip to content

A change-scoped Terraform plan gate never sees a legacy database. How do you close that gap?

level: principalimportance: should knowfreq 44%

answer

  1. the gate is not an inventory
  2. detect where you cannot afford to block
  3. same logic, two different consequences
  4. a finding needs an owner and a date
  5. widen only after the debt is gone

basics

~20 s

Keep the blocking rule scoped to changed resources, and add a scheduled evaluation over the full recorded state that reports rather than blocks. Give each finding an owner and a date; widen the blocking rule once the backlog is drained.

solid answer

~50 s

The blindness is deliberate: a rule reading `resource_changes` and skipping `no-op` entries blocks engineers only for what their change does, which is what makes it survivable. Widening it to `planned_values` — the whole post-apply state — fixes the gap and immediately blocks every team on the oldest database in the estate, which they cannot fix inside their pull request. So split the job. The preventive rule stays change-scoped and blocking, with a before-versus-after comparison so existing values cannot get worse. A detective run on a schedule applies the same logic to the full recorded state and produces a report: address, current value, owning team, date. Findings go to the team that declares the resource, not a central queue, and the trend is what you publish. When the backlog reaches zero, point the blocking rule at the whole state so the estate cannot regress, and retire the comparison.

go deeper

for a junior

Understand that a gate reading a change only sees the resources in that change, so infrastructure nobody edits is never checked by it.

for a middle

Explain the mechanism behind the blind spot: untouched resources appear as no-op entries and are filtered out, while the whole-estate view lives in a different section of the plan document.

for a senior

Design the second mechanism. Be concrete about running the same rule logic on a schedule over the full recorded state, reporting rather than blocking, and freezing the problem with a no-regression comparison meanwhile.

for a principal

Own the sequencing and the politics: why blocking people for other teams' debt destroys a policy programme, how findings get an owner and a date, and what has to be true before you widen enforcement to the whole estate.

## The blind spot is a design decision, not a bug A policy rule that iterates a plan's `resource_changes` and skips `no-op` entries evaluates exactly the resources the run touches. That is deliberate, and it is what makes the gate politically survivable: it blocks an engineer only for what their change does. The cost is precise and predictable — a database created before the policy existed, sitting at one day of backup retention with deletion protection off, appears in every plan as a `no-op` and is never judged. The gate will be green forever, and the resource will stay wrong forever. The instinct is to fix this by widening the rule to `planned_values`, the section that holds the whole post-apply state including untouched resources. Resist it, at least at first. On the day you flip that switch, the next pull request from any team is blocked by the oldest database in the estate, owned by somebody else, unrelated to the diff, and unfixable inside that change. Engineers cannot ship, the gate acquires a reputation as an obstacle rather than a guardrail, and you spend your credibility arguing about a database you actually do want fixed. Blocking people for other people's debt is how a policy programme dies. ## Separate the preventive job from the detective job The workable shape is two mechanisms with different jobs and different failure modes. **Preventive, change-scoped, blocking.** Keep the rule reading `resource_changes`, matching `create` and `update`, and denying. It guarantees the debt stops growing: every new database meets the standard, and every edited one is held to it. Add a no-regression comparison of `change.before` against `change.after` so an existing sub-standard value cannot be lowered further while you work. **Detective, estate-wide, reporting.** Run the same rule logic on a schedule against the full recorded state — the `planned_values` tree of a plan with no pending changes serves this well, since it enumerates every managed resource with its current attributes. Crucially this run **reports**; it does not block anybody. Its output is a list of addresses with a value and an owner, not a failed pipeline. The two rules should share their logic, so the standard cannot drift apart between what is blocked and what is reported. ## Turning a report into remediation A backlog nobody owns is a spreadsheet, not a plan. What makes this work in an organisation: - **Attribute every finding to a team**, using the module address and repository the resource is declared in — not to a central security queue. - **Give the list a shape and a deadline**, and size the deadline against the actual cost of the fix. Extending a backup window is cheap; some remediations require a maintenance window or a replacement, and those need scheduling rather than nagging. - **Publish the trend, not the total.** A count that only goes down is the evidence that the effort is working, and it is what you show when someone asks whether the programme is worth its cost. - **Handle the resources nobody will fix.** Some are scheduled for decommission; some belong to a system frozen for a migration. Record the decision and its owner rather than leaving the finding to rot in the report, and re-examine it on a date. ## The end state When the backlog reaches zero, widen the rule. Point the blocking rule at the full state so the estate cannot regress at all, and retire the no-regression comparison, which now protects nothing. The order matters: change-scoped enforcement first while the estate is dirty, whole-estate enforcement once it is clean. Doing it the other way round is technically stricter and practically self-defeating. ## What people get wrong - Widening to the whole estate on day one and blocking every team on the oldest resource in it. - Treating a green change-scoped gate as evidence the estate conforms; it only shows that nothing new broke the rule. - Leaving untouched legacy infrastructure permanently out of scope because no mechanism was built to see it. - Buying adoption by dropping enforcement entirely, so the debt grows while the report gets longer. - Building the detective sweep and never assigning an owner to what it finds.

  • Why not simply point the same rule at planned_values and be done with it?
    Because planned_values holds the whole post-apply state, so every pull request is judged on resources it never touched. The first team to hit it is blocked by another team's three-year-old database, which they cannot fix inside their change. The rule is stricter and the practical outcome is worse: people route around the gate.
  • How do you keep the debt from growing while you drain it?
    The change-scoped rule already stops new violations, because it enforces the standard on every create and every replacement. Add a comparison of `change.before` against `change.after` on in-place updates so an existing sub-standard value cannot be lowered further. Together they freeze the problem at its current size while remediation runs.
  • What would make you finally widen the blocking rule to the whole estate?
    The backlog reaching zero. At that point nothing is below the line, so whole-estate enforcement blocks nobody on arrival and makes regression structurally impossible; the no-regression comparison becomes redundant and can be retired. Doing it in the other order is technically stricter and self-defeating.
  • What do you do with resources nobody is going to fix?
    Record the decision and its owner rather than letting the finding rot in the report — a system frozen for a migration, a service scheduled for decommission. Attach a date to re-examine it. An unmanaged permanent finding trains everyone to ignore the report, which is worse than the original violation.

saying these in an interview costs you the question

  • Widens the rule to the whole estate and blocks every team
  • Treats a green change-scoped gate as proof the estate conforms
  • Leaves untouched legacy resources permanently out of scope
  • Drops enforcement entirely to buy adoption
  • Builds the detective sweep and assigns no owner to findings

context