skip to content

After an incident you fixed production with `terraform apply -target=module.db`. The next untargeted `terraform plan` proposes changes to resources nobody edited. What explains that, and how do you handle the plan?

level: seniorimportance: should knowfreq 40%

answer

  1. the partial run left work behind
  2. dependents were never in the run
  3. refresh only covered the target
  4. outputs can be stale for neighbours
  5. sort the diff before approving anything

basics

~20 s

A targeted apply evaluates and refreshes only the targeted subgraph, so the first untargeted plan afterwards shows the catch-up work: dependents that never picked up new attribute values, stale outputs, and real-world drift no run had refreshed since.

solid answer

~50 s

Nothing mysterious happened — the targeted run was partial by design. It included `module.db` and its dependencies, but not the resources that *depend* on the database, and it refreshed only the objects it was working with. So the first full plan is where three things surface at once: consumers of the database's attributes that need updating, root outputs left stale because they derive from untargeted resources, and genuine drift on the rest of the estate that no run has looked at since before the incident. The handling is to read it rather than approve it. I classify every diff into consequence-of-my-change, drift someone made in the console during the incident, and config merged while we were firefighting, then apply untargeted in a normal change window. What I do not do is run more targeted applies to make the diff go away.

code

bash · 4 lines
bash
terraform plan -out=cleanup.tfplan
terraform show cleanup.tfplan
terraform apply cleanup.tfplan
terraform plan

go deeper

for a junior

Understand that an apply limited to some resources leaves the rest untouched, so the next complete run will show whatever was skipped. It is expected behaviour, not a bug.

for a middle

Explain the three sources precisely: excluded dependents, refresh limited to the targeted objects, and outputs left stale — and why the first untargeted plan is where they all appear together.

for a senior

Show the triage. Save the plan, read every change, separate consequences of your fix from console drift and from merged config, and refuse to auto-approve a cleanup plan in a moving estate.

for a principal

Make the obligation structural: targeted applies are logged break-glass actions with a named owner and a same-day follow-up to an empty plan, so the reconciliation debt never lands unannounced on the next routine apply.

## Why the diff exists A targeted run reduces Terraform's graph to the addresses you named plus their dependencies. Three consequences follow, and together they account for essentially every surprise in the next full plan. **Dependents were excluded.** Targeting includes what your target depends on, never what depends on it. If the database was replaced or its endpoint changed, the application configuration, the DNS record, the security-group rule and the parameter-store entry that reference it were all outside the run. They still hold the old value in state, so the full plan proposes to update them now. **Refresh was narrowed.** Terraform read real-world attributes only for the objects in the reduced graph. Anything else in the state has not been compared against reality since the last full run — which, during a long incident with people in the console, can be a lot of accumulated drift arriving all at once. **Outputs may be stale.** Root module outputs that depend on untargeted resources are not necessarily updated by a targeted apply — Terraform's own post-run warning says as much. That matters beyond cosmetics: if another configuration reads those outputs through remote state, it has been consuming a stale value since the incident. ## Reading the plan instead of approving it The correct first move is to produce the plan and actually read it, never `-auto-approve`. Sort every proposed change into one of three buckets: 1. **Consequence of the targeted change.** Something references an attribute of the resource you fixed — a new identifier, a new endpoint, a new ARN. These are expected and are the whole point of running the full plan. 2. **Drift made by humans during the incident.** Somebody scaled an instance, opened a security-group rule, or flipped a flag in the console at 3am to stop the bleeding. Terraform is now offering to revert it. This is the dangerous bucket, because reverting a fix that is still load-bearing turns one incident into two. 3. **Configuration merged while you were firefighting.** Changes that arrived on the main branch and have simply never been applied. They are legitimate but they are not yours, and they should not ride along unnoticed in an incident cleanup. Only the first bucket is unambiguously safe to apply immediately. For the second, the decision is per resource: revert it if it was a stopgap, or write it into the configuration if it was the right answer. The third belongs to whoever merged it, applied in the normal flow. ## Sequencing the cleanup ```bash terraform plan -out=cleanup.tfplan # read it, all of it terraform apply cleanup.tfplan # apply exactly what was reviewed terraform plan # must come back empty ``` Saving the plan and applying that file matters here more than usual, because the estate is moving: a plan read at 03:20 and re-derived at 03:40 is not the same plan. The cleanup is finished when an untargeted plan reports no changes. ## What not to do The tempting mistake is another targeted apply — this time on the resources the diff surfaced — because it is smaller and feels safer at 3am. It is not safer; it defers the same reconciliation and adds a new stale slice. Occasionally it is still the right call *while the incident is live*, but then it is another entry in the same debt, not a resolution. Equally wrong is deciding Terraform is being erratic. Terraform is deterministic here: it is proposing exactly the work a deliberately partial run left behind. ## The process fix Treat a targeted apply as a break-glass action with an obligation attached. Record the exact command in the incident channel, note which addresses were targeted, and open a follow-up item to get the untargeted plan back to empty within the working day. Otherwise the debt lands on the next engineer who runs a routine apply and finds a diff nobody can explain — and that person, under time pressure, is far more likely to approve it blind than you are now.

  • How do you tell a diff caused by your targeted apply from drift somebody made in the console?
    Read it at the attribute level. Consequences of your change reference values the fixed resource produced — an identifier, an endpoint, an ARN — and Terraform is filling them in. Console drift shows up as attributes your configuration already sets being pulled back to the declared value. Your cloud audit log settles anything ambiguous by naming who changed what, and when.
  • The full plan wants to revert a firewall rule someone opened during the incident. What do you do?
    Do not apply it blind. Find out whether the rule is still holding production up: if it is a stopgap that is no longer needed, let the revert proceed; if it is the actual fix, put it in the configuration first and then apply, so the plan converges without taking the fix away. Reverting a live mitigation turns one incident into two.
  • Would you ever run a second targeted apply to clear this diff?
    Only while the incident is still live and speed genuinely outranks completeness. Outside that window it just adds another partial slice and defers the same reconciliation to someone with less context. The exit condition never changes: an untargeted plan that reports no changes.

saying these in an interview costs you the question

  • Auto-approves the first full plan after a targeted apply
  • Blames the unexpected diff on Terraform being non-deterministic
  • Assumes the targeted run refreshed the whole state anyway
  • Runs more targeted applies to make the diff disappear
  • Reverts console changes without checking whether they are still load-bearing

context