skip to content

Plan and Preview Lifecycle

A dry-run diff before any change is what makes IaC safe to point at production. I should be able to walk the refresh to diff to approve to apply lifecycle and name what can still go wrong between plan and apply.

on this pageshow

questions

5

An infrastructure-as-code change preview lists actions such as create, update in place, replace, and destroy. What does each mean, and why is replace the one to look at hardest?

level: juniorimportance: must knowfreq 72%

answer

  1. not every change can be done in place
  2. some attributes are fixed at creation
  3. one action means delete then recreate
  4. new identity, and local data is gone
  5. start the review at the destroy count

basics

~20 s

Create adds a new object, update in place modifies an existing one, destroy removes it, and replace destroys the existing object and creates a new one because an attribute changed that the provider cannot alter after creation. Replace is destructive: data and identity do not survive it.

solid answer

~40 s

A diff assigns every managed object one action. Create means it does not exist yet. Update in place means the provider can modify the changed attributes on the live object. Destroy means it is being removed, usually because you deleted it from the configuration. Replace is the dangerous one: you changed an attribute the provider's API only accepts at creation time, so the tool schedules a destroy followed by a create. The new object gets a new identity — new identifier, new address, new endpoint — and anything stored on it that was not separately persisted is gone. That is why a review should start at the counts: any non-zero replace or destroy against a stateful resource is a stop-and-look, and the preview normally annotates which attribute forced the replacement.

go deeper

for a junior

Know the four actions by name and be able to point at the destroy and replace lines in a preview and say why they are the risky ones. Do not approve a diff you have not read to the summary counts.

for a middle

Explain that replacement is forced by provider attributes that are immutable after creation, and that the default ordering is destroy-then-create. Name the opt-in reversal and what it costs.

for a senior

Show judgment about avoiding a replacement in production: snapshot first, check what holds references to the old identity, consider standing the replacement up alongside and cutting over rather than letting the tool do it in one run.

for a principal

Own the guardrails at the estate level — deletion protection on data-bearing resources, code-level rules that make a destroy fail rather than proceed, and a policy that any diff containing a replacement of stateful infrastructure routes to the owning team.

## The action vocabulary A diff is a list of managed objects, each tagged with exactly one intent: - **create** — the object is declared but does not exist. The tool will call the provider's create API. - **update in place** — the object exists and differs from the declaration in attributes the provider can change on a live object. One or more modify calls. - **destroy** — the object exists and is no longer declared (or is being removed deliberately). A delete call. - **replace** — the object exists but the change touches an attribute that cannot be modified after creation, so the tool schedules a delete and a create as one logical change. - **no change** — the declaration matches reality. Most of the estate on any given run. ## Why replacement happens This is not a tool decision, it is a provider constraint. Cloud APIs expose many attributes that are fixed at creation: which subnet or availability zone an instance lives in, the engine family of a managed database, the partition key of a table, the name of many resources. The provider integration declares which attributes force a new object, and when the diff touches one, the only way to reach your declared state is to build a new object. So the mental model is: *update in place is possible only where the vendor's API has an update call for that field*. Everything else is a rebuild. ## What a replacement actually costs Three things, all easy to underestimate: 1. **Data.** Anything living on the object that is not in a separately-managed store disappears — local disks, in-memory caches, a database whose storage is part of the instance. The tool will not migrate it for you. 2. **Identity.** The new object has a new identifier, and usually a new address or endpoint. Everything holding a reference — DNS entries, connection strings, allow-lists in other systems, other teams' hardcoded values — is pointing at something that no longer exists until it is updated. 3. **Downtime.** By default, most tools destroy first and create second, because the old object holds the name, address, or unique constraint the new one needs. That means the gap is the sum of teardown plus provisioning plus warm-up. Tools generally offer an opt-in mode that builds the replacement before removing the original, but it only works where two of the object can legally coexist, and it costs a period of double resources. ## Reading a preview safely Start at the summary counts rather than the body. `0 to destroy` on a production run is a good day. A non-zero destroy or replace count deserves a specific question: *which* object, and *which attribute* forced it — the output normally names the offending field on the replace line. Watch for one confusing case. If you rename the logical address of a resource in code without changing the real object, a naive diff reads as a destroy of the old address plus a create of the new one, even though nothing about the infrastructure needs to change. That is a bookkeeping problem, not an infrastructure one, and every mature tool has a way to record the rename instead of executing it. Approving that diff as written would delete a healthy production resource for no reason. ## Guardrails worth having - Enable the provider's own deletion protection on data-bearing resources, so a mistaken destroy fails at the API rather than succeeding. - Declare, at the code level, that specific resources may never be destroyed, so the preview errors instead of proposing it. - Require a human approval whenever the diff contains any destroy or replace, and auto-approve otherwise. - Take a snapshot before applying a replacement of anything that stores data, and confirm you can restore it. ## Interview framing Define the four actions crisply, then make the point that the tool is not choosing to be destructive — it is reporting the provider's immutability constraint. Finish with the operational consequences: new identity, lost local data, and downtime whose length is the sum of destroy and create.

  • Your preview shows a replacement and you cannot accept the downtime. What are your options?
    First, check which attribute forced it — sometimes the change is avoidable, or achievable by adding a new resource alongside instead of mutating the old one. Second, use the opt-in build-before-destroy ordering, if the resource permits two to coexist under the same names and constraints. Third, do it manually as a migration: stand up the replacement under a different name, cut traffic across, then remove the original in a later change.
  • A colleague renamed a resource's identifier in the code and the preview now proposes destroying and recreating it. Is that correct behaviour?
    It is expected behaviour but the wrong outcome. The tool tracks objects by their address in code, so a renamed address looks like one object disappearing and another appearing, even though the real infrastructure is unchanged. The fix is to record the rename at the bookkeeping level so the tool re-points its record at the existing object, and then re-run the preview until it shows no changes.
  • Why is update in place available for some attributes and not others on the same resource?
    Because it mirrors the provider's API surface exactly. If the vendor exposes a modify call for an attribute, the tool can change it live; if the attribute is only accepted in the create call, there is nothing to invoke. The integration encodes this per attribute, which is why one field on a resource is a harmless in-place edit and its neighbour forces a full rebuild.

saying these in an interview costs you the question

  • Says replace just restarts the resource without losing anything
  • Assumes the tool always creates the replacement before destroying the original
  • Believes any attribute can be changed without recreating the resource
  • Approves a preview without looking at the destroy count
  • Thinks data on a replaced resource is migrated across automatically

context

open as a page

Walk through the phases an infrastructure-as-code tool goes through when it previews a change, from reading the current world to applying it. What happens in each?

level: middleimportance: must knowfreq 78%

basics

~20 s

A preview runs in four phases: refresh, where the tool re-reads live infrastructure; diff, where it compares your declared configuration against that reality and lists proposed actions; approval, where a human or policy gate reviews them; and apply, which executes the approved actions in dependency order.

open as a page

An infrastructure-as-code preview was generated and reviewed twenty minutes ago, and the apply runs now. What can make the applied result differ from what was reviewed, and how do teams narrow that window?

level: seniorimportance: should knowfreq 52%

basics

~20 s

A preview describes the world at the moment it was computed, so anything that changes afterwards causes skew: another apply, a console edit, an autoscaler, or values that only resolve during apply. Teams narrow it by serializing applies behind a lock, applying the reviewed artifact, and keeping approval-to-apply short.

open as a page

You own the human approval gate for infrastructure-as-code changes across many teams. How do you decide which changes require an approval, and what makes such an approval meaningful rather than a rubber stamp?

level: principalimportance: should knowfreq 34%

basics

~20 s

Gate on the content of the diff rather than on every change: destructive, stateful, security-relevant and production changes need a human, additive and reversible ones do not. An approval is meaningful only when the approver can read the diff, is accountable for the resource, and can realistically decline.

open as a page

Why do mature infrastructure-as-code pipelines execute a saved preview artifact from the review step instead of recomputing the diff at apply time, and what does that approach cost?

level: seniorimportance: nice to knowfreq 42%

basics

~20 s

Saving the reviewed diff and executing exactly that artifact guarantees the change a human approved is the change that runs. Recomputing at apply time silently applies a diff nobody saw. The costs are artifact plumbing between jobs, a sensitive file to protect, staleness failures, and a hard version pin.

open as a page