skip to content

An infrastructure-as-code preview was generated and reviewed twenty minutes ago, and the apply runs now. What can make the applied result differ from what was reviewed, and how do teams narrow that window?

level: seniorimportance: should knowfreq 52%

answer

  1. the world moved between check and use
  2. time-of-check to time-of-use
  3. another writer got there first
  4. some values only resolve at apply
  5. shrink the window, freeze the diff

basics

~20 s

A preview describes the world at the moment it was computed, so anything that changes afterwards causes skew: another apply, a console edit, an autoscaler, or values that only resolve during apply. Teams narrow it by serializing applies behind a lock, applying the reviewed artifact, and keeping approval-to-apply short.

solid answer

~50 s

This is a time-of-check-to-time-of-use race. The diff is a statement about infrastructure as it was at refresh time, and the apply acts on it later. In between, another pipeline or engineer can apply a change, someone can edit the console, an autoscaler or provider-side maintenance can move an attribute, and values that were unknown at preview time only resolve during the apply itself. If the pipeline re-plans at apply time rather than executing the reviewed artifact, the code or input variables can also have moved. Teams narrow the window rather than closing it: serialize applies per state behind a lock, apply the saved plan artifact so the tool refuses when the recorded baseline has moved, auto-apply immediately once approval is recorded, and keep each state small so fewer actors touch it. What remains, you detect and repair afterwards.

go deeper

for a junior

Understand that a preview describes the moment it was run, so a diff from yesterday is not evidence about today. Re-run it before you apply.

for a middle

Explain the concrete causes: a concurrent apply, a manual console edit, an autoscaler, and values that only resolve during apply. Know that a lock serializes writers but does not cover the approval gap.

for a senior

Demonstrate having operated this: one writer per state, apply the artifact that was reviewed, short approval-to-apply windows, and a failure that stops the run rather than proceeding on a stale baseline. Say plainly that some residual risk remains.

for a principal

Own the tradeoff between review rigour and window length — long human review makes each diff better understood and more stale. Decide where to spend it, how small states should be to reduce collisions, and what detection you fund for what still gets through.

## The shape of the problem A preview is a claim about a moment: *given the world as I read it at 14:02, here is what I will do*. The apply happens at 14:22. Between those two timestamps the world is not frozen. This is the classic time-of-check-to-time-of-use race, and it is the reason the preview is strong evidence rather than a guarantee. ## What can move underneath you **Another actor applied a change.** A second pipeline run, a colleague running the tool locally, or a change from a different repository that touches overlapping resources. This is the most common and most damaging cause, because it can invalidate not just an attribute but the existence of an object. **Someone edited by hand.** A console change during an incident, or a script run against the account. The diff you reviewed compared against pre-edit values. **An autonomous system changed something.** An autoscaler adjusting a desired count, a certificate rotating, a managed service applying a maintenance-window update, another controller writing to the same object. These have no human to coordinate with and they will keep doing it. **Values were unknown at preview time.** Identifiers and computed attributes assigned by the provider at creation cannot be shown in the diff, and anything derived from them is unknown too. Those parts of the change were never actually reviewed — they resolve during apply. This is skew you cannot see coming, only bound. **The inputs differ between the two runs.** If the pipeline re-plans at apply time: the branch has moved on, a variable file changed, a lookup that resolves to "the newest image" now returns a different one, or the apply job runs different tool and provider versions than the preview job. Same code path, different answer. **The execution context differs.** The apply job may hold different credentials, a different role, or hit a quota that the read-only preview never exercised. ## How tools narrow it They do not eliminate it; they make the divergence loud rather than silent. - **A lock on the shared record** serializes applies, so two runs cannot interleave writes against the same set of resources. This turns a corrupted concurrent apply into a wait or a clear error. - **Applying a saved artifact** carries the baseline the diff was computed from. If the shared record has moved on since, the tool refuses to apply a stale artifact rather than executing a diff computed against a world that no longer exists. A refusal is the correct outcome here. - **Reading before writing.** Many providers support conditional writes keyed to a version or entity tag, so a modify call fails if the object changed since it was read. - **Failing loudly on apply.** If an object in the diff has been deleted underneath, the API call errors. The estate ends up partially applied, but you find out. ## How teams narrow it 1. **Shrink the window.** Approval should trigger the apply automatically, not queue it for a nightly batch. Minutes, not hours. 2. **One writer per state.** Pipeline concurrency of one per environment, and no local applies against production. Most skew is self-inflicted concurrency. 3. **Apply the artifact you reviewed.** If the apply job recomputes the diff, the review covered a document that no longer exists — and it will do so silently. 4. **Re-plan and compare, when you cannot ship an artifact.** Compute the diff again at apply time and hard-fail if it differs from the approved one. A weaker version of the same guarantee. 5. **Freeze the manual path.** No console write access outside break-glass, with the break-glass path alerting so you know a preview may be stale. 6. **Keep blast radius small.** A state that ten teams touch has ten sources of skew; five states with two each have far fewer collisions per run. 7. **Expire stale previews.** Refuse to apply an approval older than some threshold and require a re-run. ## What you cannot fix An apply is a sequence of independent API calls across systems with no shared transaction. There is no version of this where the reviewed diff is guaranteed to be the executed reality. The mature posture is: narrow the window, make divergence fail closed rather than proceed quietly, and rely on periodic re-checking of the estate to catch what slipped through. ## Interview framing Name it as time-of-check-to-time-of-use, list the categories of change (other actors, autonomous systems, unknown values, differing inputs), then give the mitigations in order of leverage: one writer per state, apply the reviewed artifact, short approval-to-apply window. Ending on "and you cannot eliminate it, so you also detect after the fact" is what separates a senior answer.

  • Does locking the shared record eliminate plan/apply skew?
    No. A lock serializes applies so two runs cannot write concurrently, which removes the worst corruption case, but the lock is normally held during the apply, not across the human approval. Anything that changes while a reviewer is reading — a console edit, an autoscaler, another team's apply that took the lock first and released it — still lands between the diff and your execution.
  • What is the argument against simply re-planning at apply time to get a fresh diff?
    It gives you the freshest possible view, but it breaks the link between the review and the execution: the diff that gets applied is one nobody has seen. You have not removed the risk, you have moved it from 'the reviewed diff may be stale' to 'the applied diff was never reviewed'. If you must re-plan, compare the new diff against the approved one and fail closed on any difference.
  • How would you detect that skew actually bit you on a past change?
    Compare the applied outcome against the approved diff — pipelines that keep both can diff them after the fact. Beyond that, the shared record's history shows which runs wrote and when, so an unexpected write between your preview and your apply is visible. Periodic re-checking of the estate against the declared configuration catches the residue that neither approach saw.

It is the same race as a booking site telling you a seat is free and you paying for it a minute later: the check was true when you made it, not necessarily when you acted on it.

saying these in an interview costs you the question

  • Treats the preview as a contract the apply is guaranteed to honour
  • Claims state locking removes all plan/apply skew
  • Asserts every value in a diff is known at preview time
  • Blames divergence on tool bugs rather than concurrent change
  • Says re-planning at apply time eliminates the risk entirely

context