skip to content

A colleague says an infrastructure-as-code tool "just compares my code to what exists in the cloud". Which three distinct states does a reconciliation run actually involve, and which comparisons does the tool make between them?

level: middleimportance: must knowfreq 72%

answer

  1. three states, not two
  2. code, record, reality
  3. refresh updates beliefs, not infrastructure
  4. the record is what enables deletion
  5. unrecorded objects are invisible

basics

~20 s

Three states, not two: desired state (your code), recorded state (the tool's snapshot of the objects it manages) and actual state (what the provider reports now). A run refreshes the record from reality, then diffs desired against it.

solid answer

~50 s

There are three states. **Desired** is the configuration you wrote. **Recorded** is the tool's own snapshot of the objects it believes it manages — identifiers plus the attribute values it last observed. **Actual** is whatever the provider's API returns right now. A run first refreshes: for everything in the record, it reads the real object and updates the record's view of it. Then it diffs desired against that refreshed record to decide create, update in place, replace or destroy. The record is what makes deletion possible — an object in the record that no longer appears in the configuration gets destroyed, while an object that exists in the cloud but was never recorded is simply invisible to the tool. Stateless tools collapse this to two states by re-deriving reality from the target on every run, which is exactly why they can converge a host but cannot tell that something should be removed.

go deeper

for a junior

Be able to name the three states out loud — the code you wrote, the tool's record of what it made, and what the provider actually has — and say plainly that the record is how the tool recognises resources it created earlier.

for a middle

Explain the run as refresh-then-diff, and say which pair each phase compares. The point that earns the mark is why deletion needs the record when create and update do not.

for a senior

Show what you do when the three states disagree in production: a lost or stale record, an object the tool forgot, a refresh that is slow or rate-limited. Talk about consequences and recovery, not definitions.

for a principal

Own the estate-level consequence: how many records you keep, what each one's blast radius is, and how you make sure the mapping between code and reality survives a lost file, a migrated backend or a team reorganisation.

## Why "code versus cloud" is the wrong mental model The two-state model — code on one side, the real world on the other — is the most common misunderstanding in infrastructure-as-code interviews, and almost every confusing behaviour a newcomer hits comes from it. A stateful tool works with three states, and knowing which pair is being compared at each moment explains create-versus-update, explains deletion, explains why a hand-made resource is ignored, and explains why losing the record is a catastrophe rather than an inconvenience. ## The three states **Desired state** is the declaration: the files a human wrote and reviewed, plus the input values supplied for this run. It says what should exist, not how to get there. **Recorded state** is the tool's own bookkeeping. For each object it manages, it holds the provider-side identifier and the attribute values it saw last time. This record is the only thing that ties the abstract name in the configuration to the concrete object in the provider — without it, a tool re-running the same configuration has no way to know whether the object it should create already exists. **Actual state** is ground truth: what the provider's API reports at this instant. It changes without the tool's involvement — someone edits the console, an autoscaler adjusts a capacity, a service applies a default the tool never asked for. ## The loop A reconciliation run is two comparisons, in order. First, **refresh**: for every object in the record, read the live object and update the record's view. This changes the tool's beliefs, not the infrastructure. It is also where the tool discovers an object that has been deleted out from under it and marks it as gone. Second, **diff**: compare the desired configuration against the refreshed record and derive the actions that close the gap. Roughly: - in desired, not in the record → create - in both, attributes differ → update in place, or replace if the attribute cannot be changed on a live object - in the record, not in desired → destroy - in both and identical → no action Then the actions execute against the provider, and the record is written back so the next run starts from an accurate baseline. Note that nothing in this loop scans the account for objects the tool never created — the fourth quadrant, *exists in reality but not in the record*, is deliberately outside the tool's field of view. That is what makes it possible for a hundred teams to run a hundred configurations against one account without each one trying to delete the others' work. ## Why deletion is the interesting case Create and update can be derived from desired plus actual alone: look at what should exist, look at what does, make up the difference. Deletion cannot. If you remove a block from your configuration, the desired state no longer mentions the object at all — there is nothing left to compare against reality. Only the record remembers "I made this, and it is no longer wanted". This single fact is the cleanest one-sentence answer to "why does the tool keep a state file at all", and it is why deleting the record does not delete infrastructure: it orphans it. The objects keep running, the tool forgets them, and the next run tries to create a second set. ## Stateless tools Not every tool keeps a record. A configuration-management run against a host reads the host's current condition — installed packages, file contents, running services — and compares that directly to the declaration. Two states, re-derived every time, nothing persisted between runs. The trade is exact: no state file to store, lock, corrupt or leak, but also no memory of what it created, so removing a declaration from the code means the previously configured thing simply stops being managed rather than being removed. Practitioners work around this by declaring removal explicitly — stating that a package should be *absent* rather than deleting the line that said it should be present. A control-plane service that owns the record on the server side is a third arrangement: the same three states, with the recorded one living in the provider's own database instead of a file you store. And a cluster controller sits somewhere between — the desired object and the observed status both live in the same API, so the record and the desired state are stored together while actual state is read from the world. ## Interview traps The candidate who says the record is "just a cache" gets pushed on deletion and usually fails. The candidate who says refresh "fixes" drift is confusing updating the tool's beliefs with changing infrastructure — refresh only makes the record honest; the following diff decides whether reality gets corrected.

  • If the recorded state is lost, is the infrastructure lost with it?
    No — the real objects keep running untouched. What is lost is the mapping from configuration to object, so the tool no longer knows those resources exist. The next run tries to create a duplicate set, and the originals become orphans nobody manages. Recovery means restoring the record from backup or versioned storage, or re-adopting each object into a fresh record one at a time.
  • Why can the refresh phase itself be a problem in a large estate?
    Refresh makes at least one provider API call per recorded object before any diff can be computed, so a run over thousands of resources spends most of its wall-clock time reading, and can hit provider rate limits. It also means a provider outage blocks a run that would otherwise have been a no-op. Splitting the estate into smaller records, or skipping refresh when you know nothing changed, are the usual levers.
  • A stateless tool has no record. How does it express "this should no longer exist"?
    By declaring absence rather than deleting the declaration. You keep a statement in the code saying the package, file or service should be absent, and the run converges the host to that. Removing the line instead just stops managing the thing, leaving it in place. The cost is that the codebase accumulates tombstones you eventually have to garbage-collect by hand.

saying these in an interview costs you the question

  • The tool compares code directly to the cloud; there is no record
  • The state file is only a cache, so deleting it is harmless
  • Refresh repairs infrastructure to match the code
  • A resource created by hand will be destroyed as unmanaged
  • Removing a declaration is enough for a stateless tool to delete it

context