skip to content

How does an infrastructure-as-code tool actually detect drift — what does it compare against what, and how does that differ between tools that keep a recorded state and tools that do not?

level: middleimportance: should knowfreq 58%

answer

  1. declared, recorded, observed
  2. which pair differs tells you which kind
  3. record is what catches a deletion
  4. stateless tools re-derive every run
  5. detection is only as good as the read

basics

~20 s

Detection re-reads each managed resource from the provider API and compares three things: the configuration you declared, the record of what the tool last created, and the observed live values. A difference between the record and the observation is drift.

solid answer

~50 s

Every tool detects drift the same way underneath: it asks the provider what the resource looks like right now, then compares that observation with what it declared and with what it last recorded. Three-way comparison matters, because it is what separates the two kinds of difference. Declared versus recorded is a change *you* made and have not applied. Recorded versus observed is drift — someone or something else moved reality. Tools that keep a record (a state file, or state held server-side by the platform) can make that distinction and can also notice deletions, because the record still lists an object the provider no longer returns. Stateless tools such as Ansible keep no record and re-derive the machine's condition on every run, so they see only declared versus observed: they converge what you named and are blind to anything you removed from the code. Continuously reconciling controllers do the same comparison, just on a loop.

code

json · 6 lines
json
{
  "resource": "web_server",
  "declared": { "instance_type": "t3.medium", "monitoring": true },
  "recorded": { "instance_type": "t3.medium", "monitoring": true },
  "observed": { "instance_type": "t3.large", "monitoring": true }
}

go deeper

for a junior

Know that the tool asks the cloud provider what each managed resource looks like right now, and compares that with what your files say. Say that detection is a read, not a change.

for a middle

Be able to name the three states — declared, recorded, observed — and say which pair defines drift versus a pending change. Explain what a recorded state buys: identity, deletion tracking and attribution.

for a senior

Demonstrate the limits: write-only attributes are invisible, provider normalisation creates phantom drift, refreshing a big estate costs API calls and hits throttling, and unmanaged resources are outside the comparison entirely.

for a principal

Own the tradeoff between stateful and stateless designs across the estate, and decide where detection lives: a periodic scan, a platform-side check, or a controller reconciling continuously — each with different cost, latency and failure modes.

## The three states Drift detection is a comparison between three descriptions of the same resource: - **Declared** — what the configuration files say should exist. - **Recorded** — what the tool believes it last created or last saw, kept in a state file, in a backend, or on the platform's own servers. - **Observed** — what the provider API returns when asked right now. Detection begins with the observation. The tool walks the resources it manages, calls the provider's read APIs for each one, and refreshes its picture of reality. Only then can it compare. ```json { "declared": { "instance_type": "t3.medium", "monitoring": true }, "recorded": { "instance_type": "t3.medium", "monitoring": true }, "observed": { "instance_type": "t3.large", "monitoring": true } } ``` Here `recorded` and `observed` disagree on one attribute: that is drift, and nobody edited the code. If instead `declared` and `recorded` disagreed while `recorded` matched `observed`, that would be an ordinary pending change waiting to be applied. Distinguishing those two is the entire practical value of keeping a record. ## Why the record earns its keep A recorded state buys three capabilities that pure declaration cannot provide: 1. **Identity.** The record maps a name in your code to a concrete provider identifier, so the tool knows *which* object to read. 2. **Deletion tracking.** If the record lists an object and the provider says it does not exist, the tool knows it was deleted out of band. Symmetrically, if you delete a block from the code, the record still names the object, so the tool knows to destroy it. Without a record there is nothing to compare a removal against. 3. **Change attribution.** As above — telling "you changed this" apart from "the world changed this". ## Stateless tools Configuration-management tools in the Ansible style keep no record between runs. Each run gathers the current condition of the target directly and converges only the things the playbook names. This is a perfectly coherent design: the machine itself is the source of observed truth, so there is nothing to lose or corrupt, no locking problem, and no adoption step. The costs follow directly from the missing record: the tool cannot report drift as a standalone read-only question separate from a run, cannot tell you about resources you never mentioned, and cannot clean up something you removed from the playbook — you must write an explicit "absent" instruction instead. Platform-side state is a third shape: the service keeps the record for you on its servers, so drift detection is an API call to the platform rather than a local file read, and the record cannot be lost by an engineer's laptop. Continuously reconciling controllers are a fourth: they run the same comparison on a loop and act on it within seconds instead of when a human happens to run a command. ## What limits detection Detection is only as good as the read path, and every limitation here shows up as a real-world complaint: - **Write-only and unreadable attributes.** Passwords, secret values and some generated fields are never returned by the provider. The tool cannot see whether they drifted, so those attributes are structurally invisible. - **Attributes the tool does not model.** If the provider gained a field the plugin does not know about, changes to it are not drift as far as your tooling is concerned. - **Normalisation and semantic equality.** A policy document reordered by the provider, a value case-folded, a duration expressed differently — all produce differences that are textually real and operationally meaningless. This *phantom drift* is the biggest source of distrust in drift reports. - **Cost and rate limits.** Refreshing thousands of resources is thousands of API calls. On a large estate, detection is slow and can trip provider throttling, which is why scans are often scheduled rather than run on every command. - **Eventual consistency.** A resource read moments after a change may report an intermediate value and look like drift when it is simply not settled yet. - **Scope.** Detection covers what the tool manages. A resource created by hand and never adopted is invisible: there is no record of it, so nothing compares it to anything. Finding *unmanaged* resources is a different job, usually done by an inventory or cloud-security tool scanning the account. ## The point to make in an interview Say the three states out loud, name which pair defines drift, and then show you know that detection is a *read* — its accuracy is bounded by what the provider will tell you and how faithfully the tool compares it. That framing explains phantom drift, invisible secret fields, and slow scans in one move, without depending on any particular tool's command surface.

  • Why can a tool without recorded state not clean up a resource you deleted from the configuration?
    Because deletion is only visible by comparison with a record. With no record, removing the declaration simply means the tool no longer mentions that object, and nothing is left to say it once existed. Stateless tools therefore require an explicit "ensure absent" instruction, which you must remember to write and later remove.
  • What is phantom drift and why does it matter more than it sounds?
    It is a reported difference with no operational meaning — a provider reordering a policy document, normalising case, or filling a default. It matters because teams learn to skim drift reports that are always noisy, and the one genuine change hides among them. The fix is to correct the comparison or scope the attribute out deliberately, not to ignore the report.
  • Why are drift scans usually scheduled rather than run before every command?
    Because detection is a read of every managed resource, so it costs one or more provider API calls each. On a large estate that is slow and can hit rate limits, and it makes routine commands feel unusable. Teams run a cheap comparison for day-to-day work and a full refresh on a schedule.

saying these in an interview costs you the question

  • Saying the tool compares code directly to the state file only
  • Believing drift detection scans the whole cloud account
  • Thinking a stateless tool can detect out-of-band deletions
  • Assuming every reported difference is a real configuration change
  • Claiming detection sees secret or write-only attributes

context