skip to content

State Management Theory

Some tools keep a recorded map of everything they created; others re-derive it from the cloud on every run. Understanding why that record exists, and what breaks when two people apply at once, is the deepest idea in this branch.

on this pageshow

questions

4

Infrastructure-as-code tools such as Terraform and Pulumi keep a recorded state of what they created. Why does a tool need that record at all, and what could it not do without one?

level: juniorimportance: must knowfreq 72%

answer

  1. the tool must remember what it built
  2. code names versus opaque cloud ids
  3. how it knows what to delete
  4. dependency order after the code is gone
  5. lose it and everything becomes orphans

basics

~20 s

A state record maps each declared resource to the real object the tool created. Without it the tool cannot tell an update from a create, cannot know which real objects to delete when code is removed, and cannot order a destroy.

solid answer

~50 s

Your configuration says "one database called primary"; the cloud knows only an opaque identifier like `db-prod-8f2a10c4`. The state record is the bridge: for every resource address in the code it stores the real object's identifier, plus a snapshot of its attributes and the dependencies between resources. That mapping is what makes a second run an *update* instead of a duplicate create. It is also the only place the tool learns that an object it once created is no longer in the code, so removing a block can mean "destroy this" rather than "I have never seen this". And because dependencies are recorded alongside the objects, the tool can still tear an environment down in the correct order after the code describing it has been deleted. Lose the record and the tool believes nothing exists: it plans to create everything again while the real infrastructure keeps running unmanaged.

code

json · 14 lines
json
{
  "resources": [
    {
      "address": "database.primary",
      "mode": "managed",
      "id": "db-prod-8f2a10c4",
      "dependencies": ["network.private_subnet"],
      "attributes": {
        "endpoint": "db-prod.internal:5432",
        "engine_version": "16.3"
      }
    }
  ]
}

go deeper

for a junior

Be able to say plainly that the record maps each resource in your code to the real object's id, and that it is how the tool knows what to update or delete on the next run. Never call it a disposable cache.

for a middle

Explain the three-way comparison — code, record, live provider — and why deletion is the one decision that is impossible without a record. Be ready to describe what a run does when the record is empty but the infrastructure exists.

for a senior

Show that you treat the record as production data: backed up, versioned, access-controlled, never hand-edited. Talk through recovering an environment whose record was lost, and what you would check before adopting existing objects.

for a principal

Own the policy question: who may read and write each record, how many records the estate has, what the recovery objective for one is, and what the organisation gives up by adopting a tool that keeps no client-side record at all.

## The problem the record solves Infrastructure-as-code is written in terms the author chose: a resource called `primary`, a network called `private`. The provider knows nothing about those names. It hands back opaque identifiers — an instance id, an ARN, a resource path — and those identifiers are the only handle for a later update or delete. Nothing in the cloud API links "the resource my code calls primary" to "the object with id db-prod-8f2a10c4". Some component has to remember that link, and in a tool like Terraform or Pulumi that component is the recorded state. ## Four jobs the record does **1. Identity mapping.** For each address in the configuration the record stores the real object's identifier. This is what makes runs idempotent at the *tool* level: on the second run the tool looks up the address, finds an identifier, reads that object, and decides whether an update is needed. Without the mapping the only honest interpretation of the configuration is "create this", so every run would produce another copy. ```json { "address": "database.primary", "id": "db-prod-8f2a10c4", "dependencies": ["network.private_subnet"] } ``` **2. Knowing what to delete.** Deletion is the job that cannot be done from the configuration alone, because the configuration is exactly where the evidence has just been removed. The tool compares three things: what the code declares now, what the record says it created before, and what the provider currently reports. An object present in the record but absent from the code is the definition of "you removed this, so I should destroy it". A tool with no record can only ever add or converge; it has no way to notice a subtraction. **3. Dependency order that outlives the code.** Order matters on the way down as much as on the way up: a subnet cannot be deleted while an instance still sits in it. On a create the tool derives that order from references in the configuration. On a destroy of resources whose code has already been deleted there are no references left to read, so the recorded dependency edges are what preserve a safe teardown order. **4. A cached snapshot of attributes.** The record also stores attribute values the provider returned. That lets the tool compute a diff, and lets other parts of the configuration consume values (an address, a generated name) without re-querying. This copy is the part that goes stale — the world can change underneath it — which is why tools re-read live objects before diffing, and why the gap between the record and reality is a permanent topic of its own. ## What losing the record looks like Suppose CI keeps the record in a bucket and someone empties the bucket. The next run sees an empty record, concludes that none of the declared resources exist, and proposes to create all of them. Meanwhile the real environment is still up and serving traffic. In the best case the run fails partway on a name or address that is already taken; in the worse case you end up with a duplicate environment and a set of orphans — running, billed, and now managed by nobody. Recovery means re-establishing the mapping object by object, which is slow and error-prone. That is why the record is production data: it is backed up, versioned, and access-controlled, not treated as a scratch file. The inverse mistake matters too. The record is not the desired state — the code is. The record only says what the tool believes it built. If you hand-edit it, you have not changed any infrastructure; you have changed the tool's beliefs, which is how you get a run that destroys something real or adopts something it never made. ## Sharing and blast radius One record belongs to one configuration, and everyone who applies that configuration must read and write the same copy — otherwise two people hold two divergent beliefs about the same infrastructure. That single shared record is also a single blast radius: everything it tracks is in scope of every run against it. Sharing a mutable record between humans and pipelines is precisely what creates the concurrency problem, and it is why write-bearing operations take a lock. ## Not every tool keeps one A record is a design choice, not a law. A configuration-management tool can re-derive the current state from each host on every run; a hosted service can keep the equivalent record server-side; a controller can read live objects continuously. Each of those alternatives gives up one of the four jobs above — usually deletion tracking — in exchange for having nothing to lose, lock, or leak.

  • If the tool re-reads the provider before every diff, why can't it just skip the record and discover its resources by querying the account?
    Because a query returns objects, not ownership. Nothing in the account says which objects this configuration created, which resource address each belongs to, or which ones a human made by hand. Tag conventions are an approximation that breaks the moment someone forgets a tag or another team copies it, and adopting whatever it finds would let one configuration silently take over another team's infrastructure.
  • Someone hand-edits the record to remove an entry so a run stops trying to change a resource. What have they actually done?
    They have made the tool forget the object, not stop managing reality. The resource keeps running and billing, unmanaged, and the next run will propose creating a replacement — often colliding on a name that is already taken. The legitimate operations are to change the code, or to use the tool's own supported way of adopting or forgetting a resource, both of which keep the record consistent.
  • Why does the record store attribute values at all, rather than only the identifier?
    So the tool can compute a diff and resolve references cheaply: the previous values are one half of the comparison, and other resources consume outputs like an endpoint or generated name without another API call. The cost is that this snapshot can go stale relative to the live object, which is why a run refreshes it before diffing.

It is a receipt book. The catalogue says what you wanted; the receipts say which specific serial-numbered items you actually took home, which is the only way to return one later.

saying these in an interview costs you the question

  • The state file is just a cache; deleting it is harmless
  • State stores your cloud credentials
  • The tool can always rediscover its resources by scanning the account
  • State is the desired state, so editing it changes infrastructure
  • State only matters once you work in a team

context

open as a page

An IaC tool creates a database with a generated password. Why does that plaintext password end up in the tool's recorded state, and what follows for how you store and access that record?

level: middleimportance: should knowfreq 48%

basics

~20 s

The record stores every attribute the provider returned, and a generated password is just another attribute. Any recorded state therefore contains secrets in the clear, so its storage must be encrypted, tightly access-controlled and audited exactly like a secret store.

open as a page

Two engineers start an apply at the same time against the same shared IaC state record. What can go wrong, and why is locking the standard answer?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Each run reads the record, changes real infrastructure, then writes the record back. Interleaved, the later write overwrites the earlier one, so resources the first run created vanish from the record and become unmanaged orphans. A mutual-exclusion lock serialises the write-bearing operation.

open as a page

Ansible keeps no state file, while Terraform and Pulumi record one. How does a tool with no recorded state decide what to do on each run, and what does each approach trade away?

level: middleimportance: nice to knowfreq 38%

basics

~20 s

A stateless tool re-derives reality from the target on every run: each step reads the current condition and acts only if it differs. It has nothing to lose, lock or leak, but it cannot notice that something you deleted from the code should be removed.

open as a page