skip to content

What State Stores

State maps each address in my config to a real object in the cloud, which is how Terraform knows what to update and what to delete. If I can explain that mapping, most other state questions answer themselves.

part ofTerraformoverview, primer and where to startread it →
on this pageshow

questions

6

A Terraform root module's state file is gone and there is no backup, but everything it created is still running in the cloud. What happens on the next terraform plan and apply, and how do you recover?

level: middleimportance: must knowfreq 66%

answer

  1. every address misses, so everything is an add
  2. the old objects don't disappear, they detach
  3. destroy is driven by the record, not the config
  4. unique names fail instead of duplicating
  5. re-adopting is per-object, per-ID work

basics

~20 s

Terraform sees every declared resource as new and plans to create a second copy of everything, while the existing objects become orphans it can no longer read, change or destroy. Recovery means restoring a snapshot or re-adopting each object into state.

solid answer

~50 s

Plan shows a full create of every resource in the configuration, because Terraform matches config to reality only through state and there is now nothing to match. Apply then builds duplicates — a second VPC, a second database — and the originals keep running, unreferenced and unmanaged: Terraform cannot update them and, crucially, cannot destroy them, because a destroy is driven by the state entry that names the object ID. Some resources will fail instead of duplicating, where a name or CIDR is unique, so you typically get a half-created mess rather than a clean second copy. Recovery is either restoring an earlier state snapshot if your storage keeps versions, or re-adopting each object into a fresh state one by one. The second path is slow and manual, which is why state storage is treated as production data.

go deeper

for a junior

Know that state is not a disposable cache: if it is lost, Terraform no longer recognises what it built and will try to build it again. Say that the running infrastructure itself is unaffected.

for a middle

Walk through the plan and apply mechanics — every address misses, so everything is an add; unique names fail; the originals become unmanageable, including undeletable. Name both recovery paths.

for a senior

Show the operational judgment: read the plan as the alarm, prefer the versioned snapshot, and reconcile what changed since it was taken. Explain why state storage is treated as production data with versioning and restricted access.

for a principal

Frame it as a blast-radius decision: how much of the estate one lost or corrupted state file can strand, who is allowed to write it, and what the documented recovery drill is before the incident rather than during it.

## The mechanics Terraform decides what to do for each resource by looking it up in state by address. The loop is: read the configuration, read the state, for each configured address find the matching state entry, refresh that entry from the provider, and diff. When state is empty, every address misses. A miss is not an error — it is the definition of "needs to be created." So `terraform plan` reports the whole configuration as adds, with a plan summary of the form `Plan: 42 to add, 0 to change, 0 to destroy`. ## What apply actually does Apply then does exactly what the plan said. Two things happen at once: **Duplicates.** Every resource that *can* be created twice is created twice. You end up with two VPCs, two clusters, two load balancers, and a bill to match. The new objects go into the new state; the old ones exist nowhere in Terraform's world. **Partial failure.** Every resource that *cannot* be created twice fails at the provider — an S3 bucket whose name is globally unique, an IAM role with a fixed name, a subnet CIDR that collides inside the existing VPC. Terraform applies in dependency order and stops the affected branches on error, so the realistic outcome is a partially-created parallel estate plus a stack of provider errors, which is messier to clean up than either a clean duplicate or a clean failure. ## The orphans The original objects are the sharper problem. They are still running and still serving traffic, but Terraform has no entry pointing at them, so: - it cannot show them in a plan; - it cannot change them, because a change is expressed as an update to a recorded object; - it **cannot destroy them**. This is the part candidates miss. Destroy is not "find everything matching the config and delete it" — it is "for each object recorded in state, call the provider's delete with the recorded ID." With no record, there is no delete. Those resources now have to be cleaned up by hand in the console or the cloud CLI, or re-adopted first. ## Recovery, in order of preference **Restore a snapshot.** If the state was stored somewhere with version history, the previous snapshot is the fastest route back. Restore it, then run a plan and read it very carefully: the snapshot is from before whatever writes happened since, so anything created after it is missing from the restored ledger and will look new all over again, and anything destroyed after it will appear in state but be gone in reality — the refresh drops those entries and the plan proposes recreating them. **Re-adopt object by object.** Failing a snapshot, you rebuild the ledger by pointing Terraform at each existing object's real ID so it records the binding without creating anything. This is per-resource, per-instance work: dozens of IDs looked up in the console, a plan run after each batch, and every plan read to confirm it says no changes rather than proposing a replacement. For a large estate this is days, not hours. **Rebuild from scratch.** Sometimes honest and fastest for a small, stateless environment: create the parallel estate deliberately, cut over, then delete the orphans manually. Never an option where the estate holds data. ## Reading the aftermath The tell that this has happened, before you apply, is a plan that proposes creating things you know already exist. Treat any plan whose add count matches the size of your configuration as a stop signal — it almost always means Terraform is looking at the wrong state or no state at all, not that your infrastructure vanished. ## Why interviewers ask it The question tests whether you understand that state is *identity*, not a cache. A candidate who thinks state is an optimisation says "Terraform would just re-read everything from the API." A candidate who understands the ledger says "Terraform has no way to know those objects are its own, so it makes new ones and abandons the old." That difference also explains why state storage gets versioning, backups and restricted access — it is the one artifact whose loss cannot be reconstructed from your repository.

  • Does losing state break the running infrastructure itself?
    No. The cloud objects keep running exactly as they were — nothing about them depends on Terraform. What is lost is Terraform's ability to identify, change or delete them. That is precisely why it is dangerous: nothing alerts, nothing goes down, and the damage only appears the next time someone applies.
  • You restore a state snapshot from last week. What should you check before applying?
    Read the plan, do not trust it. Anything created since that snapshot is absent from the ledger and will show as a create against objects that already exist; anything destroyed since will show as a recreate. Reconcile those two lists by hand — re-adopt the missing ones and confirm the deleted ones — before any apply.
  • Why does a plan against an empty state not just error out?
    Because "no state entry for this address" is the normal, expected signal for a brand-new resource — it is how the very first apply of any configuration works. Terraform cannot distinguish a first run from a lost ledger, which is exactly why the safeguard has to be operational: versioned storage and restricted access, not a CLI check.

saying these in an interview costs you the question

  • Terraform would re-discover the resources from the cloud API
  • Just run destroy and start over
  • Losing state only costs you a slower next plan
  • The configuration is the source of truth, so state can be regenerated
  • Terraform errors out rather than planning duplicates

context

open as a page

Terraform writes a state file (terraform.tfstate) after every apply. What does that file actually record, and why can't Terraform work from your configuration plus live cloud API queries alone?

level: middleimportance: must knowfreq 82%

basics

~20 s

Terraform state binds each configuration address, such as aws_instance.web, to the real object ID the provider returned, and stores that object's last-known attributes plus its recorded dependencies. Without the binding, Terraform cannot tell which cloud objects are its own.

open as a page

Terraform state is plain JSON, so why is opening terraform.tfstate in an editor and fixing it by hand considered dangerous, and what is the supported alternative?

level: juniorimportance: should knowfreq 52%

basics

~20 s

Being readable does not make the format a public interface: it is versioned, internally consistent, and rewritten by Terraform on every run. Hand edits bypass the bookkeeping and corrupt silently, so Terraform provides dedicated state subcommands for every legitimate change.

open as a page

You delete a resource block from your Terraform configuration and run apply. The config no longer says anything about what that resource depended on, yet Terraform still destroys it in a safe order. How?

level: middleimportance: should knowfreq 38%

basics

~20 s

Each instance in Terraform's state carries a dependencies list of the addresses it referenced at the last apply. For a resource whose block is gone, Terraform works entirely from that recorded entry and reverses the recorded edges to order the destroy.

open as a page

Terraform's state stores the full attribute values of each resource, not just its ID. What does it need those cached values for, and how can they mislead a plan when they go stale?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Cached attributes are the prior side of every diff, the source for cross-resource references and outputs at plan time, and the only home for values the API never returns. When they are stale, the plan is computed against a fiction and can show no changes while reality differs.

open as a page

A Terraform state file carries top-level serial and lineage fields. What is each one for, and what does Terraform do when the lineage of the state it is about to write does not match the one it read?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Serial is a counter Terraform increments on every state write, so it can tell which snapshot is newer. Lineage is a UUID minted when the state was first created, identifying the line of descent; a mismatch means two unrelated states, and Terraform refuses rather than overwriting.

open as a page