A Terraform root module's state file is gone and there is no backup, but everything it created is still running in the cloud. What happens on the next terraform plan and apply, and how do you recover?
answer
- every address misses, so everything is an add
- the old objects don't disappear, they detach
- destroy is driven by the record, not the config
- unique names fail instead of duplicating
- re-adopting is per-object, per-ID work
basics
~20 sTerraform sees every declared resource as new and plans to create a second copy of everything, while the existing objects become orphans it can no longer read, change or destroy. Recovery means restoring a snapshot or re-adopting each object into state.
solid answer
~50 sPlan shows a full create of every resource in the configuration, because Terraform matches config to reality only through state and there is now nothing to match. Apply then builds duplicates — a second VPC, a second database — and the originals keep running, unreferenced and unmanaged: Terraform cannot update them and, crucially, cannot destroy them, because a destroy is driven by the state entry that names the object ID. Some resources will fail instead of duplicating, where a name or CIDR is unique, so you typically get a half-created mess rather than a clean second copy. Recovery is either restoring an earlier state snapshot if your storage keeps versions, or re-adopting each object into a fresh state one by one. The second path is slow and manual, which is why state storage is treated as production data.
go deeper
Know that state is not a disposable cache: if it is lost, Terraform no longer recognises what it built and will try to build it again. Say that the running infrastructure itself is unaffected.
Walk through the plan and apply mechanics — every address misses, so everything is an add; unique names fail; the originals become unmanageable, including undeletable. Name both recovery paths.
Show the operational judgment: read the plan as the alarm, prefer the versioned snapshot, and reconcile what changed since it was taken. Explain why state storage is treated as production data with versioning and restricted access.
Frame it as a blast-radius decision: how much of the estate one lost or corrupted state file can strand, who is allowed to write it, and what the documented recovery drill is before the incident rather than during it.
## The mechanics Terraform decides what to do for each resource by looking it up in state by address. The loop is: read the configuration, read the state, for each configured address find the matching state entry, refresh that entry from the provider, and diff. When state is empty, every address misses. A miss is not an error — it is the definition of "needs to be created." So `terraform plan` reports the whole configuration as adds, with a plan summary of the form `Plan: 42 to add, 0 to change, 0 to destroy`. ## What apply actually does Apply then does exactly what the plan said. Two things happen at once: **Duplicates.** Every resource that *can* be created twice is created twice. You end up with two VPCs, two clusters, two load balancers, and a bill to match. The new objects go into the new state; the old ones exist nowhere in Terraform's world. **Partial failure.** Every resource that *cannot* be created twice fails at the provider — an S3 bucket whose name is globally unique, an IAM role with a fixed name, a subnet CIDR that collides inside the existing VPC. Terraform applies in dependency order and stops the affected branches on error, so the realistic outcome is a partially-created parallel estate plus a stack of provider errors, which is messier to clean up than either a clean duplicate or a clean failure. ## The orphans The original objects are the sharper problem. They are still running and still serving traffic, but Terraform has no entry pointing at them, so: - it cannot show them in a plan; - it cannot change them, because a change is expressed as an update to a recorded object; - it **cannot destroy them**. This is the part candidates miss. Destroy is not "find everything matching the config and delete it" — it is "for each object recorded in state, call the provider's delete with the recorded ID." With no record, there is no delete. Those resources now have to be cleaned up by hand in the console or the cloud CLI, or re-adopted first. ## Recovery, in order of preference **Restore a snapshot.** If the state was stored somewhere with version history, the previous snapshot is the fastest route back. Restore it, then run a plan and read it very carefully: the snapshot is from before whatever writes happened since, so anything created after it is missing from the restored ledger and will look new all over again, and anything destroyed after it will appear in state but be gone in reality — the refresh drops those entries and the plan proposes recreating them. **Re-adopt object by object.** Failing a snapshot, you rebuild the ledger by pointing Terraform at each existing object's real ID so it records the binding without creating anything. This is per-resource, per-instance work: dozens of IDs looked up in the console, a plan run after each batch, and every plan read to confirm it says no changes rather than proposing a replacement. For a large estate this is days, not hours. **Rebuild from scratch.** Sometimes honest and fastest for a small, stateless environment: create the parallel estate deliberately, cut over, then delete the orphans manually. Never an option where the estate holds data. ## Reading the aftermath The tell that this has happened, before you apply, is a plan that proposes creating things you know already exist. Treat any plan whose add count matches the size of your configuration as a stop signal — it almost always means Terraform is looking at the wrong state or no state at all, not that your infrastructure vanished. ## Why interviewers ask it The question tests whether you understand that state is *identity*, not a cache. A candidate who thinks state is an optimisation says "Terraform would just re-read everything from the API." A candidate who understands the ledger says "Terraform has no way to know those objects are its own, so it makes new ones and abandons the old." That difference also explains why state storage gets versioning, backups and restricted access — it is the one artifact whose loss cannot be reconstructed from your repository.
- Does losing state break the running infrastructure itself?No. The cloud objects keep running exactly as they were — nothing about them depends on Terraform. What is lost is Terraform's ability to identify, change or delete them. That is precisely why it is dangerous: nothing alerts, nothing goes down, and the damage only appears the next time someone applies.
- You restore a state snapshot from last week. What should you check before applying?Read the plan, do not trust it. Anything created since that snapshot is absent from the ledger and will show as a create against objects that already exist; anything destroyed since will show as a recreate. Reconcile those two lists by hand — re-adopt the missing ones and confirm the deleted ones — before any apply.
- Why does a plan against an empty state not just error out?Because "no state entry for this address" is the normal, expected signal for a brand-new resource — it is how the very first apply of any configuration works. Terraform cannot distinguish a first run from a lost ledger, which is exactly why the safeguard has to be operational: versioned storage and restricted access, not a CLI check.
saying these in an interview costs you the question
- Terraform would re-discover the resources from the cloud API
- Just run destroy and start over
- Losing state only costs you a slower next plan
- The configuration is the source of truth, so state can be regenerated
- Terraform errors out rather than planning duplicates