A failed pipeline run leaves the Terraform state object in your S3 backend truncated and unusable. How does bucket versioning let you recover, and what do you check after restoring?
answer
- state is one overwritten object
- previous copies still exist
- pick the newest intact version
- serial and lineage must match
- plan afterwards, creates mean orphans
basics
~20 sS3 bucket versioning keeps every previous copy of the state object, so recovery means listing object versions, identifying the last good one, and restoring it as the current version. Afterwards run terraform plan: unexpected creates mean resources built after that snapshot need importing.
solid answer
~50 sThe state object is a single mutable key, so without versioning a bad write is final. With versioning enabled, every overwrite keeps the prior copy, and recovery is: `aws s3api list-object-versions` on that key, download the candidates, pick the last one that parses and has the highest sane `serial`, then restore it — either by copying that version back over the key, or with `terraform state push`. Then the real work: `terraform plan`. A clean plan means you are whole. Creates in the plan mean the failed apply actually created resources that the restored snapshot predates, so those are now unmanaged and need importing rather than creating. Deletes mean the restored state is *newer* than reality. I would also check whether the crashed run left a lock behind, and treat the whole thing as an argument for versioning plus a restrictive bucket policy being non-negotiable.
code
bash · 15 lines# List every retained version of the state object
aws s3api list-object-versions \
--bucket acme-tfstate \
--prefix prod/network/terraform.tfstate
# Pull a candidate down and inspect it before trusting it
aws s3api get-object \
--bucket acme-tfstate \
--key prod/network/terraform.tfstate \
--version-id 3HL4kqtJlcpXroDTDmjVBH40Nrjfkd \
restored.tfstate
# Restore, then audit
terraform state push restored.tfstate
terraform plango deeper
Know that state lives as a single object that gets overwritten, and that enabling versioning on the state bucket is what makes recovering a previous copy possible at all.
Explain the mechanics: list object versions, choose the newest intact one, check serial and lineage before restoring, and understand that terraform state push has safety checks a direct object copy skips.
Show the incident judgment — freeze the pipeline first, minimise the reconciliation window by taking the newest good version, and read the post-restore plan line by line because creates mean orphaned real resources, not work to do.
Own the prevention posture: versioning and retention policy on every state bucket by default, delete permissions restricted to break-glass, one key per root module to bound blast radius, and apply jobs that cannot be killed halfway.
## Why versioning is the whole safety net Terraform state in an object-storage backend is one key that is overwritten in place on every apply. A crashed process, a truncated upload, a mistaken `aws s3 cp`, or a state operation gone wrong replaces the authoritative record of your production estate with garbage — and Terraform cannot rebuild it, because it has no way to enumerate which cloud objects were "its". Versioning on the bucket converts that from a catastrophe into an inconvenience, which is why it belongs in the bucket's definition from day one rather than after the first incident. The equivalents exist elsewhere: object versioning on GCS buckets, and blob versioning or soft delete on the Azure storage account used by the `azurerm` backend. Pair versioning with a noncurrent-version lifecycle rule that keeps old versions long enough to survive a slow-to-notice corruption — days, not hours — and with encryption at rest and a policy that restricts who can write or delete the key at all. ## The recovery procedure **1. Stop the bleeding.** Freeze the pipeline. Any further apply against a corrupt state can create duplicates or destroy real resources, and every write buries the good version deeper. **2. Enumerate versions.** ```bash aws s3api list-object-versions \ --bucket acme-tfstate \ --prefix prod/network/terraform.tfstate ``` You get version IDs with timestamps. The most recent is the broken one. **3. Fetch and inspect candidates.** Download each candidate with `aws s3api get-object --version-id ...` and check two things inside the JSON: it parses at all, and the `serial` and `lineage` fields look right. Terraform increments `serial` on each write and keeps `lineage` constant for the life of a state; a candidate with a different lineage belongs to a different state entirely and must not be restored here. **4. Restore.** Either copy the chosen version back over the key, making it the new current version, or use `terraform state push` with the downloaded file. Note that `state push` performs safety checks — it refuses when the pushed serial is lower than the remote one or the lineage differs, and `-force` bypasses those checks. Copying the object directly in S3 bypasses the checks entirely, which is faster and blinder; either way the verification step below is what actually protects you. **5. Verify with a plan.** `terraform plan` is the audit: - *No changes* — the restored snapshot matches reality; you are done. - *Creates* — the failed run genuinely created resources after the snapshot was taken. They exist in the cloud but not in state, so applying would build duplicates. Adopt them into state instead of creating them. - *Destroys* — the restored state is ahead of reality, or the failed run deleted things; check each one before letting an apply act on it. - *In-place updates* — usually harmless drift in cached attributes that the next refresh reconciles. **6. Clear any leftover lock.** A process that died mid-apply may not have released the backend lock, so the first recovery run can fail to acquire it. That is a separate mechanism with its own safety rules — never break a lock you have not confirmed is stale. ## Judgment, not just commands The hard part is step 5. Restoring an old state is not "undo": infrastructure kept moving while state did not, so what you have restored is a snapshot of the world as of a moment that has passed. Reconciling the difference between the snapshot and reality is the actual recovery, and it must be done by reading the plan resource by resource rather than by trusting the exit code. The safest posture is to restore the *newest* version that is intact, minimising the window you have to reconcile by hand. It is also worth asking why the truncated write happened. Common causes: a CI runner killed mid-apply by a job timeout or a spot reclaim; an engineer running a state operation against the wrong key; a network failure during upload. The first is best fixed by making apply jobs uninterruptible and generously timed, the second by tightening who can write to state at all. ## Prevention that costs nothing Versioning on. Encryption on. A bucket policy that denies deletes to everything except a break-glass role. One state key per root module so a bad operation cannot take out several estates. And treat the state bucket itself as infrastructure that is created once and never casually re-provisioned. ## How to say it in an interview "Versioning means the previous state object is still there. List versions, take the newest intact one, restore it, and then run a plan — creates in that plan are resources the failed run made that the snapshot predates, and those get imported, not created."
- Is the DynamoDB table used by the S3 backend a second copy of the state you could restore from?No. The table holds the lock item and a digest of the state used to detect a mismatched write — not the state document itself. Restoring from it is impossible. The only backup is the bucket's own object versions, plus whatever local or exported copies happen to exist, which is exactly why versioning is mandatory rather than nice to have.
- After restoring, terraform plan shows creates for three resources that clearly already exist. What now?They were created by the failed run after the restored snapshot, so they are real but unmanaged. Adopt them into state rather than letting the plan build duplicates — Terraform's import mechanism maps an existing object to a configuration address. Only after a subsequent plan comes back clean should the pipeline be unfrozen.
- How do you keep a state object from being casually deleted in the first place?Restrict it at the bucket: deny DeleteObject and DeleteObjectVersion to everything except a break-glass role, keep versioning plus a noncurrent-version retention window measured in days, and give day-to-day pipeline roles write access only to their own key prefix. Read access matters too, since state contains secrets in plaintext.
saying these in an interview costs you the question
- Believes Terraform can rebuild lost state from the cloud account
- Thinks the DynamoDB lock table stores a state backup
- Applies immediately after restoring without reading the plan
- Restores an old version and calls the estate recovered
- Enables bucket versioning only after the first incident