skip to content

State & Backends

State is Terraform's record of what it created, and it holds the hardest questions: remote backends, locking, drift, import, and the fact that secrets land in it in plaintext. Expect at least one state question in any Terraform interview.

part ofTerraformoverview, primer and where to startread it →
on this pageshow

questions

30

Why does Terraform lock the state before a run, and what actually goes wrong if two applies run against the same state at the same time?

level: juniorimportance: must knowfreq 72%

answer

  1. one writer at a time
  2. read-modify-write on one document
  3. last writer wins, silently
  4. orphaned resources with no record
  5. plan takes the lock too

basics

~20 s

Terraform locks state so only one write-capable run touches it at a time. Two concurrent applies read the same starting snapshot and then overwrite each other's write, so real resources end up with no record in state and later runs recreate or destroy them.

solid answer

~50 s

Terraform's state file is a single document that every run does a read-modify-write on: read the current state, refresh and change infrastructure, write the whole file back. If two applies overlap, both read the same starting snapshot, both make real API calls, and whichever finishes last writes a state that never saw the other's changes. The resources the loser created are still real but untracked, so the next plan proposes creating them again — or a `destroy` never removes them and they bill forever. The lock closes that window by making the backend refuse a second write-capable run while one is in progress; the second run fails with `Error acquiring the state lock` rather than silently racing. Locking is automatic where the backend supports it, and it applies to `plan` as well as `apply`, because a refresh can persist an updated state too.

go deeper

for a junior

Know that Terraform takes a lock before a run and that the second run fails with an error instead of racing. Be able to say plainly that two applies at once can leave real resources with no entry in state.

for a middle

Explain the read-modify-write cycle and where the race window sits: both runs read the same serial, both act, the last write wins. Mention that plan locks as well, because refresh can persist state.

for a senior

Show the operational consequence — untracked resources that keep billing, plans proposing to recreate what exists, name collisions — and describe how you would detect and repair it rather than only preventing it.

for a principal

Own the argument that a single lock per state means concurrency is a property of how you split state and sequence pipelines. Be able to weigh serialised safety against teams blocking each other, and where you would put that boundary.

## The thing being protected Terraform records what it manages in a state file: a mapping from the addresses in your configuration (`aws_instance.web`) to the real object IDs at the provider, plus cached attributes and the dependency order it used. That file is the *only* place that mapping exists. Nothing in AWS or GCP tells Terraform "this instance is yours and it is called `aws_instance.web`" — the state file says so. Because it is one document, every run performs a read-modify-write against it: 1. read the current state, 2. refresh it against the real world, decide a plan, and execute it, 3. write the whole updated state back to the backend. Step 2 can take minutes. The window between the read and the write is where two runs can collide. ## What a collision produces Suppose two engineers, or two CI jobs, apply the same root module at once, and locking is off or bypassed. Both read serial N. Run A creates a load balancer; run B creates a security group. Run A finishes and writes serial N+1 containing the load balancer. Run B finishes a minute later and writes its own serial N+1 (or N+2) containing the security group — computed from a state that never contained the load balancer. The load balancer is now real, costing money, and absent from state. Terraform will not destroy it on `terraform destroy`, and the next `plan` will propose creating a *second* one — which then collides on a unique name, or silently duplicates. This is not corruption in the sense of invalid JSON; it is worse, because the file parses fine and looks authoritative. A nastier variant happens when both runs touch the *same* resource. Two applies can issue conflicting API calls to the same object — one updating an autoscaling group while the other replaces it — and the provider may return errors that look like transient cloud failures rather than a self-inflicted race. ## How the lock closes the window Before any operation that could write state, Terraform asks the backend to take an exclusive lock and records metadata with it: a lock ID, the path, the operation, who is running it, and when it started. The lock is released when the run finishes, successfully or not. A second run that finds the lock held fails immediately: ``` Error: Error acquiring the state lock Lock Info: ID: 4f2c9c17-9a6d-4c94-a09b-0e4a0dd9e0b5 Operation: OperationTypeApply Who: runner@ci-7f4b Created: 2026-03-04 10:12:41 UTC ``` That metadata is the point: it turns "something is wrong" into "this specific run, started by this principal, at this time". The lock is *advisory in scope*, not a distributed transaction. It stops a second Terraform run; it does not stop a human clicking in the console, and it does not roll anything back if a run dies. It only serialises Terraform against itself for one state. ## Which operations lock Everything that could persist state: `apply`, `destroy`, `import`, the `state` subcommands — and `plan`, which surprises people. `plan` refreshes resources and can write the refreshed state back, so it takes the lock too. That is why a long plan in CI can block a colleague's apply, and why "it was only a plan" is not a reason to bypass locking. ## Where you get no protection Locking is a backend feature, not a Terraform-core feature, so what you get depends on where state lives. Backends that implement it lock automatically with no configuration. A state file sitting on each engineer's laptop has no shared lock to take at all — the local backend takes an operating-system file lock, which serialises two runs on *one machine* and does nothing about two machines. And `-lock=false` disables the mechanism outright on any backend. ## What it does not solve Locking prevents concurrent *writes*. It does not prevent drift, it does not prevent someone editing infrastructure outside Terraform, and it does not merge anything: there is no three-way merge of two state files. Recovery from an overlapping apply is manual — usually importing the orphaned resources back or restoring a previous state version from the backend's history. Preventing the race is enormously cheaper than repairing it, which is why the lock is on by default.

  • Does terraform plan take the state lock as well, or only apply?
    Plan takes it too. A plan refreshes resources against the provider and can persist that refreshed state, so it is a write-capable operation. Practically, a slow plan in CI can block a teammate's apply on the same state — which is a reason to split state or shorten plans, never a reason to run plans with locking disabled.
  • If two applies did overlap, how would you notice afterwards?
    The next plan proposes creating resources that visibly already exist, or an apply fails on a duplicate name. You may also see cloud objects with your tagging conventions that `terraform state list` does not know about. Confirm by diffing the backend's state versions — the serial numbers show two writes derived from the same parent.
  • Once resources are orphaned by a race, how do you get them back under management?
    Either import them back to their configuration addresses, or destroy them out of band if they were surplus. Restoring an older state version only helps if nothing has been applied since; otherwise it re-orphans whatever the newer state recorded. Decide per resource rather than rolling the whole file back reflexively.

saying these in an interview costs you the question

  • Thinks state locking prevents drift or console edits
  • Believes only apply locks state, never plan
  • Assumes a local state file protects a whole team
  • Says overlapping applies are fine because refresh rebuilds state
  • Thinks Terraform merges two concurrent state writes

context

open as a page

Terraform writes terraform.tfstate to the local working directory by default. What breaks once a second engineer starts running applies on the same project, and what does moving to a remote backend actually fix?

level: juniorimportance: must knowfreq 82%

basics

~20 s

A remote backend keeps Terraform state in shared, durable storage such as an S3 bucket instead of one laptop's disk. Every engineer and CI job then reads and writes the same state, so runs cannot silently diverge, duplicate resources, or destroy each other's work.

open as a page

Which files produced by a Terraform run must never be committed to a Git repository, and why is a saved plan file created with terraform plan -out=tfplan as sensitive as the state file itself?

level: juniorimportance: must knowfreq 65%

basics

~20 s

Never commit terraform.tfstate, terraform.tfstate.backup, the .terraform directory, saved plan files, crash logs, or tfvars files holding secrets. A saved plan embeds the prior state plus planned attribute values, so it carries the same plain-text credentials state does.

open as a page

In Terraform, what does importing a resource actually do to state and configuration, and how does the `import` block differ from the `terraform import` command?

level: middleimportance: must knowfreq 78%

basics

~20 s

Importing binds an already-existing object to a resource address in Terraform state. It creates no infrastructure and writes no configuration for you. The terraform import command needs the resource block written first; a Terraform 1.5+ import block runs inside plan and apply.

open as a page

A Terraform root module's state file is gone and there is no backup, but everything it created is still running in the cloud. What happens on the next terraform plan and apply, and how do you recover?

level: middleimportance: must knowfreq 66%

basics

~20 s

Terraform sees every declared resource as new and plans to create a second copy of everything, while the existing objects become orphans it can no longer read, change or destroy. Recovery means restoring a snapshot or re-adopting each object into state.

open as a page

Terraform writes a state file (terraform.tfstate) after every apply. What does that file actually record, and why can't Terraform work from your configuration plus live cloud API queries alone?

level: middleimportance: must knowfreq 82%

basics

~20 s

Terraform state binds each configuration address, such as aws_instance.web, to the real object ID the provider returned, and stores that object's last-known attributes plus its recorded dependencies. Without the binding, Terraform cannot tell which cloud objects are its own.

open as a page

You inherit a Terraform project managing about 60 live resources whose state is a local terraform.tfstate file. How do you move it onto an S3 backend without destroying or recreating anything?

level: middleimportance: must knowfreq 62%

basics

~20 s

Create the state bucket first with versioning enabled, add a backend "s3" block to the terraform block, then run terraform init -migrate-state to copy the existing state up. Confirm success by running terraform plan and seeing no changes.

open as a page

A Terraform configuration generates a database password with the random_password resource and passes it to an RDS instance. After apply, where does that password exist in plaintext, and why does marking the corresponding output sensitive = true not change that?

level: middleimportance: must knowfreq 75%

basics

~20 s

Terraform state caches resource attributes exactly as configured or returned, so the generated password sits in plaintext in the state file, its local backup and any saved plan. sensitive = true only redacts CLI output; it never encrypts or omits storage.

open as a page

How do you decide where to split a Terraform estate into separate state files, and what does splitting actually buy you?

level: middleimportance: must knowfreq 68%

basics

~20 s

Split along blast radius, change cadence and ownership: long-lived networking in one state, fast-moving application resources in another. The payoff is shorter plans, less lock contention, narrower credentials per state, and a mistake that cannot reach the whole estate.

open as a page

What does `terraform state rm` do to the real cloud object, and when is dropping a resource from Terraform state the right move rather than a mistake?

level: seniorimportance: must knowfreq 55%

basics

~20 s

terraform state rm deletes only the state entry: the cloud object keeps running, now managed by nobody. It is right when handing an object to another state or deliberately releasing it from Terraform; since Terraform 1.7 a removed block does the same thing declaratively.

open as a page

A CI job was killed mid-apply and every Terraform run since fails with 'Error acquiring the state lock'. How do you clear it, and what must you verify before running terraform force-unlock?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Read the Lock Info block, then prove the run that took the lock is really dead — check the CI job and the named principal — before running terraform force-unlock with that exact lock ID. Force-unlocking a live apply lets a second run race it and lose state writes.

open as a page

In Terraform, what do `terraform state list` and `terraform state show` tell you, and how do they differ from the other `terraform state` subcommands?

level: juniorimportance: should knowfreq 50%

basics

~10 s

terraform state list prints the resource addresses Terraform is tracking; terraform state show <address> prints that one resource's recorded attributes. Both only read state. The other subcommands, such as mv and rm, rewrite it.

open as a page

Terraform state is plain JSON, so why is opening terraform.tfstate in an editor and fixing it by hand considered dangerous, and what is the supported alternative?

level: juniorimportance: should knowfreq 52%

basics

~20 s

Being readable does not make the format a public interface: it is versioned, internally consistent, and rewritten by Terraform on every run. Hand edits bypass the bookkeeping and corrupt silently, so Terraform provides dedicated state subcommands for every legitimate change.

open as a page

In Terraform, how does one root configuration read a value produced by a different configuration's state, and what has to exist on the producing side?

level: juniorimportance: should knowfreq 60%

basics

~10 s

The producing configuration must declare a root-level output. The consumer adds a terraform_remote_state data source pointing at the producer's backend and reads data.terraform_remote_state.NAME.outputs.KEY. Only declared root outputs are exposed, never arbitrary resource attributes.

open as a page

You rename a Terraform resource and move it inside a module, and the plan now shows a destroy and a create. How do you make the refactor without touching the real infrastructure?

level: middleimportance: should knowfreq 62%

basics

~20 s

Add a moved block with the old address in from and the new address in to. Terraform then rewrites the state mapping during plan and apply instead of destroying and recreating the object. The alternative, terraform state mv, does the same thing imperatively and unreviewably.

open as a page

Where does a Terraform state lock physically live for the S3, GCS and local backends, and which of those actually protects a team?

level: middleimportance: should knowfreq 52%

basics

~20 s

The lock is a backend feature, not a Terraform one. The S3 backend writes a .tflock object next to the state (use_lockfile) or, historically, a DynamoDB item; GCS writes a lock object; the local backend takes an OS file lock that only covers one machine.

open as a page

You delete a resource block from your Terraform configuration and run apply. The config no longer says anything about what that resource depended on, yet Terraform still destroys it in a safe order. How?

level: middleimportance: should knowfreq 38%

basics

~20 s

Each instance in Terraform's state carries a dependencies list of the addresses it referenced at the last apply. For a resource whose block is gone, Terraform works entirely from that recorded entry and reverses the recorded edges to order the destroy.

open as a page

Why does Terraform reject an input variable inside a backend block, and how do you point the same configuration at a different state bucket per environment given that restriction?

level: middleimportance: should knowfreq 50%

basics

~20 s

Backend settings are read at init, before variables, locals or providers are evaluated, so they must be literal values. The workaround is partial configuration: omit the varying arguments and supply them at init with -backend-config files or key=value flags.

open as a page

You inherit a production environment that was built by hand in a cloud console, and you must bring it under Terraform without recreating anything. How do you approach it?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Adopt in small slices ordered by blast radius, write or generate configuration for each slice, import it, and treat a plan reporting no changes as the acceptance test before moving on. Never apply a post-import plan that still proposes changes.

open as a page

Pipeline runs keep failing with a held Terraform state lock and a colleague proposes adding -lock=false. Why is that dangerous, and what should you do instead?

level: seniorimportance: should knowfreq 46%

basics

~20 s

-lock=false disables locking entirely, converting a loud failure into a silent race that can strand real resources outside state. Use -lock-timeout so a queued run waits for the lock instead, and serialise the pipeline so only one run per state executes at a time.

open as a page

Terraform's state stores the full attribute values of each resource, not just its ID. What does it need those cached values for, and how can they mislead a plan when they go stale?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Cached attributes are the prior side of every diff, the source for cross-resource references and outputs at plan time, and the only home for values the API never returns. When they are stale, the plan is computed against a fiction and can show no changes while reality differs.

open as a page

A failed pipeline run leaves the Terraform state object in your S3 backend truncated and unusable. How does bucket versioning let you recover, and what do you check after restoring?

level: seniorimportance: should knowfreq 45%

basics

~20 s

S3 bucket versioning keeps every previous copy of the state object, so recovery means listing object versions, identifying the last good one, and restoring it as the current version. Afterwards run terraform plan: unexpected creates mean resources built after that snapshot need importing.

open as a page

A security review is satisfied because the S3 bucket holding Terraform state has SSE-KMS enabled and the S3 backend block sets encrypt = true. Which threats does that actually close, and what remains wide open?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Server-side encryption protects state only against someone reading the storage underneath the API — stolen media, a mis-scoped bucket listing, an unauthorised backup copy. Any identity allowed to fetch the object gets it decrypted, so authorization, not encryption, is the real control.

open as a page

A repository is found to contain a committed terraform.tfstate holding a live production database password, and the same state also lives in a versioned S3 bucket. What does remediation actually require, and why are deleting the file and running terraform state rm not part of it?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Rotate the credential first — that is the only action that revokes access. Copies persist in Git history and clones, prior S3 object versions, local backups and CI artifacts, so removing files cannot restore confidentiality, and terraform state rm only makes Terraform forget the resource.

open as a page

Compare the three ways one Terraform configuration can consume another's results: the terraform_remote_state data source, a normal provider data lookup, and a value published to a store such as SSM Parameter Store.

level: seniorimportance: should knowfreq 50%

basics

~20 s

terraform_remote_state reads the producer's entire state file — simple, but it needs broad read access. A provider data lookup queries the live API by tag or name — looser, but it relies on naming discipline. Publishing to SSM Parameter Store makes the contract explicit and narrowly scoped.

open as a page

Terraform 1.10 introduced ephemeral values and 1.11 added write-only resource arguments. What problem do they solve that marking a value sensitive never could, and what constraint does a write-only argument place on your configuration?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

They keep a secret out of state and out of plan files entirely, rather than merely redacting it in output. The tradeoff: a write-only argument is never stored, so Terraform cannot diff it — you bump a paired version argument to make it re-send the value.

open as a page

A Terraform state file carries top-level serial and lineage fields. What is each one for, and what does Terraform do when the lineage of the state it is about to write does not match the one it read?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Serial is a counter Terraform increments on every state write, so it can tell which snapshot is newer. Lineage is a UUID minted when the state was first created, identifying the line of descent; a mismatch means two unrelated states, and Terraform refuses rather than overwriting.

open as a page

As the author of a shared Terraform module, how do you restructure its internal resource addresses without forcing every consumer into a destroy and create?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Ship moved blocks inside the module itself. They travel with the module version, so each consumer's next plan re-keys their own state entries automatically. Keep them for a documented deprecation window and drop them only at a major version boundary.

open as a page

When would you choose HCP Terraform's cloud block over a self-managed S3 backend for a team's Terraform state, and what does each choice cost you?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

A self-managed object-storage backend stores state cheaply and keeps everything inside your account, but you build locking, versioning, access control and the pipeline yourself. HCP Terraform provides state, locking, run history, RBAC and optional remote execution as a product, at per-resource cost and outside your network.

open as a page

Your team has split a Terraform estate into about thirty state files, one per component. What does that granularity start costing you, and how do you judge whether it has gone too far?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

Every seam adds a handoff: no single plan shows a cross-cutting change, applies must run in a chosen order, and one edit ripples through several pipelines. Split for blast radius and ownership; stop when routine changes need three states coordinated.

open as a page