skip to content

Why does Terraform lock the state before a run, and what actually goes wrong if two applies run against the same state at the same time?

level: juniorimportance: must knowfreq 72%

answer

  1. one writer at a time
  2. read-modify-write on one document
  3. last writer wins, silently
  4. orphaned resources with no record
  5. plan takes the lock too

basics

~20 s

Terraform locks state so only one write-capable run touches it at a time. Two concurrent applies read the same starting snapshot and then overwrite each other's write, so real resources end up with no record in state and later runs recreate or destroy them.

solid answer

~50 s

Terraform's state file is a single document that every run does a read-modify-write on: read the current state, refresh and change infrastructure, write the whole file back. If two applies overlap, both read the same starting snapshot, both make real API calls, and whichever finishes last writes a state that never saw the other's changes. The resources the loser created are still real but untracked, so the next plan proposes creating them again — or a `destroy` never removes them and they bill forever. The lock closes that window by making the backend refuse a second write-capable run while one is in progress; the second run fails with `Error acquiring the state lock` rather than silently racing. Locking is automatic where the backend supports it, and it applies to `plan` as well as `apply`, because a refresh can persist an updated state too.

go deeper

for a junior

Know that Terraform takes a lock before a run and that the second run fails with an error instead of racing. Be able to say plainly that two applies at once can leave real resources with no entry in state.

for a middle

Explain the read-modify-write cycle and where the race window sits: both runs read the same serial, both act, the last write wins. Mention that plan locks as well, because refresh can persist state.

for a senior

Show the operational consequence — untracked resources that keep billing, plans proposing to recreate what exists, name collisions — and describe how you would detect and repair it rather than only preventing it.

for a principal

Own the argument that a single lock per state means concurrency is a property of how you split state and sequence pipelines. Be able to weigh serialised safety against teams blocking each other, and where you would put that boundary.

## The thing being protected Terraform records what it manages in a state file: a mapping from the addresses in your configuration (`aws_instance.web`) to the real object IDs at the provider, plus cached attributes and the dependency order it used. That file is the *only* place that mapping exists. Nothing in AWS or GCP tells Terraform "this instance is yours and it is called `aws_instance.web`" — the state file says so. Because it is one document, every run performs a read-modify-write against it: 1. read the current state, 2. refresh it against the real world, decide a plan, and execute it, 3. write the whole updated state back to the backend. Step 2 can take minutes. The window between the read and the write is where two runs can collide. ## What a collision produces Suppose two engineers, or two CI jobs, apply the same root module at once, and locking is off or bypassed. Both read serial N. Run A creates a load balancer; run B creates a security group. Run A finishes and writes serial N+1 containing the load balancer. Run B finishes a minute later and writes its own serial N+1 (or N+2) containing the security group — computed from a state that never contained the load balancer. The load balancer is now real, costing money, and absent from state. Terraform will not destroy it on `terraform destroy`, and the next `plan` will propose creating a *second* one — which then collides on a unique name, or silently duplicates. This is not corruption in the sense of invalid JSON; it is worse, because the file parses fine and looks authoritative. A nastier variant happens when both runs touch the *same* resource. Two applies can issue conflicting API calls to the same object — one updating an autoscaling group while the other replaces it — and the provider may return errors that look like transient cloud failures rather than a self-inflicted race. ## How the lock closes the window Before any operation that could write state, Terraform asks the backend to take an exclusive lock and records metadata with it: a lock ID, the path, the operation, who is running it, and when it started. The lock is released when the run finishes, successfully or not. A second run that finds the lock held fails immediately: ``` Error: Error acquiring the state lock Lock Info: ID: 4f2c9c17-9a6d-4c94-a09b-0e4a0dd9e0b5 Operation: OperationTypeApply Who: runner@ci-7f4b Created: 2026-03-04 10:12:41 UTC ``` That metadata is the point: it turns "something is wrong" into "this specific run, started by this principal, at this time". The lock is *advisory in scope*, not a distributed transaction. It stops a second Terraform run; it does not stop a human clicking in the console, and it does not roll anything back if a run dies. It only serialises Terraform against itself for one state. ## Which operations lock Everything that could persist state: `apply`, `destroy`, `import`, the `state` subcommands — and `plan`, which surprises people. `plan` refreshes resources and can write the refreshed state back, so it takes the lock too. That is why a long plan in CI can block a colleague's apply, and why "it was only a plan" is not a reason to bypass locking. ## Where you get no protection Locking is a backend feature, not a Terraform-core feature, so what you get depends on where state lives. Backends that implement it lock automatically with no configuration. A state file sitting on each engineer's laptop has no shared lock to take at all — the local backend takes an operating-system file lock, which serialises two runs on *one machine* and does nothing about two machines. And `-lock=false` disables the mechanism outright on any backend. ## What it does not solve Locking prevents concurrent *writes*. It does not prevent drift, it does not prevent someone editing infrastructure outside Terraform, and it does not merge anything: there is no three-way merge of two state files. Recovery from an overlapping apply is manual — usually importing the orphaned resources back or restoring a previous state version from the backend's history. Preventing the race is enormously cheaper than repairing it, which is why the lock is on by default.

  • Does terraform plan take the state lock as well, or only apply?
    Plan takes it too. A plan refreshes resources against the provider and can persist that refreshed state, so it is a write-capable operation. Practically, a slow plan in CI can block a teammate's apply on the same state — which is a reason to split state or shorten plans, never a reason to run plans with locking disabled.
  • If two applies did overlap, how would you notice afterwards?
    The next plan proposes creating resources that visibly already exist, or an apply fails on a duplicate name. You may also see cloud objects with your tagging conventions that `terraform state list` does not know about. Confirm by diffing the backend's state versions — the serial numbers show two writes derived from the same parent.
  • Once resources are orphaned by a race, how do you get them back under management?
    Either import them back to their configuration addresses, or destroy them out of band if they were surplus. Restoring an older state version only helps if nothing has been applied since; otherwise it re-orphans whatever the newer state recorded. Decide per resource rather than rolling the whole file back reflexively.

saying these in an interview costs you the question

  • Thinks state locking prevents drift or console edits
  • Believes only apply locks state, never plan
  • Assumes a local state file protects a whole team
  • Says overlapping applies are fine because refresh rebuilds state
  • Thinks Terraform merges two concurrent state writes

context