skip to content

Where does a Terraform state lock physically live for the S3, GCS and local backends, and which of those actually protects a team?

level: middleimportance: should knowfreq 52%

answer

  1. the backend implements it, not core
  2. atomic create-if-absent somewhere
  3. an object beside the state key
  4. DynamoDB item keyed LockID
  5. a queue instead of an object

basics

~20 s

The lock is a backend feature, not a Terraform one. The S3 backend writes a .tflock object next to the state (use_lockfile) or, historically, a DynamoDB item; GCS writes a lock object; the local backend takes an OS file lock that only covers one machine.

solid answer

~50 s

Terraform core asks the backend to lock; each backend implements that with whatever primitive its storage offers. Since Terraform 1.10 the S3 backend can set `use_lockfile = true`, which uses S3 conditional writes to create a `.tflock` object beside the state key; before that, and still supported for a while, you pointed `dynamodb_table` at a table whose partition key is `LockID` and the lock was a conditional item write. GCS writes a lock object in the same bucket. The azurerm backend takes a blob lease. HCP Terraform does not need an object at all — it queues runs per workspace, so serialisation is a property of the run queue. The local backend takes an operating-system file lock, which stops two processes on one machine and does nothing about two laptops. That last case is the one that bites teams: local state feels like it locks, and across people it does not.

code

hcl · 9 lines
hcl
terraform {
  backend "s3" {
    bucket       = "example-tfstate"
    key          = "prod/network/terraform.tfstate"
    region       = "us-east-1"
    encrypt      = true
    use_lockfile = true
  }
}

go deeper

for a junior

Know that the lock lives wherever the state lives, and that remote backends like S3 and GCS handle it for you once configured. Be able to name one concrete example.

for a middle

Explain the primitive each backend uses — a conditional object write, a DynamoDB conditional item, a blob lease — and why an atomic create-if-absent is the requirement. Say clearly what the local backend does and does not cover.

for a senior

Be ready to audit an inherited repository: confirm the backend locks, that callers have permission to write the lock, and that no pipeline bypasses it. Distinguish a permissions failure from real contention by the error text.

for a principal

Own the migration decision — moving an estate from DynamoDB locking to native lock files, or off object-store backends onto a run-queue platform — and articulate what each choice costs in operational surface and in blast radius when a lock goes wrong.

## Locking is delegated, not implemented in core Terraform core defines an interface — take a lock, release a lock, report who holds it — and each backend implements it with whatever atomic primitive its storage system provides. There is no Terraform lock server. That is why the answer to "how does locking work" is always "which backend?", and why a backend that cannot offer an atomic conditional write cannot offer locking. ## S3: the lock object, and the DynamoDB era S3 stores objects; for a long time it could not do a conditional create, so the S3 backend borrowed one from DynamoDB. You created a table with a string partition key named exactly `LockID`, pointed the backend at it, and Terraform wrote an item keyed by `<bucket>/<key>` with a conditional expression that fails if the item already exists. Losing the race meant the conditional write failed, which surfaced as `Error acquiring the state lock`. The same table also held a `<bucket>/<key>-md5` digest item that Terraform used to notice a state file changed underneath it. S3 later gained conditional writes, and Terraform 1.10 added `use_lockfile = true` to the S3 backend: the lock becomes an object at `<key>.tflock` in the same bucket, created with an if-none-match precondition. No second service, no extra IAM surface, no table to forget to provision: ```hcl terraform { backend "s3" { bucket = "example-tfstate" key = "prod/network/terraform.tfstate" region = "us-east-1" use_lockfile = true } } ``` The DynamoDB path still works in existing configurations, but new state should use the native lock file; HashiCorp has moved the `dynamodb_table` attribute onto the deprecation path. A practical migration note: while you switch, run with both configured for one cycle so a run using the old mechanism and one using the new do not both think they hold the lock. ## GCS and azurerm The GCS backend writes a lock object alongside the state in the same bucket, using object-generation preconditions to make creation atomic. The azurerm backend uses a different primitive again: it takes a *lease* on the state blob. A lease is time-bounded at the storage layer, which is why a killed run against azurerm can leave a blob you cannot overwrite until the lease is broken. The shape is the same everywhere — an atomic create-if-absent on some object that stands for "a run is in progress", carrying the metadata Terraform prints when acquisition fails. ## HCP Terraform: a queue instead of an object HCP Terraform (and Terraform Enterprise) hold state server-side and queue runs per workspace. Only one run is active at a time by construction, so there is no lock object to inspect or delete; the equivalent of a stuck lock is a run stuck in the queue, and the equivalent of breaking the lock is cancelling that run through the API or UI. If you are asked to compare, this is the cleanest contrast: an object-store backend simulates a mutex with a conditional write, while a run-queue backend never allows two concurrent executions in the first place. ## The local backend, and why it misleads people The local backend is not lock-free: it takes an operating-system file lock on the state file. That genuinely prevents two `terraform apply` processes on the *same* machine from interleaving. What it cannot do is coordinate across machines — two engineers each holding their own copy of `terraform.tfstate` are not sharing anything to lock, so they are not racing on one file so much as maintaining two divergent records of the same infrastructure, which is strictly worse. Local state on a shared network filesystem is the worst of both: file locking semantics over NFS are unreliable, so the lock may appear to succeed for both runs. ## What this means in practice Two checks are worth making on any repository you inherit. First, does the configured backend lock at all, and is anything passing `-lock=false`? Second, can everyone who runs Terraform actually write the lock — an engineer with read-only access to the DynamoDB table or without `s3:PutObject` on the `.tflock` key gets an acquisition failure that looks like contention but is really an IAM problem. Reading the exact error text distinguishes them: a permissions failure names the API call, while genuine contention prints the `Lock Info` block with somebody else's identity in it.

  • What did the DynamoDB table need to look like for S3 backend locking?
    A table with a single string partition key named exactly `LockID`, and nothing else required. Terraform wrote a conditional item keyed by `<bucket>/<key>` for the lock, plus a `-md5` digest item to detect the state changing underneath it. On-demand capacity is fine — the write volume is one item per run.
  • Why can a killed run against the azurerm backend leave state unwritable for a while?
    That backend locks by taking a lease on the state blob, and a lease is enforced by the storage service with its own duration. If the process dies without releasing it, writes are refused until the lease expires or is broken, so recovery involves the storage layer, not just deleting an object.

saying these in an interview costs you the question

  • Thinks Terraform runs a central lock service of its own
  • Believes DynamoDB stores the state, not just the lock
  • Says the local backend does no locking whatsoever
  • Assumes every backend supports locking automatically
  • Confuses an IAM write denial with lock contention

context