A CI job was killed mid-apply and every Terraform run since fails with 'Error acquiring the state lock'. How do you clear it, and what must you verify before running terraform force-unlock?
answer
- read Lock Info before acting
- who and created, not elapsed time
- prove the CI job is terminal
- it only deletes the record
- plan after, never blind re-apply
basics
~20 sRead the Lock Info block, then prove the run that took the lock is really dead — check the CI job and the named principal — before running terraform force-unlock with that exact lock ID. Force-unlocking a live apply lets a second run race it and lose state writes.
solid answer
~50 sThe error prints a `Lock Info` block: lock ID, operation, who, and when it was created. Treat that as evidence, not noise. Go find that run: was the CI job cancelled or is it still executing, is the runner process gone, is the human named in `Who` mid-apply on a laptop? Only when you can say the holder is definitively dead do you run `terraform force-unlock <ID>` from the same working directory and backend configuration, so it targets the right state. force-unlock only deletes the lock record — it does not roll back, finish, or undo anything the dead run did, so the next step is a plan to see what state the infrastructure was left in. If you break a lock while an apply is genuinely running, you get exactly the race the lock exists to prevent: two runs mutating the same infrastructure, and the survivor's state write clobbering or being clobbered by the other's.
code
bash · 10 lines# 1. Confirm nothing is holding the lock for real: inspect the lock record itself
aws dynamodb get-item \
--table-name terraform-locks \
--key '{"LockID":{"S":"example-tfstate/prod/network/terraform.tfstate"}}'
# 2. Only after confirming the holder is dead, break it with the exact ID
terraform force-unlock 4f2c9c17-9a6d-4c94-a09b-0e4a0dd9e0b5
# 3. Never re-apply blind - read the plan first
terraform plango deeper
Know that the error prints a lock ID and that terraform force-unlock <ID> clears it — and that you check with whoever or whatever holds it first rather than clearing it yourself.
Explain each field of the Lock Info block and why the ID must match. Be clear that force-unlock only removes the record and changes nothing about the infrastructure or the state contents.
Demonstrate the verification discipline: positively confirm the holding run is terminal, then unlock, then plan and reconcile before applying. Describe concretely what goes wrong if you break a live lock and how you would recover.
Own the systemic fix — graceful cancellation so runs release their locks, per-state run serialisation so jobs queue instead of failing, and a documented runbook so a stuck lock is not improvised at 3am by whoever is on call.
## Read the lock before you break it A held lock is not an obstruction to clear reflexively; it is a claim by another run, with the evidence attached: ``` Error: Error acquiring the state lock Lock Info: ID: 4f2c9c17-9a6d-4c94-a09b-0e4a0dd9e0b5 Path: example-tfstate/prod/network/terraform.tfstate Operation: OperationTypeApply Who: runner@ci-7f4b Created: 2026-03-04 10:12:41 UTC Version: 1.11.2 ``` Four fields drive the decision. `Operation` tells you whether real changes were in flight (`OperationTypeApply`) or only a refresh (`OperationTypePlan`) — an abandoned plan lock is nearly always safe to clear, an apply lock deserves care. `Who` names the machine and user. `Created` tells you how long it has been held: three minutes into a fifteen-minute apply is a very different situation from eleven hours. `ID` is the argument you will need. ## Prove the holder is dead This is the one verification the question is really about. Locks do not time out on their own for the common object-store backends — Terraform releases the lock when the run ends, so a process killed with SIGKILL, an evicted container, a runner that lost its host, or a laptop that closed mid-apply leaves the record behind forever. Nothing expires it for you. So establish death positively, not by assumption: - Find the CI run that matches `Who` and `Created` and confirm it reached a terminal state. "The job page shows cancelled" is good evidence; "nobody replied in chat" is not. - If it was a human, ask them. A colleague halfway through a fifteen-minute apply is the classic case where the timestamp looks stale but the run is alive. - If you can, look at the infrastructure: a still-running apply is making API calls, so recent CloudTrail activity from the CI role, or a resource visibly mid-creation, argues the run lives. - Never use elapsed time alone as proof. Long applies exist. "It has been an hour, it must be dead" is precisely how a lock gets broken on a live run. ## Break it correctly ```bash terraform force-unlock 4f2c9c17-9a6d-4c94-a09b-0e4a0dd9e0b5 ``` Run it from the same working directory, after `terraform init` against the same backend configuration and the same workspace — force-unlock acts on the state your configuration points at, so running it from the wrong directory either errors or, worse, unlocks a different state. Terraform prompts for confirmation; `-force` skips the prompt and is for scripts, not for hurrying. The ID must match the held lock. That mismatch check is a safety feature: if the ID you pass is not the one currently held, Terraform refuses, which is what stops you from clearing a *newer* lock taken by a healthy run that started while you were investigating. Deleting the backing object by hand — removing the DynamoDB item or the `.tflock` object — does the same thing with none of those checks. Keep it as a genuine last resort, for example when the recorded lock ID has been lost or the backend is in a state force-unlock cannot address. ## What force-unlock does not do It deletes a lock record. That is all. It does not roll back partial changes, does not resume the dead apply, and does not repair state. The dead run may have created resources and died before persisting them, in which case state is behind reality. When an apply cannot write state to the backend at all, Terraform writes the result locally as `errored.tfstate` and tells you to push it — on a CI runner that file usually dies with the container, which is another reason a crashed apply needs a plan afterwards rather than a re-run on faith. So the sequence is: verify, unlock, then `terraform plan` and *read it*. Expect creations of things that already exist (state missed them), or no changes at all (the apply had finished its work and died during cleanup). Reconcile before re-applying — a blind re-apply after a crash is how duplicate resources get made. ## If you break a live lock Both runs now mutate the same infrastructure. They will issue conflicting API calls, and both will try to persist state at the end. One write lands and the other overwrites it, so a whole run's worth of resource records vanishes from state while the resources remain real and billed. The still-running original may also fail its final state write outright and leave `errored.tfstate` behind. Recovery means restoring a state version from the backend's history if nothing has been applied since, or importing the orphans back individually — hours of work to save the ten minutes it would have taken to check whether the job was still running. ## Preventing the recurrence Stuck locks are a symptom. Make CI cancellation graceful where you can, so Terraform gets a chance to release the lock instead of being killed outright; serialise runs per state so a queued job waits instead of failing; and make the lock error surface the `Lock Info` in the job log rather than being swallowed, so the next person starts the investigation with the evidence already in hand.
- Why does terraform force-unlock require the lock ID rather than just clearing whatever is there?Because the ID pins the operation to the specific lock you investigated. If a healthy run acquired the lock while you were checking, your ID no longer matches and Terraform refuses — so you cannot accidentally break a run that started after the dead one. Deleting the backend object by hand throws that check away.
- After clearing a lock left by a crashed apply, what do you expect the next plan to show?Anything from no changes to several creates. If the apply created resources but died before persisting state, the plan proposes creating them again — that is your cue to import rather than apply. Read the plan against what the crashed run's log got through; do not assume the crash means nothing happened.
- Do state locks expire on their own if a run never releases one?Not for the common object-store backends — the record simply persists until a run releases it or someone force-unlocks. The azurerm backend is the notable exception, because it locks with a blob lease that the storage service enforces and expires. Never rely on expiry as a recovery mechanism.
saying these in an interview costs you the question
- Force-unlocks immediately because the lock 'looks old'
- Uses elapsed time as proof the holder is dead
- Thinks force-unlock rolls back the failed apply
- Deletes the DynamoDB item as the first step
- Re-runs apply straight after unlocking without a plan