Two pipeline runs for the same Terraform root module start at the same time and the second one fails because it cannot acquire the state lock. How should the pipeline be designed so this stops being a problem?
answer
- the lock is a guard, not a scheduler
- one lane per state, not per repo
- queue on apply, cancel only on plan
- killed apply leaves an orphan lock
- contention is a state-layout smell
basics
~20 sSerialise runs per state in the pipeline itself: one concurrency lane per root module, queueing rather than cancelling, so a second run waits its turn instead of racing. The state lock is a last-resort safety net, not a scheduler.
solid answer
~50 sThe lock did its job — it stopped two writers corrupting one state — but relying on it as your queue means every collision surfaces as a red build. Give each root module its own concurrency group in the pipeline so runs against the same state queue behind each other, and let a short `-lock-timeout` absorb the small overlaps rather than failing instantly. The important rule is never to cancel an apply that is already running: killing the process mid-apply can leave the lock held by a job that no longer exists, and possibly resources created that the state never recorded. So queue-and-wait is safe for apply, cancel-in-progress is only ever acceptable for plan-only runs on superseded commits. If throughput genuinely matters, the real fix is architectural — split the estate into more, smaller states so independent changes stop contending for the same lock at all.
go deeper
Know that Terraform locks state during plan and apply so two runs cannot write it at once, and that a lock error usually means someone else's run is in progress — you wait, you do not disable locking.
Explain how the pipeline can queue runs per state and what -lock-timeout changes, and why -lock=false is not an acceptable way to make the red build go away.
Show you have cleaned this up: never cancelling an apply, the orphaned lock and partially applied resources a killed run leaves, and force-unlock as a human runbook step with a careful plan afterwards.
Own the layout that removes contention — how finely to split states by lifecycle and ownership, what queue depth you consider acceptable per team, and the runner guarantees an apply job requires.
## What the failed run is telling you Terraform takes a lock on the state before any operation that could write it, and refuses to proceed while another operation holds it. A second pipeline run failing to acquire it is not a defect: it is the mechanism working. The defect is that your pipeline treated the lock as its scheduling primitive. Locks are designed to prevent corruption, and their failure mode — an immediate error — is a terrible user experience for something as routine as two people merging within a minute of each other. ## Serialise before the lock, not at it Every mainstream CI system can constrain concurrency by a key. The key you want is the identity of the **state**, not the repository and not the branch: one lane per root module (and per workspace, if you use them), so runs that touch the same state form an orderly queue while runs that touch different states stay fully parallel. Get the key wrong in either direction and you either serialise the whole repository unnecessarily, or let two lanes converge on one state anyway. A second, complementary knob is `-lock-timeout`, which tells Terraform to retry acquiring the lock for a period instead of giving up at once: ``` terraform apply -input=false -lock-timeout=5m tfplan ``` That absorbs the overlap where one run is finishing as another starts. It is a smoother, not a substitute: with no pipeline-level queue, a busy state just moves the failure from "immediately" to "five minutes later". ## The cancel rule Many pipelines are configured to cancel a superseded in-flight run when a new commit arrives. For plan-only runs that is fine and even desirable — the old plan is worthless. For an **apply** it is dangerous, and this is the part interviewers are listening for. Killing Terraform mid-apply can leave two kinds of mess. The lock may remain held, because the process that owned it was killed before it could release it, so every later run is blocked by a lock whose owner does not exist. And resources may already have been created or modified by the provider without those results being written back to state, so the next plan proposes to create things that are already there — usually surfacing as a name-already-exists error from the API rather than as anything Terraform can explain. So: never cancel apply. Let it finish, then run the next one. If the runner itself dies — a spot instance reclaimed, an agent crash — you land in the same place involuntarily, which is why apply jobs belong on stable runners with generous timeouts. ## Recovering a stuck lock When a lock is genuinely orphaned, `terraform force-unlock` with the lock ID from the error message releases it. The discipline around it matters more than the command: confirm from the CI run history that no job is still executing before you unlock, because force-unlocking a live run is precisely how two concurrent writers end up in the state you were trying to prevent. Treat it as a manual, human-only action — a pipeline that automatically force-unlocks on failure has disabled its own safety mechanism. After unlocking, run a plan and read it carefully, because the interrupted apply may have left reality ahead of state. ## The architectural answer Contention is a symptom of a state that too many changes touch. If a single root module holds the network, the clusters, the databases and every application's resources, every merge in the repository queues behind every other merge, and the queue grows with the team. Splitting the estate into smaller root modules — by lifecycle and by ownership, with the slow-changing foundations separate from the fast-moving application layer — removes the contention rather than managing it, and shortens plan times as a side effect. That is the fix a senior candidate should reach for once queueing is no longer enough. ## What good looks like A healthy pipeline: one concurrency lane per state, queue rather than cancel on apply, a modest `-lock-timeout` for the edges, apply jobs on runners that will not be reclaimed mid-run, force-unlock as a documented human runbook step rather than an automated retry, and a state layout small enough that most merges do not queue at all.
- Why is cancelling an in-flight apply worse than cancelling an in-flight plan?A plan writes nothing, so cancelling it loses only compute. An apply is mid-flight against real APIs: killing it can leave the state lock held by a dead process, and can leave resources created or modified without those results written back to state. The next plan then proposes to create things that already exist, and the provider fails with a conflict the plan output cannot explain.
- Your pipeline hits a stuck lock every few days after runner crashes. Should it force-unlock automatically?No. Automatic force-unlock removes the very protection the lock provides — if the previous run is actually still alive, you now have two writers on one state. Keep it a documented manual step: check the run history, confirm nothing is executing, unlock with the ID from the error, then plan and read the diff for work the interrupted run half-finished. Fix the crash cause instead.
- Queueing has made merges wait twenty minutes for their turn. What now?Stop treating it as a scheduling problem. A queue that long means one state covers too much: split the root module along lifecycle and ownership lines so the slow foundations and the fast-moving application resources have separate states and separate lanes. Independent changes then run in parallel, and plans get shorter as a bonus.
saying these in an interview costs you the question
- Just disable locking in CI with -lock=false
- Cancel the running apply so the newest commit wins
- Auto force-unlock in a retry step when apply fails
- One concurrency lane for the whole repository is fine
- A lock error means the backend is misconfigured