Pipeline runs keep failing with a held Terraform state lock and a colleague proposes adding -lock=false. Why is that dangerous, and what should you do instead?
answer
- the two flags are not variants
- timeout waits, false bypasses
- a loud failure beats a silent race
- contention is information, not noise
- serialise per state, then split it
basics
~20 s-lock=false disables locking entirely, converting a loud failure into a silent race that can strand real resources outside state. Use -lock-timeout so a queued run waits for the lock instead, and serialise the pipeline so only one run per state executes at a time.
solid answer
~50 s`-lock=false` does not make the lock wait or retry — it stops Terraform taking one at all, so the run proceeds straight into the window the lock exists to close. A concurrent apply then overwrites its state write, leaving real resources untracked. Worse, it hides the symptom: contention keeps happening, you just stop hearing about it until a plan proposes recreating production. The right flag is `-lock-timeout`, for example `terraform apply -lock-timeout=10m`, which makes the run retry acquisition for that duration and only then fail — perfect for a job that is queued behind a legitimate run. Above the flag, fix the design: serialise the pipeline so one run per state is in flight, and if jobs genuinely need to run in parallel, split the state so they are not contending for the same lock at all. The narrow legitimate use of `-lock=false` is a read-only operation against a backend where you cannot obtain a write lock, and even there it is a deliberate exception, not a default.
code
bash · 5 lines# Wrong: races any concurrent run, silently
terraform apply -input=false -auto-approve -lock=false
# Right: queue behind the current holder, fail only if it never frees
terraform apply -input=false -auto-approve -lock-timeout=10mgo deeper
Know that -lock=false turns locking off rather than making the run wait, and that -lock-timeout is the flag for a job that is queued behind another run.
Explain the default of 0s for -lock-timeout and why an unconfigured pipeline fails instantly. Describe the concrete outcome of a bypassed apply: untracked resources and a later plan proposing to recreate them.
Diagnose the contention rather than the error — live holder, dead holder, or two jobs racing — and pick the matching fix: timeout, deliberate unlock, or pipeline serialisation. Explain why the silent failure mode is worse than the loud one.
Own the position that persistent contention is a state-decomposition signal, and set the guardrail: read-only commands may run unlocked, apply may not, enforced by CI role or policy rather than by convention.
## The two flags do opposite things They look like variations on a theme and are not. `-lock-timeout=DURATION` keeps locking on and changes only the failure behaviour: instead of failing the instant the lock is held, Terraform retries acquisition until the duration elapses. The default is `0s`, which is why an unconfigured pipeline fails immediately. `-lock=false` removes the lock from the operation entirely — no acquisition, no wait, no record, no protection. So the response to "our jobs keep colliding" is `-lock-timeout=10m`. The response `-lock=false` is not a fix for collisions; it is a decision to allow them. ## Why bypassing is worse than failing A failed acquisition is the system working. It is loud, it names the holder, and nothing is damaged. Bypassing converts that into two applies against one state, which is the race described by every state-locking question: both read the same serial, both change infrastructure, the last writer persists a state that never saw the other's work. The resources the loser created stay real, keep billing, and are invisible to `terraform destroy`. The next plan proposes creating them again and either duplicates them or fails on a name collision. The damage is also delayed. The bad apply usually succeeds — output is green, the pipeline is unblocked, the change ships. Discovery comes days later when somebody's unrelated plan shows creations nobody asked for. By then nobody connects it to the flag that was added to "unblock CI", and the flag has usually spread by copy-paste to every other pipeline in the repository. There is a second-order effect worth naming in an interview: `-lock=false` removes the only signal that your pipeline concurrency model is wrong. Contention is information. Silencing it means you never learn that two teams share one state, or that a scheduled drift-detection job overlaps the deploy job every Tuesday. ## Fix the contention, not the symptom Work through the causes in order: 1. **Is the lock held by a live run?** Then `-lock-timeout` is the whole answer. A ten-minute window covers a normal apply; pick a value slightly above your p99 run time so a genuine deadlock still fails rather than hanging the pipeline forever. 2. **Is it held by a dead run?** That is a stuck lock, cleared deliberately after verifying the holder is gone — not something a flag should paper over on every subsequent run. 3. **Do two pipeline jobs genuinely run at once against the same state?** Serialise them. Every CI system offers some form of one-run-at-a-time grouping; the grouping key should be the state, not the repository, or you needlessly block unrelated stacks. 4. **Is contention structural — several teams, one state, all day?** No flag helps. That is a state-splitting problem, and the lock is telling you the blast radius is drawn wrong. ## Where -lock=false is defensible Read-only work against a backend where you cannot or should not take a write lock: an auditor with read access to the state bucket inspecting outputs, or a `terraform plan` in an environment where the principal deliberately has no write permission on the lock table. Even then, understand what you have given up — a plan taken without the lock can be computed against a state that another run is rewriting underneath you, so treat its output as advisory. If your team decides to allow it, allow it for read-only commands and make `apply` without a lock impossible, whether by policy check or by the CI role simply not being able to run it. ## What to say in the interview Name the difference between the flags first, because that is the technical content. Then make the operational point: the error is not the problem, the concurrency is, and a flag that removes the error while leaving the concurrency in place trades a visible failure for an invisible one. Finish with the sequencing fix — one run per state, timeout for the queue, split state when contention is structural. That progression is what distinguishes an answer that has operated a pipeline from one that has read the flag reference.
- How would you choose a value for -lock-timeout?A little above the p99 duration of the runs that hold the lock, so a job queued behind a legitimate apply waits it out while a genuinely stuck lock still fails within a reasonable time. Ten minutes suits most stacks. Very long values turn a stuck lock into a hung pipeline that nobody notices.
- Is there any case where you would accept -lock=false in a pipeline?Only for read-only work where a write lock is unobtainable by design — an audit or reporting job whose principal has no write access to the lock. Treat its plan output as advisory, since state may change underneath it, and make apply-without-lock impossible for everyone else.
saying these in an interview costs you the question
- Thinks -lock=false makes Terraform wait for the lock
- Adds -lock=false to unblock CI and leaves it there
- Believes a bypassed apply is safe if the plan looked clean
- Sets an hours-long lock timeout to avoid all failures
- Treats lock contention as noise rather than a design signal