Two engineers start an apply at the same time against the same shared IaC state record. What can go wrong, and why is locking the standard answer?
answer
- read, change the world, write back
- both hold the same stale snapshot
- last writer wins, first writer orphans
- mutual exclusion for the write-bearing run
- protects the record, not the cloud
basics
~20 sEach run reads the record, changes real infrastructure, then writes the record back. Interleaved, the later write overwrites the earlier one, so resources the first run created vanish from the record and become unmanaged orphans. A mutual-exclusion lock serialises the write-bearing operation.
solid answer
~50 sA run is a read-modify-write cycle over shared mutable data, so it has the classic lost-update problem. Run A and run B both read the record; A creates a load balancer and writes the record; B, still holding its stale copy, writes over it. The load balancer now exists, is billed, and appears nowhere — the next run will happily create a second one. Worse variants have both runs acting on the *same* resource concurrently, so the provider sees conflicting updates and the record ends up describing neither outcome. The fix is mutual exclusion: before any operation that can write the record, the tool acquires an exclusive lock in whatever system holds it, keeps it for the duration, and releases it at the end. Everyone else waits or fails fast. The important caveat is scope — the lock protects the record, not the infrastructure. Anything changing those resources by another route is unaffected.
go deeper
Know that the record is shared and that two runs at once can damage it, and that tools take a lock so only one apply proceeds at a time. Never force a lock you did not investigate.
Walk through the interleaving explicitly: both runs read the same snapshot, the later write wins, and resources created by the earlier run are left out of the record as orphans. Explain why the lock lives with the record's storage.
Show operational judgment on a stuck lock — identify the holder, confirm the process is dead, then reconcile the record afterwards — and be clear that the lock bounds record corruption, not who can change the infrastructure.
Own the design: who is allowed to write each record, how pipeline concurrency and the lock layer, how coarse a lock the estate can tolerate before unrelated changes queue behind each other, and what the break-glass path looks like during an incident.
## Why a run is a critical section Strip an apply down and it is three steps: read the recorded state, make real changes to the world, write the record back. That is a read-modify-write over shared mutable data — the shape that always needs concurrency control. What makes it nastier than a database row is that the "modify" step has external side effects that cannot be rolled back by discarding a write. If the record write is lost, the infrastructure it described still exists. ## The interleaving, concretely Run A and run B start seconds apart against the same record. 1. A reads the record; B reads the record. Both hold the same snapshot. 2. A creates a load balancer and a security group, then writes the updated record. 3. B creates a queue and writes *its* record — derived from the snapshot it read in step 1, which knows nothing about A's two resources. The result: two real resources exist that no record mentions. They are billed, they are in the traffic path, and nothing manages them. The next run sees them missing from the record, concludes they were never created, and proposes to create duplicates — which may then fail on a name collision, leaving the pipeline red for reasons that have nothing to do with the change being deployed. A second, uglier variant: both runs decide to modify the same resource. The provider applies whatever arrives, in whatever order, sometimes rejecting the second call because the object is mid-update. Now the record describes an outcome that never existed, and the next diff proposes a change that makes no sense to the person reading it. ## What locking actually is Before any operation that may write the record, the tool acquires an exclusive lock — implemented in whatever system stores the record, using whatever primitive that system offers: a conditional write, a lease, a row, a stack-level operation guard on a hosted service. It holds the lock for the whole run and releases it at the end. A second run either blocks for a bounded time or fails immediately with "the state is locked", usually printing who holds it, since when, and from where. Two design points are worth naming in an interview. First, the lock lives with the record, not with the tool: it must be visible to every client that could write, which is why a record on a laptop cannot be safely shared. Second, read-only operations that never write the record do not need the lock, which is what lets a plan-style preview run alongside a queued apply. ## What a lock does not cover **Anything not going through this record.** A console edit, another team's pipeline, an autoscaler adjusting a capacity attribute — none of these touch your record, so none of them are excluded. Concurrency control here is about corrupting the record, not about being the only actor on the infrastructure. **Another configuration managing the same object.** If two records both believe they own a resource, both runs hold their own lock and fight anyway. That is an ownership defect, not a locking one. **Everything downstream of your lock's scope.** The lock is as coarse as the record: every resource tracked in it is serialised together, so an unrelated change waits behind a slow one. ## Stale locks A CI job killed mid-apply — cancelled build, evicted runner, expired credential — can leave a lock held by nobody. Every later run then fails to acquire it. Tools provide an escape hatch to release a lock by force, and it is genuinely dangerous: if the original process is still alive, forcing the lock recreates the exact race the lock existed to prevent. The correct sequence is to identify the holder from the lock metadata, confirm that process is really dead (check the CI run, the host, the user), and only then force the release — followed by a careful review of the record, because the interrupted run may have created resources it never recorded. ## Serialising above the tool Mature setups do not rely on the lock as the primary mechanism; they arrange for the collision not to happen. Pipelines use a concurrency group per environment so only one apply for a given record runs at a time, and queue the rest. Humans are kept out of applying to production entirely, so the pipeline is the only writer. The lock stays as the backstop that catches what process forgot — including the case where someone applies from a laptop during an incident. Stateless tools have no record to corrupt, but they are not exempt: two concurrent configuration runs against one host interleave file writes and service restarts, producing a half-applied configuration. The remedy is the same, moved up a level — serialise the runs.
- Your CI job was killed mid-apply and every later run now fails with "state is locked". What do you do?Read the lock metadata first: who acquired it, when, and from which job or host. Confirm that process is genuinely dead — a cancelled build page, a terminated runner — before forcing the release, because forcing a live lock recreates the race. After releasing, review the record against reality: an interrupted run may have created resources it never got to record, and those need adopting or removing before the next apply.
- Does a lock protect you from someone changing the same resources in the cloud console?No. The lock is a mutual exclusion over the record, so it only excludes other clients of that record. A console edit, another team's pipeline with its own record, or an autoscaler adjusting a field are all invisible to it. Those changes surface later as drift, which is a detection-and-remediation problem, not a locking one.
- Why can a preview run alongside a queued apply, but two applies cannot run together?Because the lock is needed only by operations that write the record. A read-only preview reads the record and the provider and produces a diff without persisting anything, so it cannot lose another run's write. Its result can of course be invalidated by whatever the concurrent apply is doing, which is why a preview is a proposal rather than a promise.
- If pipeline concurrency groups already serialise applies, is locking redundant?No — they defend different failure modes. The concurrency group prevents the collision inside one pipeline; the lock catches everything outside it: an engineer applying from a laptop during an incident, a second pipeline in another repository, a manual re-run of an old job. Process prevents the common case, the lock is the backstop for the case process did not anticipate.
saying these in an interview costs you the question
- The tool merges two concurrent state writes automatically
- The cloud API rejects the second run, so nothing is lost
- A lock stops anyone from changing the infrastructure
- Just force-release a stuck lock and re-run
- Concurrency isn't an issue if everyone applies from CI