A cron-style job runs on three nodes and must execute on exactly one. Explain how acquiring a Consul KV key with a session — `PUT /v1/kv/service/job/leader?acquire=<sessionID>` — elects a leader, what happens to that lock when the session is invalidated, and what LockDelay protects against.
answer
- the key is contested, the session is the lease
- liveness comes from checks, not from the holder
- invalidation is not the same as stepping down
- a pause before the next holder may start
- nothing stops a process that ignores it
basics
~20 sEach node creates a session and tries to acquire the same key; exactly one acquire returns true and that node is leader. If the session is invalidated — the agent fails, a check goes critical, or the TTL lapses — Consul releases or deletes the key, and LockDelay blocks re-acquisition briefly so the old holder cannot still be acting.
solid answer
~50 sEach candidate creates a session with `PUT /v1/session/create`, then does `PUT /v1/kv/service/job/leader?acquire=<sessionID>`. Exactly one gets `true`; the key records that session and its `LockIndex` increments. The losers keep a blocking query on the key and take over when it frees. What makes it a *lease* rather than a flag is the session: Consul invalidates it when the agent's `serfHealth` check fails, when an associated health check goes critical, or when a TTL is not renewed — and on invalidation the key is released (`Behavior: release`) or deleted (`Behavior: delete`) without the holder's cooperation. `LockDelay`, 15 seconds by default, then blocks re-acquisition for that window *after an invalidation*, though not after a clean release, so a partitioned old leader that has not yet noticed cannot overlap with a new one. Critically, these locks are **advisory**: Consul stops nobody from doing work. Make the job idempotent, or carry `LockIndex` as a fencing token the protected resource checks.
code
bash · 27 lines#!/usr/bin/env bash
# Elect a leader with a session-backed lock, by hand.
set -euo pipefail
ADDR=http://127.0.0.1:8500
KEY=service/job/leader
sid=$(curl -s -X PUT --data '{
"Name":"nightly-report","TTL":"15s","LockDelay":"15s",
"Behavior":"release","Checks":["serfHealth"]}' \
"${ADDR}/v1/session/create" | jq -r .ID)
cleanup() { curl -s -X PUT "${ADDR}/v1/session/destroy/${sid}" >/dev/null; }
trap cleanup EXIT
# 200 for both outcomes: the body is true or false.
if [ "$(curl -s -X PUT --data "$(hostname)" \
"${ADDR}/v1/kv/${KEY}?acquire=${sid}")" != "true" ]; then
echo "not the leader; standing by"; exit 0
fi
# Renew for as long as we hold it.
while true; do
sleep 5
curl -s -X PUT "${ADDR}/v1/session/renew/${sid}" >/dev/null
done &
./run-report.shgo deeper
Know that Consul can elect a leader by having each candidate try to acquire the same KV key with a session, and that exactly one acquire succeeds while the others wait for the key to free.
Explain the session as a lease: it is bound to checks such as serfHealth or to a TTL, and Consul invalidates it without the holder's cooperation, releasing or deleting the key according to the session's behavior.
Show why LockDelay exists — invalidation means unreachable, not stopped — and that the locks are advisory, so real safety comes from idempotent work, re-checking leadership before side effects, or a fencing token built on LockIndex.
Own the numbers and the fallback: the TTL plus its grace plus LockDelay is the window in which nobody runs the job, and the same window is where two holders can overlap. Decide whether the workload should depend on distributed locking at all.
## The two objects Leader election in Consul combines a KV key with a *session*. The key is the contested resource; the session is the liveness handle that lets Consul take the key back from a dead holder. ```json PUT /v1/session/create { "Name": "nightly-report", "TTL": "15s", "LockDelay": "15s", "Behavior": "release", "Checks": ["serfHealth"] } ``` The session is bound to the agent that created it and to the checks you list. By default that includes `serfHealth`, the agent's own gossip-derived liveness check, so a node that dies or is partitioned loses its sessions without doing anything. You can add your application's own health check to the list, which is stronger: a process that is running but broken loses the lock too. A `TTL` adds a renewal contract on top — the holder must call `PUT /v1/session/renew/<id>` periodically, and Consul may take up to twice the TTL to invalidate a lapsed one, so size TTLs with that grace in mind. ## Acquiring ``` PUT /v1/kv/service/job/leader?acquire=<sessionID> ``` The body is whatever you want the key to hold — usually the node's identity, so observers can see who leads. The response is `true` or `false`, both under HTTP 200, so parse the body rather than the status. On success the entry's `Session` field names the holder and `LockIndex` increments; `LockIndex` therefore counts how many times the lock has changed hands over the key's lifetime and only ever moves forward. The losing candidates do not spin. They put a blocking query on the key and are woken the instant the `Session` field clears, at which point they race to acquire again. That gives you failover in roughly the time it takes to detect the failure, not a polling interval. ## Losing the lock There are two very different exits. A **clean release** — `PUT /v1/kv/service/job/leader?release=<sessionID>`, or destroying the session — frees the key immediately. The holder chose to step down, so there is no ambiguity about whether it is still working, and re-acquisition can happen at once. An **invalidation** is Consul deciding the holder is gone: `serfHealth` failing, a listed check going critical, or the TTL lapsing. What happens to the key then depends on the session's `Behavior`. With `release` (the default) the lock is freed and the value is kept, which is what you want for leader election. With `delete` the key is removed entirely, which is how you build ephemeral registrations that disappear with their owner. ## Why LockDelay exists This is the part candidates most often miss. When Consul invalidates a session, it has concluded the holder is unreachable — it has not confirmed the holder has stopped. A partitioned leader may still be mid-task, still believing it holds the lock, because it has not yet observed that its own agent lost contact. If a new leader could acquire instantly, both would run. `LockDelay` (15 seconds by default) blocks any acquisition of that key for the delay window following an invalidation. It does not apply to a clean release. The window is a heuristic, not a proof: it gives the old holder time to notice and stop before a new one starts. Setting it to `0` disables it, which is defensible only when the work itself is safe to run twice. ## Advisory, not enforced Consul locks are advisory. Nothing prevents a process that never asked for the lock — or one whose lock was silently invalidated — from writing to the database, sending the emails, or running the report. The lock coordinates only among participants that agree to check it. Three ways to close the gap, in increasing strength: 1. **Make the work idempotent.** Running the job twice produces the same end state. Easiest and most robust. 2. **Re-check leadership before each side effect,** not just once at startup. A long job that verified leadership an hour ago is asserting nothing about now. 3. **Fence.** Carry the key's `LockIndex` — which only increases — into the protected resource, and have that resource reject writes stamped with an index lower than the highest it has seen. This is the only approach that actually stops a stalled old leader that wakes up and resumes, but it requires the downstream system to cooperate. ## What you usually run instead The CLI wraps the whole protocol: `consul lock service/job/leader ./run-report.sh` creates the session, acquires, runs the child while holding the lock, and terminates the child when the lock is lost. `consul lock -n 3 <prefix> <cmd>` builds a semaphore instead, allowing up to that many concurrent holders. Most language clients ship an equivalent helper. Writing the loop by hand is worth doing once to understand renewal and invalidation, and worth avoiding afterwards. ## Sizing the knobs A short TTL fails over fast and risks losing the lock during a stop-the-world pause or a brief network hiccup; a long one leaves the job stalled for longer after a genuine crash. The numbers to reason about together are the TTL plus its 2x grace, then LockDelay on top — that sum is the worst-case gap during which nobody is running the job, and simultaneously the window a wrongly-invalidated leader has in which to notice and stop.
- What is the difference between Behavior release and Behavior delete on a session?On invalidation, release frees any keys the session holds while leaving their values intact — the right choice for leader election, since the key survives to record who leads next. Delete removes those keys outright, which makes them ephemeral: they exist exactly as long as their owner does. That is how you build a presence list where entries vanish when the process behind them dies.
- Your job holds the lock and its node is partitioned from the cluster. What does each side believe?The cluster sees serfHealth fail, invalidates the session and frees the key, so after LockDelay a new leader acquires it. The partitioned node still believes it is leader until its own agent tells it otherwise. That overlap is real, and it is why LockDelay is a mitigation rather than a guarantee — safety has to come from idempotent work, re-checking leadership before each side effect, or fencing.
- How does LockIndex work as a fencing token?It increments on every successful acquire of that key and never decreases, so each leadership term has a strictly higher number than every term before it. If the leader stamps its writes with that number and the protected resource rejects anything stamped lower than the highest it has accepted, a stalled old leader that resumes is refused. It only works if the downstream system participates in the check.
- How would you choose the session TTL?Reason about the worst-case gap: Consul may take up to twice the TTL to invalidate a lapsed session, then LockDelay applies on top, and that total is how long nobody runs the job after a crash. Shorter TTLs fail over faster but risk losing the lock to a garbage-collection pause or a brief network stall, costing an unnecessary handover. Pick from the job's tolerance for delay versus for churn.
A session is a lease on a hotel room, not a purchase: you keep it by continuing to check in, and if you stop, the front desk empties the room whether or not you have actually left. LockDelay is the cleaning window before the next guest gets a key.
saying these in an interview costs you the question
- Calls Consul locks mandatory and assumes nobody else can act
- Treats a false acquire response as an error rather than losing the race
- Believes the lock frees instantly when a node dies, with no delay
- Never renews a TTL session and wonders why leadership drops
- Confuses check-and-set on a key with mutual exclusion