skip to content

Why does a lease key that lets only one worker run a job carry a lifetime instead of living until its holder releases it?

level: juniorimportance: must knowfreq 68%

answer

  1. release is only the happy path
  2. holders die between claim and release
  3. a claim with no deadline is permanent
  4. the lifetime bounds the stall, not the job

basics

~20 s

A lifetime bounds a crash. If the holder dies between claiming and releasing, an entry that only a release deletes blocks the job forever; the lifetime lets the store drop the claim so another worker can take it.

solid answer

~40 s

The claim is made with a **conditional create** — a write that succeeds only if the key is absent, in one operation — so exactly one worker wins it and the others go away. Deleting the key on completion is the happy path, and the happy path is not guaranteed to run: a worker can be killed mid-job, lose the network, or hang. A claim with no lifetime survives all of that, so the job silently never runs again until somebody deletes the key by hand. Attaching a lifetime turns a permanent outage into a bounded stall: at worst one lifetime after the crash the claim is gone and the next worker takes it. The claim disappearing is not evidence the job finished — it says only that nobody renewed it.

go deeper

for a junior

Remember the two states: a claim that exists means somebody is working, and it must be able to disappear on its own. Be ready to say what happens if the holder is killed before it releases.

for a middle

Explain that the claim is made with a conditional create so only one worker wins, and that the lifetime exists because release code is not guaranteed to run.

for a senior

Say how you pick the number — from the longest observed run and the stall you can tolerate — and say out loud that renewal is the alternative to guessing high.

for a principal

Frame the lifetime as the price of a bounded outage: long hides a stalled job, short buys concurrent runs. Decide which of the two the business can absorb, and say what each costs.

## What a lease is here A **lease** is an entry in a shared volatile tier whose existence means *somebody is working on this*. A worker that wants the job attempts a **conditional create** — a write that succeeds only if the key is absent, in one operation. Whoever wins the create runs the job; everyone else sees the create fail and exits. Nothing about the entry's contents matters yet: its presence is the whole signal. The obvious next step is to delete the key when the job ends, and that is indeed the happy path. The interesting question is what happens when the happy path does not run. ## Release is the path you cannot count on A release is ordinary application code, and ordinary application code does not always execute: 1. The process is killed outright — the host is reclaimed, the container is evicted, a deploy rolls, a supervisor times it out. No cleanup block runs. 2. The process survives but the path to the store does not, and the delete never lands. 3. The process neither dies nor finishes: it hangs on a read with no timeout and sits there for hours. 4. The job throws somewhere the author did not wrap, and the delete is skipped. In each case an entry created with no deadline stays in the keyspace indefinitely. The failure that produces is unusually quiet: nothing crashes. Every later worker attempts the conditional create, fails, logs *somebody else holds it*, and exits successfully. The job simply stops happening. The first real signal is downstream and late — a report that stopped updating, a cleanup that stopped running, a backlog that stops draining — and the repair is a human deleting a key by hand. ## What the lifetime actually promises Attaching a lifetime to the claim converts that permanent, silent outage into a **bounded stall**: - The store, not the worker, measures the deadline. It passes whether or not the job finished, whether or not the worker is healthy, and whether or not anybody is watching. - After a crashed holder, the work resumes at worst one lifetime later, with no human involved. - The lifetime belongs to the entry, not to the job. It does not stretch because the holder is busy, and it does not pause while the holder is blocked. - The claim disappearing is **not** evidence that the job completed. It says only that nobody renewed it. That last point is the one candidates miss. "The claim is gone" and "the job is done" are different statements, and conflating them is how a team ends up treating the tier as a record of work rather than as a hint about who is working. ## Choosing the number The lifetime is a single knob with a cost on both sides. | Lifetime chosen | After a crashed holder | While an unusually slow run is in progress | |---|---|---| | Far longer than the longest run | The job does not run for up to that long, silently | Comfortable: the claim outlives the work | | A little longer than the longest observed run | Stall is short and bounded | One slow run loses its claim mid-job | | Shorter than a normal run | Barely any stall | Routinely two workers on one job | No single value is good in both columns, which is why the honest answer at any scale is a short lifetime that the holder **renews** while it is healthy. Renewal moves the requirement from *longer than the worst run* to *longer than the worst pause between renewals*, which is a smaller and much more stable number. ## What the store has to offer for the recipe to hold The recipe assumes more than every store of this class provides, and it is worth naming which parts: - **A create conditional on absence, in one operation.** Where a store has no such operation, a read followed by a write leaves a gap in which two workers both see an absent key and both write. - **A lifetime attached in the same operation as the create.** Where the lifetime must be a second call, a worker that dies between the two calls leaves exactly the permanent claim the lifetime was meant to prevent, so the two-call version is weaker than it looks. - **A tier that will not remove the entry early.** Where the tier is configured to drop entries once it reaches its memory ceiling, a claim that is still in force can vanish and the holder is told nothing. ## The consequence when the claim is not there When the claim is absent while a worker is still running, the loss is specific: **two workers run the same job at the same time**. If that is duplicated effort — a report built twice, a scan run twice — the cost is waste and the design is fine. If the job has an outside effect that must not happen twice, the claim is an optimisation and the real guard has to live where that effect is made.

  • What is the visible symptom of a claim with no lifetime whose holder was killed?
    Nothing crashes. The job stops happening: every later conditional create fails, each worker logs that somebody else holds it and exits cleanly, and the first real signal is a stale downstream artefact — a report that stopped updating, a backlog that stops draining. Silence is the symptom, which is why the lifetime is not optional.
  • Is it acceptable to create the key first and attach the lifetime with a second call?
    It is much weaker. If the worker dies between the two calls, the claim has no deadline and is permanent — the exact failure the lifetime was meant to prevent. The claim and its lifetime should be one operation. Stores differ in whether they offer that, and where a store does not, the gap is a real argument for coordinating somewhere else.

saying these in an interview costs you the question

  • Says a cleanup block on exit makes a lifetime unnecessary
  • Treats the claim's disappearance as proof the job completed
  • Picks a lifetime from habit rather than from the longest observed run
  • Assumes the deadline extends while the holder is still busy
  • Claims one lifetime value is right for every job