skip to content

In a shared volatile tier, how can a worker deleting its own lease key release a claim that another worker now holds?

level: middleimportance: must knowfreq 60%

answer

  1. delete by name forgets who wrote it
  2. the deadline can pass mid-job
  3. store a token with the claim
  4. delete only if the token still matches

basics

~20 s

A blind delete removes whatever is under the key, not the claim you made. If your lease reached its deadline mid-job and a second worker claimed that key, your release frees the second worker's claim.

solid answer

~50 s

Deleting by key name has no memory of who wrote the entry. If the first worker's claim reaches its deadline while its job is still running, a second worker wins its own conditional create under the same key; when the first worker finally finishes and deletes, it frees the second worker's claim, and a third worker starts on a job that is already running. The repair is a **holder token**: a fresh, unguessable value written with the claim, plus a release that checks the holder first — delete only if the stored token still matches. Stores differ in whether that comparison and the delete can run as one server-side step; where they cannot, a read followed by a delete leaves a small gap, and the practical answer is a short, renewed claim rather than trust in the delete.

code

pseudocode · 9 lines
pseudocode
token = freshRandomValue()

if createIfAbsent("job:nightly-rollup", token, lifetime = 5 minutes):
    try:
        runTheJob()
    finally:
        deleteIfHolder("job:nightly-rollup", token)
else:
    exit()

go deeper

for a junior

Hold on to one idea: deleting a key removes whatever is under that name right now, which is not necessarily the claim you created. That is why a claim carries something identifying its claimant.

for a middle

Walk the sequence out loud — deadline passes, second worker claims, first worker deletes — and then name the repair as a fresh token plus a release that deletes only while the stored token matches.

for a senior

Say that the comparison and the delete must happen as one step, that stores differ in whether they can, and that a release deleting nothing is a signal your lifetime is too short rather than a harmless outcome.

for a principal

Decide what the system does with the overlap the token cannot prevent. If two concurrent runs are merely wasteful, this is enough; if they are not, the guard belongs where the job's effects are made durable.

## The delete is broader than the claim A lease is a key whose presence means *somebody is working on this*, created with a **conditional create** — a write that succeeds only if the key is absent. Releasing it looks symmetrical: you created a key, so you delete a key. It is not symmetrical. The create was conditional on the key's state; the delete, written naively, is conditional on nothing. It removes whatever is under that name at the moment it arrives, and by then that may be somebody else's claim. ## The sequence that produces it The claim is timed, and the timing is what opens the gap. Take a claim written with a five-minute lifetime on a job that usually takes two minutes and occasionally takes eight: | Time | Worker A | Worker B | The key in the store | |---|---|---|---| | 0:00 | wins the conditional create, starts the job | — | A's claim, deadline 5:00 | | 5:00 | still working | — | absent: the deadline passed | | 5:01 | still working | wins its own conditional create, starts the same job | B's claim, deadline 10:01 | | 8:00 | finishes, deletes by key name | still working | absent: A deleted B's claim | | 8:01 | done | still working | — | | 8:02 | — | still working | C's claim: a third worker walks in | Notice what the naive delete did. The first problem — two workers on one job — was caused by the deadline. The second problem, a third worker, was caused by the release, and it is entirely avoidable. ## The holder token The fix is to make the claim identify its claimant. Along with the key, write a **holder token**: a value only that claim attempt knows. - It must be **fresh for every attempt**, not fixed per worker. A worker that loses a claim and later wins a new one would otherwise be unable to tell its own two claims apart, and a stale release would still match. - It must be **unguessable**, so no other worker can reproduce it by accident or by construction. A host name, a process identifier or a job name all fail this: they repeat. - It is stored **with the claim**, not beside it. A token in a second key can be lost, expire separately, or be updated out of step with the claim it describes. ## Release that checks the holder first The release becomes conditional: *delete this key only if the stored token still equals mine*. Read as behaviour rather than as any particular operation, that gives three outcomes: 1. The token matches — the claim is still yours, the delete succeeds, the next worker may start immediately. 2. The token differs — your claim already reached its deadline and somebody else holds the key. The delete does nothing, which is exactly right. 3. The key is absent — your claim reached its deadline and nobody has re-claimed it yet. Nothing to do. The comparison and the delete need to happen **together**. If you read the token, compare it in your application, and then issue a delete, the claim can reach its deadline and be re-taken in between, and you are back to deleting somebody else's claim through a narrower gap. Stores of this class differ here: some can evaluate a condition and apply a change in one server-side step, others can only read and write, and in those the gap cannot be closed from the client side at all. ## What the token does and does not buy A holder token makes the *release* safe: you can no longer free a claim that is not yours. It does nothing about the underlying cause, which is that the deadline passed while you were working. Two workers were already on the job before the delete happened; the token simply stops a third from joining them. The cause is addressed separately, by renewing the claim while the holder is healthy, or by accepting the overlap and putting the real guard where the job's effects land. ## The signal you should not swallow A release that deletes nothing is information. It says: *my claim was gone before I finished, so somebody else may have run this job*. Log it with the job identity, count it, and alert when it stops being rare. In most systems it is the only cheap evidence available that a lifetime is sized too short, and teams routinely discard it because the release "succeeded" from the application's point of view.

  • Why must the token be new for every claim attempt rather than fixed per worker?
    Because the same worker can lose one claim and later win another. With a fixed per-worker value, a release left over from the lost claim still matches the new one and deletes it, which is the very failure the token was added to prevent. A fresh value per attempt makes each claim distinguishable, and it must be unguessable so no other worker can produce the same value.
  • What should a worker do when its release that checks the holder deletes nothing?
    Treat it as a signal rather than an error to swallow. It means the claim was gone before the work finished, so another worker may have run, or may still be running, the same job. Log it with the job identity, count it, and alert if it stops being rare — it is usually the cheapest evidence that the lifetime is undersized.
  • Does the holder token stop two workers running the job at once?
    No. By the time a release is refused, the overlap has already happened: the deadline passed and a second worker claimed the key while the first was still working. The token only prevents a third worker joining because the first freed a claim that was not its own.

saying these in an interview costs you the question

  • Releases the claim with a plain delete by key name
  • Uses the host name or process identifier as the holder token
  • Believes a read, then a compare, then a delete closes the race
  • Assumes a claim can only be gone if its holder released it
  • Thinks a longer lifetime removes the need for a holder token