skip to content

In a background-job platform, a worker stalls past its lease and the job is reassigned; how do you stop the stalled worker committing a stale result?

level: seniorimportance: must knowfreq 55%

answer

  1. a pause can land anywhere
  2. check-then-act race
  3. monotonic number per claim
  4. zero rows changed means you lost
  5. external writes need the key too

basics

~20 s

Give every claim a monotonically increasing attempt number, a fencing token, and make every commit conditional on it still being current. The stalled worker's write matches zero rows, so it aborts; external side effects must check the token too or be idempotent.

solid answer

~50 s

Checking 'do I still hold the lease?' just before writing is not enough, because the pause can land between the check and the write. Instead, each claim increments `attempt` on the job row and the worker carries that number. Completion is one compare-and-set: `UPDATE jobs SET state = 'succeeded' WHERE job_id = ? AND attempt = ?`. If the job was reassigned, `attempt` has moved on, zero rows change, and the stalled worker discards its work. Outputs are written to an attempt-scoped location and promoted by that same conditional write. External side effects, such as a customer email, cannot see the job row, so either pass the token to a system that rejects older tokens or key the effect by job id so a second execution is a no-op. Fencing prevents a double *commit*, not a double *execution*, so job handlers must be idempotent.

code

pseudocode · 9 lines
pseudocode
attempt = claim(job)              // attempt counter incremented atomically
output = render(job)
ref = writeStaging(job.id, attempt, output)
changed = commitIfAttempt(job.id, attempt, ref)
if changed == 0:
    deleteStaging(ref)            // lease moved on; discard this result
    return LOST_LEASE
sendEmail(key = job.id + ":notify")  // same key on every attempt
return DONE

go deeper

for a junior

Know that a paused worker can wake up after its job was reassigned, and that each claim gets a new, larger attempt number.

for a middle

Explain why checking the lease before writing still races, and show the conditional update that compares the attempt number inside the store.

for a senior

Cover every write path: staged outputs promoted atomically, external effects keyed by job id, zero-row heartbeats stopping work, and metrics on rejected stale commits.

for a principal

Be explicit that the platform gives at-least-once execution, and decide per downstream system whether to require token checks, idempotency keys, or reconciliation.

## The zombie worker problem In a background-job platform that hands out time-limited leases, a worker that stops heartbeating loses its job to another worker. That is correct when the worker is dead, but a worker can also **stall** and come back: a long stop-the-world pause, a machine suspended by the hypervisor, or a network partition that heals. The returning worker is a **zombie**: it still believes it owns the job. A typical timeline: 1. Worker A claims a render job as attempt 4 and starts working. 2. A stalls for 90 seconds; its 60-second lease expires. 3. Worker B reclaims the job as attempt 5 and starts over. 4. A wakes up, finishes its render and tries to mark the job succeeded with its output. 5. B finishes later and writes its own output. Without protection, A's result is committed first and B then overwrites it, or both write different files and the job ends up pointing at whichever wrote last. ## Why check-then-write fails The obvious fix is for A to check its lease before committing. That is a **check-then-act race**: the pause can happen after the check succeeds and before the write lands. No amount of checking closes that window, because the stall can happen at any instruction. The protection has to live in the **write itself**, enforced by the store that everyone writes through. ## Attempt numbers as fencing tokens A **fencing token** is a number that strictly increases every time ownership changes. In a job platform the attempt counter already is one: each claim increments it atomically. The worker keeps its number and presents it on every write: ```sql UPDATE jobs SET state = 'succeeded', result_ref = :output_ref, finished_at = CURRENT_TIMESTAMP WHERE job_id = :job_id AND state = 'running' AND attempt = :my_attempt; -- zombie A presents attempt 4, the row says 5: 0 rows changed, A aborts ``` Key properties: - The comparison happens **inside** the store, atomically with the write, so no pause can split them. - The token must be **monotonic**. A worker's wall-clock timestamp is not safe, because clocks skew and can move backwards. - Every state-changing write carries the guard: progress checkpoints, heartbeats, partial results and the final commit. ## Protecting outputs and side effects The job row is only one of the places a job writes to. Each kind of effect needs its own protection: | Effect | Protection | |---|---| | Job state and result pointer | conditional update on `attempt` | | Output files | write to a path that includes the attempt, then point the job at it with the conditional commit; clean up losers later | | Rows in another store the platform controls | store the latest accepted token per job and reject writes carrying a lower one | | External calls such as email or payment | an **idempotency key** derived from job id and step, never from the attempt number | The last row is the subtle one. If the key included the attempt number, attempts 4 and 5 would look like two different requests and the customer would get two emails. Keyed by job id and step, the second send is recognised as a duplicate. ## At-least-once execution, exactly-once effect Fencing guarantees that **at most one attempt's commit is accepted**. It does not stop both attempts from **executing**: during the overlap, A and B genuinely run at the same time. The platform therefore offers **at-least-once execution**, and the observable effect is made exactly-once by combining: 1. fenced commits for state the platform owns, 2. attempt-scoped staging for outputs, promoted atomically, 3. idempotency keys or token checks for anything external. Where a downstream system can do neither, the remaining options are to record the intent under the fenced write before calling out and reconcile afterwards, or to accept a rare duplicate and make it detectable. ## Operational signals - Count **rejected stale commits**. A rising rate means leases are too short for real pause behaviour, or workers are overloaded. - Log the attempt numbers on both sides of a rejection so the overlap can be reconstructed from the run history. - Alert on orphaned attempt-scoped outputs that were never cleaned up; they show wasted compute. - Make the zombie stop quickly: a heartbeat that changes zero rows should cancel the work, not just the commit.

  • Why derive a side effect's idempotency key from the job id rather than the attempt number?
    The point of the key is to make two executions of the same logical work look identical to the receiver. Attempt numbers differ between the zombie and the new worker, so keys built from them would let both sends through. A key from job id plus step name is the same on every attempt, so the receiver performs the effect once and treats the other as a duplicate.
  • What if a downstream system can neither check tokens nor deduplicate requests?
    Record the intended effect in the job store under the fenced write before calling out, so only the current attempt can register it, then have one relay perform registered effects and mark them done. A crash between the call and the mark can still duplicate, so pair it with reconciliation that detects and corrects duplicates. Sometimes the honest answer is that exactly-once is not achievable there.
  • How should a worker react when its final commit changes zero rows?
    Treat it as a lost lease, not an error to retry. Delete its attempt-scoped output, skip any external effects not yet performed, record a stale-commit metric with its attempt number, and return. Retrying the commit or forcing it through would overwrite the current attempt's work.

saying these in an interview costs you the question

  • Checking the lease just before writing removes the race.
  • Once the lease moves, the old worker is guaranteed to have stopped.
  • Fencing tokens ensure the job body can never run twice.
  • A worker's wall-clock timestamp works just as well as a counter token.
  • Idempotency keys for side effects should include the attempt number.