In a job scheduler that hands out time-limited worker leases, how does it detect a worker that died mid-job and reclaim that job?
answer
- time-bound ownership, not a flag
- renewal pushes the expiry forward
- silence is the failure signal
- expired lease means claimable again
- compare against one clock only
basics
~20 sEach claim records a lease owner and expiry, and the worker extends the expiry with periodic heartbeats. When heartbeats stop, the lease lapses, and a reaper or the claim query itself treats the job as claimable and starts a new attempt.
solid answer
~50 sWhen a worker claims a job it writes `lease_owner` and `lease_expires_at = now + T` in one conditional update. While running, it heartbeats every `T/3` or so, pushing the expiry forward. A worker that crashes, is killed or loses its network simply stops renewing, so within `T` of its last renewal the lease lapses. Recovery is then just a query: a periodic reaper, or the normal claim query that accepts `queued` jobs *or* `running` jobs with an expired lease, picks the orphan up, increments the attempt number and hands it to a live worker. Expiry should be compared against one authoritative clock, the job store's, not each worker's. Lease length is a trade-off: short leases recover orphans quickly but misfire on long pauses or slow networks, long leases strand dead jobs. Long jobs keep a short lease and renew it rather than taking one lease as long as the job.
code
sql · 8 linesUPDATE jobs
SET state = 'running',
lease_owner = :worker_id,
lease_expires_at = CURRENT_TIMESTAMP + INTERVAL '60' SECOND,
attempt = attempt + 1
WHERE job_id = :job_id
AND (state = 'queued'
OR (state = 'running' AND lease_expires_at < CURRENT_TIMESTAMP));go deeper
Remember the three pieces: a lease with an expiry, heartbeats that renew it, and a reclaim when the expiry passes without renewal.
Walk through the atomic claim, the heartbeat as a conditional renewal, and how the reaper or claim query turns an expired lease into a new attempt.
Show how you pick lease length and heartbeat interval from observed pause and network behaviour, compute worst-case detection delay, and handle a mass worker failure without a reclaim storm.
Frame the lease length as a business trade-off between recovery time and duplicate executions, and decide which job types justify the extra write load of frequent heartbeats.
## Why a lease instead of a 'running' flag In a background-job platform, a worker that dies rarely announces it. A process can be killed by the operating system for using too much memory, the machine can lose power, or a network partition can cut it off while it keeps running. If a job were simply marked `running`, nothing would ever move it again: it would be an **orphan**, stuck forever. A **lease** fixes this by making ownership **time-bound**. The worker owns the job only until `lease_expires_at`. To keep owning it, the worker must keep proving it is alive by **heartbeating**, which means renewing the lease. Silence is the failure signal, and silence covers every failure type at once: crash, hang, kill and partition. ## Claiming and renewing Claiming is one atomic conditional update, so two workers can never both win the same job: ```sql UPDATE jobs SET state = 'running', lease_owner = :worker_id, lease_expires_at = CURRENT_TIMESTAMP + INTERVAL '60' SECOND, attempt = attempt + 1 WHERE job_id = :job_id AND (state = 'queued' OR (state = 'running' AND lease_expires_at < CURRENT_TIMESTAMP)); -- 1 row changed: claim won. 0 rows: someone else holds a live lease. UPDATE jobs SET lease_expires_at = CURRENT_TIMESTAMP + INTERVAL '60' SECOND WHERE job_id = :job_id AND attempt = :my_attempt AND state = 'running'; -- heartbeat; 0 rows changed means the lease was lost ``` Note that `CURRENT_TIMESTAMP` is evaluated by the job store, so every expiry is written and compared on **one clock**. Workers never compare their own wall clock with the stored expiry. ## Reclaiming orphans There are three common ways an orphaned job becomes runnable again: 1. **Reaper scan.** A periodic task finds `running` jobs whose lease has expired, records a `lease_expired` outcome in the run history, and moves them to `queued` or `retry_wait`. 2. **Claim-time takeover.** The claim query itself accepts expired `running` jobs, as in the example above. No separate reaper is needed, although a reaper is still useful for recording outcomes and metrics. 3. **Graceful release.** A worker that is shutting down on purpose clears its lease immediately, so the job does not wait out the full lease. Every reclaim increments `attempt`. That number is what later lets the platform reject writes from the previous holder. ## Choosing lease length and heartbeat interval | Choice | Benefit | Cost | |---|---|---| | Short lease (tens of seconds) | orphans recovered quickly | a long pause or slow network causes a false reclaim and a duplicate run | | Long lease (hours) | almost no false reclaims | a crashed worker's job sits idle for hours | | Heartbeat about every T/3 | a worker can miss a renewal or two and keep the job | more writes to the job store per running job | With illustrative numbers, a 60-second lease and a reaper scanning every 15 seconds, the worst-case detection delay is about 75 seconds: up to 60 for the lease to lapse after the last renewal, plus up to 15 before the next scan. For jobs that run for hours, the right answer is a **short lease with renewal**, not a lease as long as the job. ## Heartbeats in practice - Run the heartbeat on its own timer, separate from the work, so a busy job still renews; but make it stop if the work is truly stuck, or a hung job will hold its lease forever. A progress counter included in the heartbeat lets the platform spot that case. - Treat a heartbeat that changes **zero rows** as a lost lease: stop the work and do not commit. - Batch heartbeats from one worker that runs many jobs into a single write to reduce load. - When a whole pool dies at once, reclaim in batches so survivors are not flooded. ## What a lease does not guarantee - An expired lease means the platform **treats** the job as orphaned. It does not prove the old worker stopped: a paused worker can wake up and try to finish. Guarding commits with the attempt number, known as fencing, handles that. - Because a false reclaim can happen, the same job can execute twice, so job side effects must be idempotent. - A lease detects silence, not wrong answers. A worker that heartbeats while producing garbage needs separate validation.
- What should a worker do when its heartbeat renewal changes zero rows?Treat it as having lost the lease: the job was reclaimed, cancelled or finished by someone else. Stop the work at the next safe point, do not attempt the final commit, clean up any attempt-scoped temporary output, and record a metric. Retrying the renewal or pressing on hoping to finish first just creates a second, conflicting run.
- Why not rely only on the process supervisor to report dead workers?A supervisor sees a process exit, but not a hung process, a worker cut off by a network partition, or a machine that vanished together with its supervisor. Lease expiry covers all of these with one mechanism. Supervisor signals are still useful to release a lease early and shorten recovery, but they cannot replace expiry.
- How do you keep a large burst of reclaims from overwhelming the surviving workers?Workers claim only as many jobs as they have free slots, so reclaimed jobs queue rather than pile onto survivors. The reaper moves expired jobs in bounded batches, and reclaimed jobs can get a small randomised `not_before` so they do not all become due in the same instant.
saying these in an interview costs you the question
- A crashing worker will always mark its job failed before exiting.
- Set the lease to the job's maximum runtime so it never expires early.
- Compare the stored lease expiry against each worker's local clock.
- An expired lease proves the previous worker has stopped running.
- One heartbeat per full lease length leaves enough safety margin.