A job outruns its lease's deadline while still working — what does renewing the lease buy, and what can renewal not prevent?
answer
- the lifetime now covers the pause, not the job
- extend only while still the holder
- several renewals inside one lifetime
- a failed renewal means stop, not retry
basics
~20 sRenewal replaces one guess with a smaller one: the lifetime need only exceed the worst pause between renewals, not the worst run. It cannot stop a stalled holder losing a claim it still believes it holds.
solid answer
~50 sWithout renewal the lifetime has to cover the slowest run you will ever have, which means a crashed holder stalls the job for that long. With renewal the holder extends its own claim while it is healthy, so the lifetime only has to cover the longest gap between two successful renewals — seconds rather than minutes. What renewal cannot do is make the claim follow the worker's reality: if the holder is paused, descheduled, or cut off for longer than the remaining lifetime, the claim ends on the store's clock, another worker claims the key, and the first worker is told nothing and carries on. So renewal must itself be conditional on still being the holder, the interval must leave room for several failures inside one lifetime, and a failed renewal is an instruction to stop acting as the holder — not a transient error to retry through.
go deeper
The idea to keep is that a claim ends on time whether or not the work is finished. Renewal is the holder saying 'still here' before the deadline arrives.
Explain the arithmetic: the lifetime must cover the gap between renewals, and the interval should be a fraction of it so a couple of failures are survivable.
Walk the paused-holder timeline and state plainly that neither worker learns of the other, then give the rule that a failed renewal means stop acting as the holder.
Treat the lifetime as a tuning knob between stall time after a crash and overlap after a pause, and decide which one the workload can absorb before choosing a number.
## What renewal changes A claim with a fixed lifetime forces one uncomfortable choice: the lifetime must exceed the slowest run the job will ever have, and that same number is how long the job is stalled after a holder is killed. Renewal breaks the coupling. The holder extends its own claim periodically while it is healthy, so: - the lifetime now only has to cover **the longest pause between two successful renewals**, which is a property of the runtime and the network rather than of the job; - a crashed holder is noticed within roughly one lifetime, so the stall shrinks from minutes to seconds; - an unusually slow run no longer loses its claim just for being slow. Renewal is not free — it is a write per interval per live job — but the volume is trivial next to what it buys. ## Renewal must be conditional too An unconditional extension of the key has the same defect as an unconditional delete: it acts on whatever is under the name. If the deadline already passed and a second worker claimed the key, a blind renewal by the first worker **extends the second worker's claim** — and now the wrong worker's deadline is being pushed forward by a process that no longer holds anything. The renewal must check the holder first: extend only while the stored token still matches. When it does not match, the answer is not to retry harder; it is that the claim is gone. ## The pause renewal cannot cover Here is the sequence that a senior interview is really asking about. The claim carries a thirty-second lifetime and the holder renews every ten seconds: | Time | Worker A | The key in the store | Worker B | |---|---|---|---| | 0:00 | claims, starts work | A's claim, deadline 0:30 | — | | 0:10 | renews | A's claim, deadline 0:40 | — | | 0:20 | host pauses — no renewal leaves the process | A's claim, deadline 0:40 | — | | 0:40 | still paused | absent: the deadline passed | — | | 0:41 | still paused | B's claim, deadline 1:11 | claims, starts the same work | | 1:05 | resumes, believing it holds the claim | B's claim | still working | Nothing here is broken. Each component did exactly what it promised. The store enforced a deadline it was asked to enforce; the second worker won a create against an absent key; the first worker was never contacted, because the store is not in a position to interrupt a caller that is not asking it anything. The result is two workers doing one job, and **neither of them knows it**. The pause can be anything that stops a renewal landing for longer than the remaining lifetime: a long garbage-collection pause, a descheduled container, a host suspended for migration, a network path that drops for twenty seconds, or the store itself being slow to answer. The deadline is measured where the entry lives, not where the worker lives, and beliefs on the worker's side are always an estimate of a decision taken elsewhere. ## Sizing the interval 1. **Fit several renewals inside one lifetime.** An interval of about a third of the lifetime or less means two consecutive failures still leave time to recover. An interval equal to the lifetime means one slow round trip loses the claim. 2. **Set the lifetime from the worst tolerable pause, not from the job.** Ask how long the runtime can stop the process, or the network can be gone, before you would rather somebody else took over. 3. **Treat a renewal attempt like any call to the tier**: give it a timeout shorter than the interval, so a hanging renewal does not silently become no renewal. ## The rule the holder must obey The moment a renewal fails — refused because the token no longer matches, or simply not answered — the worker must stop acting as the holder: - stop starting new outside effects; - checkpoint or abandon the work rather than push it to completion; - do **not** release with a blind delete, because the key may now hold somebody else's claim; - record the event, because it is the measurement that tells you your lifetime is mis-sized. The failure mode teams actually ship is the opposite: a renewal loop that logs a warning and keeps working, on the assumption that the failure was transient. That converts a recoverable overlap into an overlap nobody can see. ## What still varies Stores of this class differ in what they let you express: some can evaluate a condition and extend a lifetime in one server-side step, others cannot, and in those the renewal is a read and a write with a gap between them. Where the condition cannot be expressed at all, renewal is best-effort and the design should lean harder on the downstream guard.
- How should the renewal interval relate to the lifetime?Several renewals should fit inside one lifetime — an interval of about a third or less is a common shape, so two consecutive failures still leave room to recover. An interval equal to the lifetime means a single slow round trip costs the claim. Give each renewal a timeout shorter than the interval so a hung call does not silently become a missed renewal.
- What must a worker do the moment a renewal of its claim fails?Stop treating itself as the holder: issue no new outside effects, checkpoint or abandon the work, and do not release with a blind delete, since the key may now hold another worker's claim. Record the event so the mis-sizing is visible. The one thing it must not do is keep going on the assumption the failure was transient.
- Why is renewing without checking the holder worse than not renewing at all?Because a blind extension acts on whatever is under the key. If the deadline already passed and a second worker claimed it, the first worker is now pushing the second worker's deadline forward while holding nothing — so the overlap lasts longer and the real holder's timing is being driven by a process with no stake in it.
A meeting room booked until three o'clock. At one minute past, the next group walks in — the booking does not stretch because you are still talking, and nobody in the room is consulted. If you want the room for longer you must extend the booking before three, and you cannot do that while you are stuck in the corridor with no signal.
saying these in an interview costs you the question
- Believes a renewal loop makes the claim safe against any pause
- Renews the key without checking the claim is still theirs
- Sets the renewal interval equal to the lifetime
- Keeps working normally after a renewal has failed
- Thinks the worker's own clock decides when its claim ends