skip to content

A worker holding claimed queue messages and an exclusive lease is asked to stop - what must it release before exiting?

level: seniorimportance: nice to knowfreq 32%

answer

  1. held on behalf of others
  2. in-flight is not only requests
  3. exiting does not give things back
  4. a stall the length of a timeout
  5. return the claim, release the lease

basics

~20 s

Everything it holds on behalf of others: claimed messages returned so another consumer can take them, and the exclusive lease given up so a replacement can acquire it. A worker that just exits leaves both held until their timeouts expire.

solid answer

~50 s

In-flight work is not only requests in handlers. A worker that has claimed messages from a queue has made them invisible to every other consumer until the claim expires, and a worker holding an exclusive lease is blocking whoever would take over. If it exits without releasing them, nothing is corrupt but everything is stalled: the replacement starts, finds the messages still claimed and the lease still held, and does nothing useful until those timeouts run out. So the shutdown sequence for a worker adds two steps - stop claiming new messages first, then finish or explicitly return what is already claimed, and release the lease before exiting. The timeouts remain the safety net for the case where the forced kill arrives first, which is why they should be short enough to bound the stall.

code

pseudocode · 13 lines
pseudocode
on stopRequest:
    claimingEnabled = false            # 1. stop taking on new messages

    deadline = now() + gracePeriod - safetyMargin
    for each message in claimedMessages:
        if now() < deadline:
            process(message)
            acknowledge(message)       # 2. finished: it is gone for good
        else:
            returnToQueue(message)     # 3. not finished: visible again now

    releaseLease(role)                 # 4. replacement can take over at once
    exit()

go deeper

for a junior

Know that a message taken from a queue is usually hidden from other consumers rather than deleted, so a worker that vanishes takes that work out of circulation until a timeout brings it back.

for a middle

Explain the two kinds of in-flight state - work inside the process and things held on its behalf outside it - and why only the second kind needs an explicit release step during shutdown.

for a senior

Recognise the symptom in the wild: consumption stalls for exactly the length of some timeout after every deploy, and the fix is returning claims and releasing leases rather than tuning the timeout down.

for a principal

Frame the timeouts as a bound on the ungraceful case and the explicit release as the design for the graceful one, and decide how that contract is made consistent across every worker on the estate.

## In-flight work is not only requests Graceful shutdown is usually taught with a request-serving example, where "in flight" means a request sitting in a handler. A worker consuming from a queue has a second, less visible kind of in-flight state: **things it holds that other instances are therefore denied**. - **A claimed message.** When a consumer takes a message from a work queue, the queue typically hides it from other consumers for a claim period rather than deleting it. Until the consumer acknowledges it, or the claim expires, nobody else can process it. - **An exclusive lease.** A worker that owns a shard, a partition, a scheduled job or a singleton role usually holds it by renewing a lease with a term. Nobody else may take that role while the lease is live. - **A reserved slot** in some external limiter or pool, held for the duration of the work. All three share a property: they are **held on behalf of the system, not the process**, and the process exiting does not automatically give them back. ## What happens if the worker just exits Nothing is corrupted. The queue is doing exactly what it was designed to do, and the lease is doing exactly what it was designed to do. What you get instead is a **stall whose length is set by a timeout, not by your shutdown**: 1. The worker is asked to stop and exits promptly, feeling well behaved. 2. Its replacement starts and connects to the queue. 3. The messages the old worker had claimed are still invisible. The replacement sees an emptier queue than reality and idles. 4. The lease is still live for the rest of its term, so the replacement cannot take over the role at all. 5. Minutes later the claim period and the lease term expire, the messages reappear, the lease is acquirable, and work resumes. On a scale-down of one replica this is an odd latency blip. On a rolling replacement of every worker at once it is a service that appears to stop consuming for the length of a timeout nobody remembers configuring. ## The sequence a claim-holding worker needs | Step | Why it is in that position | |---|---| | Stop claiming new messages | Otherwise the set you must finish keeps growing | | Finish the messages already claimed, if the budget allows | Completing is always better than returning | | Explicitly return the ones you cannot finish | Makes them visible again immediately, not at timeout | | Release the lease | Lets the replacement take the role now rather than at term end | | Exit | Do not wait out the remaining window | The ordering mirrors request draining exactly: stop taking on, then finish, then give back, then leave. The difference is the explicit **return** step, which has no equivalent in request serving - a request you cannot finish is simply cut, but a message you cannot finish can be handed back. ## The timeouts are the safety net, not the plan The claim period and the lease term exist so that a worker which crashes, is force-killed, or loses the network does not strand work forever. That is a correctness guarantee and you should keep it. But treating it as the shutdown mechanism means every planned shutdown pays the full timeout, which is the difference between a two-second handover and a two-minute one. The practical consequence is a sizing rule of its own: a lease term should be **long enough that ordinary renewal delays do not cause a spurious handover, and short enough that an ungraceful loss is bounded to something you can tolerate**. Same for the claim period. Neither number should ever be reached during a normal shutdown, because a normal shutdown returns things explicitly. ## Returning work means it will be processed again A message you hand back is redelivered, and a message that was force-killed mid-processing is also redelivered once its claim expires. Either way, the handler must be safe to run twice on the same message - partially applied effects from the interrupted attempt are the thing that makes redelivery dangerous rather than merely wasteful. That is the same requirement graceful shutdown puts on multi-step request work, arriving through a different door. ## Why this is asked Because it separates candidates who have read about graceful shutdown from candidates who have operated a worker fleet. Request draining is in every article; the queue claim and the lease are what actually made the incident, and the symptom - consumption stops for exactly the length of some timeout after a deploy - is one you recognise instantly or not at all.

  • What happens if the worker is force-killed before it can release anything?
    The claim period and the lease term are the safety net: the messages become visible again and the lease becomes acquirable once those expire, so nothing is lost. The cost is that the handover stalls for the length of those timeouts, which is precisely what an explicit release avoids on a planned shutdown.
  • How should those timeouts be sized, given that shutdown should never reach them?
    Long enough that ordinary renewal or processing delays do not trigger a spurious handover while the worker is healthy, and short enough that an ungraceful loss - a crash or a forced kill - strands the work for an interval you can tolerate. They are a bound on the bad case, not a parameter of the good one.
  • Why does returning a claimed message put a requirement on the handler?
    Because the message will be delivered again, to this worker's replacement. If the interrupted attempt already applied part of its effect, the second attempt must not apply it twice. That makes safe re-processing a precondition for returning work rather than an optional refinement.

saying these in an interview costs you the question

  • Treats in-flight work as only the requests in handlers
  • Assumes exiting the process releases claims and leases
  • Relies on claim timeouts as the normal handover mechanism
  • Keeps claiming new messages while shutting down
  • Returns claimed work without checking it is safe to reprocess