A worker pops a job from a Redis list with LPOP and crashes while processing it. Describe how the LMOVE and BLMOVE commands let you build a queue that survives this, and what that design still leaves you to handle yourself.
answer
- LPOP = at-most-once, job vanishes on crash
- BLMOVE jobs -> processing:<worker>, atomic
- ack = LREM after success
- no visibility timeout: build heartbeat + reaper
- at-least-once => idempotent handlers, attempt counter
basics
~20 sWith LPOP the job is gone the moment it is delivered, so a crash loses it. Instead use BLMOVE queue processing:<worker> LEFT RIGHT: it atomically pops and appends to a per-worker processing list, so the job stays visible. Remove it with LREM after success. You must still write the reaper for orphaned entries and make handlers idempotent.
solid answer
~50 sLPOP is at-most-once: the element leaves the list as it is written to the socket, so a crash between delivery and completion loses the job silently. LMOVE source destination LEFT|RIGHT LEFT|RIGHT pops and pushes in one atomic step, and BLMOVE is its blocking form, replacing the deprecated RPOPLPUSH/BRPOPLPUSH since Redis 6.2. The pattern: each worker runs BLMOVE jobs processing:w7 LEFT RIGHT 5. The job now sits in that worker's own processing list. On success the worker runs LREM processing:w7 1 <job> to acknowledge. If it crashes, the job is still in processing:w7 and can be recovered. What Redis does not give you: no visibility timeout and no automatic redelivery. You must run a reaper that detects dead workers, typically via a heartbeat key, and moves stranded entries back with LMOVE. Delivery becomes at-least-once, so handlers must be idempotent, and a poison job retries forever unless you count attempts.
code
text · 9 lines# consume
BLMOVE jobs processing:w7 LEFT RIGHT 5
"job:a91:resize"
# acknowledge after success
LREM processing:w7 1 "job:a91:resize"
# reaper, once worker w7's heartbeat key is gone
LMOVE processing:w7 jobs RIGHT LEFTgo deeper
Know that LPOP loses the job on a crash and that LMOVE keeps it in a second list until the work is acknowledged.
Describe the full loop — BLMOVE, process, LREM — and explain why the move is atomic and the acknowledgement is a separate step.
Own the recovery design: heartbeats, reaper, attempt counters, dead-letter list, LREM cost, and the at-least-once consequences.
Decide whether to build this at all versus using a structure with native acknowledgement, and state what asynchronous replication does to the guarantee across a failover.
## The failure the pattern fixes A naive Redis queue is RPUSH to produce and BLPOP to consume. The flaw is that popping is destructive and unacknowledged: once the reply leaves the server the element exists nowhere. If the worker dies, the network drops the reply, or the handler throws after receiving it, the job is gone with no trace. That is at-most-once delivery, fine for disposable work like cache warm-ups and unacceptable for anything a user paid for. ## The reliable-queue construction LMOVE source destination wherefrom whereto atomically removes an element from one end of the source list and pushes it onto one end of the destination. Atomically matters: there is no window in which the element exists in neither list nor in both. BLMOVE adds the blocking wait, so the consumer loop is a single command. The standard loop is: 1. job = BLMOVE jobs processing:w7 LEFT RIGHT 5 — take from the head of the shared queue, append to this worker's private in-flight list. 2. Process the job. 3. LREM processing:w7 1 job — acknowledge by removing exactly one matching entry. Use a per-worker processing list rather than one shared one. With a shared list you cannot tell which entries belonged to a worker that died, so recovery has to guess; with per-worker lists the recovery rule is simply 'everything in the list of a worker known to be dead'. ## What the crash looks like now If the worker dies after step 1, the job remains in processing:w7 forever and nothing in Redis notices. The recovery machinery is entirely yours to build, and describing it is what separates a senior answer from a textbook one: - Liveness: each worker maintains a heartbeat such as SET worker:w7:alive 1 EX 30, refreshed periodically. A reaper looks for processing:* lists whose worker heartbeat is missing. - Requeue: for each stranded entry, LMOVE processing:w7 jobs RIGHT LEFT to push it back for another attempt, or move it to a dead-letter list after too many tries. - Attempt counting: bare list entries carry no metadata, so the payload itself should include an id and an attempt counter, or a parallel hash keyed by job id should track attempts. Without this a poison job loops forever, burning a worker each pass. - Idempotency: because a job can be redelivered after a crash that happened late in processing, handlers must tolerate running twice. Deduplicate on a job id, or make effects naturally idempotent. ## Costs and caveats LREM is O(N) in the length of the processing list, which is fine when that list holds one or a few in-flight jobs per worker and terrible if it accumulates thousands of stranded entries. Monitor LLEN on processing lists; growth means the reaper is not keeping up. Acknowledging by value also makes duplicate identical payloads ambiguous — LREM removes some matching entry, not necessarily yours. Include a unique id in every payload to keep entries distinct. The move and the acknowledgement are separate round trips, so the guarantee is at-least-once, never exactly-once. And because Redis replication is asynchronous, an element moved on a primary that fails over before the replica caught up can be redelivered or lost — the pattern makes worker crashes survivable, not node failovers. ## When to stop If you find yourself building heartbeats, attempt counters, dead-letter lists and a reaper, you are reimplementing consumer groups. Redis Streams provide pending-entry tracking and acknowledgement natively and are usually the better structure for a durable work queue. The list pattern earns its place when the workload is simple, volume is modest, and you want no extra concepts.
- Why give each worker its own processing list instead of one shared one?Because recovery needs to know which in-flight entries belonged to the process that died. In a shared list, stranded entries are indistinguishable from entries healthy workers are still processing, so a reaper would have to guess and risk duplicate execution. Per-worker lists make the recovery rule trivial.
- Does this pattern give exactly-once processing?No. The move and the acknowledgement are separate operations, so a crash after processing but before LREM causes redelivery. The guarantee is at-least-once, which is why handlers must be idempotent, typically by deduplicating on a job id.
saying these in an interview costs you the question
- Claiming BLMOVE alone makes the queue reliable with no reaper
- Assuming Redis redelivers stranded entries automatically after a timeout
- Using one shared processing list and expecting clean recovery
- Calling the result exactly-once delivery
- Ignoring that LREM is O(N) and that processing lists must stay short