skip to content

A worker consuming a Redis Stream with XREADGROUP crashed while holding entries it had not acknowledged. What happens to those entries, and how do you get them processed by a surviving worker?

level: seniorimportance: must knowfreq 45%

answer

  1. no visibility timeout — nothing redelivers by itself
  2. XPENDING IDLE finds them, XCLAIM/XAUTOCLAIM take them
  3. min-idle-time = atomic race guard; never 0 in a loop
  4. XAUTOCLAIM 7.0 returns cursor + entries + deleted IDs
  5. delivery counter climbing = poison → manual dead-letter stream

basics

~20 s

Nothing happens automatically: they stay in the group's pending list owned by the dead consumer's name forever. Another consumer must take ownership — XAUTOCLAIM key group new-consumer min-idle-time 0 (or XCLAIM on specific IDs found via XPENDING IDLE). The min-idle-time guard prevents stealing from a merely slow worker.

solid answer

~60 s

Redis has no delivery timeout. Unacknowledged entries stay in the group's PEL, still owned by the crashed consumer's name, indefinitely — recovery is something your workers must actively perform. The modern way is `XAUTOCLAIM mystream mygroup worker-2 60000 0 COUNT 50`, run periodically by every live worker. It scans the PEL from the given cursor, transfers ownership of every entry idle for at least 60s to `worker-2`, and returns a cursor to continue from plus the claimed entries themselves, ready to process. Since 7.0 it also returns IDs whose underlying stream entries no longer exist and drops them from the PEL. The older, more surgical route is `XPENDING mystream mygroup IDLE 60000 - + 10` to find candidates, then `XCLAIM mystream mygroup worker-2 60000 <id...>` for chosen IDs. The `min-idle-time` argument is the safety mechanism: it is checked atomically, so two workers racing to claim the same entry cannot both win, and a slow-but-alive worker is not robbed while it is mid-job. Claiming resets idle time and bumps the delivery counter, which is how you detect poison entries.

code

text · 14 lines
text
# 1. what has been stuck for over a minute?
XPENDING orders billing IDLE 60000 - + 10

# 2. sweep and take ownership, processing what comes back
XAUTOCLAIM orders billing worker-2 60000 0 COUNT 50
# 1) "0-0"                 <- next cursor (0-0 = scan complete)
# 2) 1) 1) "1712-0" 2) 1) "orderId" 2) "A17"
# 3) 1) "1699-0"           <- PEL records whose entries no longer exist (7.0+)

# 3. surgical alternative for known IDs
XCLAIM orders billing worker-2 60000 1712345678901-0

# 4. once its work is claimed away, retire the dead consumer
XGROUP DELCONSUMER orders billing worker-7

go deeper

for a junior

Know that nothing is automatic: a crashed consumer's entries stay pending until another consumer explicitly claims them.

for a middle

Name the commands and their roles — XPENDING IDLE to find, XCLAIM for specific IDs, XAUTOCLAIM to sweep — and explain what min-idle-time is for.

for a senior

Describe the full worker loop (own pending, sweep, then new work), how to choose min-idle-time against p99 job duration, and the delivery-counter-driven dead-letter policy.

for a principal

Reason about the recovery model as a whole: recovery latency vs double-processing risk, why no leader is needed, dangling PEL records from trimming, consumer-name lifecycle in an autoscaled fleet, and where idempotency has to sit.

## There is no automatic redelivery The single most important fact: Redis Streams do not have a visibility timeout. If a consumer reads entries and dies, those entries sit in the group's Pending Entries List forever, tagged with the dead consumer's name, and no other consumer will ever be offered them by `XREADGROUP ... >` — `>` only serves entries never delivered to the group at all. Recovery is a job your application must run. This differs from queue systems that lease messages and re-queue them after a timeout, and candidates who assume a lease model design pipelines that quietly leak work. ## Finding the abandoned entries The diagnostic query is the extended XPENDING with an idle filter: ``` XPENDING mystream mygroup IDLE 60000 - + 10 ``` Each row is `entry-id, consumer, idle-ms, delivery-count`. Idle time is measured from the last delivery or claim of that entry, so anything idle far beyond your normal processing time is either owned by a dead process or by a wedged one. `XINFO CONSUMERS mystream mygroup` gives the complementary view: a consumer whose own idle time is large and whose pending count is non-zero is almost certainly gone. ## XCLAIM: surgical transfer of ownership ``` XCLAIM mystream mygroup worker-2 60000 1712345678901-0 ``` The arguments are key, group, the *new* owner, `min-idle-time` in milliseconds, and one or more entry IDs. Redis transfers each entry only if its current idle time is at least `min-idle-time`; otherwise it is skipped and simply not returned. Because the check and the transfer happen inside one command on the server, two workers issuing the same claim concurrently cannot both succeed — the first resets idle time to zero, so the second's guard fails. That is the whole race-safety story, and it is why you never claim with `min-idle-time 0` in a competitive loop. By default XCLAIM returns the entries with their fields, so the claiming worker can process them immediately. Useful options: - `JUSTID` returns only IDs and does **not** increment the delivery counter (handy for bookkeeping sweeps that are not going to process the data). - `IDLE ms` / `TIME unix-ms` set the new idle value explicitly, e.g. to schedule a retry backoff. - `RETRYCOUNT n` overrides the delivery counter. - `FORCE` creates a PEL entry even if none exists, provided the entry is still in the stream — a repair tool, not a normal path. ## XAUTOCLAIM: the sweep you actually schedule XCLAIM requires you to know the IDs. XAUTOCLAIM (Redis 6.2+) folds the scan and the transfer into one cursor-based command: ``` XAUTOCLAIM mystream mygroup worker-2 60000 0 COUNT 50 ``` It starts at the given ID cursor (`0` for the beginning of the PEL), claims up to COUNT entries idle for at least 60s, and returns three things in Redis 7.0+: the next cursor to resume from (`0-0` when the scan wrapped), the claimed entries with their fields, and a list of IDs that were deleted from the PEL because the underlying stream entry no longer exists. That third element matters: trimming with MAXLEN can evict entries that are still pending, leaving dangling PEL records, and XAUTOCLAIM is the mechanism that reaps them. The normal deployment pattern is that every worker runs the same loop: drain its own pending history with `XREADGROUP ... STREAMS key 0`, then run one XAUTOCLAIM pass, then read new work with `>`. No leader election is needed — the min-idle-time guard makes concurrent sweeps safe, and duplicated effort is bounded. ## Choosing min-idle-time Set it well above the p99 processing time for one entry, including retries and any external call timeouts. Too low and you steal entries from healthy-but-slow workers, producing genuine concurrent double-processing. Too high and failed work sits idle for that long before anyone picks it up. A few multiples of the worst realistic job duration, or of your job timeout, is the usual framing. ## Poison entries and the delivery counter Every claim (and every redelivery) bumps the entry's delivery counter, visible in XPENDING. An entry with a counter climbing into double digits is not a transient failure — it is a message that crashes whatever touches it. Redis has no built-in dead-letter queue, so the pattern is explicit: when the counter exceeds a threshold, XADD the payload plus the failure reason to a separate dead-letter stream, then XACK the original so it leaves the PEL, and alert. Without that, a poison entry is claimed forever and can wedge a whole pipeline. ## Cleaning up dead consumers After its pending entries have been claimed away, the dead consumer name still lingers in the group. `XGROUP DELCONSUMER mystream mygroup worker-7` removes it and returns how many pending entries it still held — a non-zero return means you just discarded work, so delete only after claiming. Periodic cleanup keeps `XINFO CONSUMERS` readable in environments with churning hostnames or pod names, which is also the argument for stable, slot-based consumer names rather than random ones.

  • Why does XCLAIM take a min-idle-time argument instead of just transferring ownership unconditionally?
    It is the concurrency guard. The idle check and the ownership transfer happen atomically inside one command, so if two recovery sweeps target the same entry, the first resets its idle time to zero and the second's guard fails — only one claimant wins. It also protects a slow-but-alive worker from having its in-flight entry stolen mid-job, which would cause real double processing.
  • An entry's delivery count in XPENDING keeps climbing. What is happening and what do you do?
    It is a poison entry: every consumer that claims it fails and never acks, so it is claimed again and again, burning capacity. Redis has no automatic dead-letter mechanism, so implement one — above a retry threshold, XADD the payload and the error to a dead-letter stream, XACK the original to clear the PEL, and alert. Handle it explicitly or it will be redelivered forever.
  • Trimming removed a stream entry that was still pending. What is the state, and how is it cleaned up?
    The PEL keeps a record pointing at an entry that no longer exists, so it can never be delivered again but still counts toward pending. On Redis 7.0+, XAUTOCLAIM removes such records and returns their IDs in its third reply element. On older versions you find them via XPENDING and XACK them explicitly. Sizing trimming bounds well above the pending range avoids the situation.

Orders on a dead colleague's clipboard do not walk back to the wall. Someone has to check the clipboard, confirm nobody has touched it for a while, and move the orders onto their own.

saying these in an interview costs you the question

  • Expecting Redis to redeliver unacknowledged entries automatically after some timeout
  • Running XCLAIM or XAUTOCLAIM with min-idle-time 0, which lets workers steal in-flight entries from each other
  • Assuming XREADGROUP with `>` will eventually re-serve a dead consumer's pending entries
  • Calling XGROUP DELCONSUMER on a dead worker before its pending entries have been claimed, silently dropping work
  • Believing Redis Streams provide a built-in dead-letter queue for repeatedly failing entries

context