skip to content

For an at-least-once job queue where every job must eventually be processed, why is Redis Pub/Sub the wrong primitive, and what do Redis Streams provide instead?

level: middleimportance: must knowfreq 50%

answer

  1. at-least-once = retention + ownership + acknowledgement
  2. PUBLISH copies to current subscribers, then forgets
  3. lock bolt-on buys exclusivity only, still loses the job
  4. XREADGROUP `>` assigns; PEL holds until XACK
  5. duplicates guaranteed → idempotent handlers

basics

~20 s

Pub/Sub keeps nothing: it broadcasts once and forgets, so a worker that is restarting or crashes mid-job loses that job. A Redis Streams consumer group retains each entry, gives one consumer ownership of it, and keeps it pending and recoverable until XACK.

solid answer

~60 s

At-least-once needs three things Redis Pub/Sub does not have: **retention**, **ownership**, and **acknowledgement**. `PUBLISH` copies the payload to whoever is subscribed at that instant and forgets it. A worker restarting during a deploy misses those jobs permanently and cannot even detect the gap; a worker that crashes halfway through has nothing to recover from, because no copy exists; a slow worker is disconnected once it exceeds `client-output-buffer-limit pubsub` and loses everything buffered. And every subscriber gets a copy, so ten workers means ten executions. The usual bolt-on — each worker races `SET lock:<job-id> 1 NX EX 60` and only the winner processes — buys exclusivity and nothing else. A crash after taking the lock still loses the job, because nothing retained it. Redis Streams give all three natively: `XADD` stores the entry, `XREADGROUP ... STREAMS jobs >` hands it to exactly one consumer in the group and records it in that group's pending entries list (PEL), and `XACK` clears it. Anything unacked stays in the PEL and can be reclaimed by another consumer, which is what makes at-least-once real — and why handlers must be idempotent.

code

text · 8 lines
text
# Pub/Sub + a lock: exclusivity bolted on, still not a queue
PUBLISH jobs "{\"id\":1234,\"type\":\"resize\"}"
# every subscribed worker receives it; each then races:
SET lock:job:1234 worker-3 NX EX 60
# winner processes. If it is killed now, the job exists nowhere:
#  - the message was never retained
#  - the lock expires and simply disappears
#  - no record says "1234 was started and not finished"

go deeper

for a junior

Recall the core contrast: PUBLISH sends to whoever is listening right now and forgets; a stream stores the job and a consumer group makes one worker own it until it acknowledges with XACK. Name retention, ownership, acknowledgement.

for a middle

Be concrete about the failure modes — restart during a deploy, crash mid-job, slow consumer disconnected by client-output-buffer-limit pubsub — and show why a per-job lock fixes only the ownership third. Name XADD, XREADGROUP with >, the PEL and XACK.

for a senior

Price the bolt-on honestly: second store, non-atomic write-then-publish, lock TTL tuned against job duration, a sweeper and attempt counters — i.e. a hand-rolled queue with Pub/Sub demoted to a wake-up ping. State plainly that at-least-once means duplicates and that handlers must be idempotent.

for a principal

Frame it as choosing a primitive by the guarantees it owns rather than by what is already deployed: whether the durability, ordering and backlog bounds of a stream are enough for this workload, what the memory cost of retention and trimming is, and when the answer is neither Pub/Sub nor Streams but a dedicated durable log.

## What "at-least-once" actually demands "At-least-once" means: for every job enqueued, at least one successful execution eventually happens, across worker crashes, restarts and deploys. To satisfy that, the system must be able to answer at any moment, "which jobs were handed out but never confirmed finished?" — and to hand those out again. That requires three properties: 1. **Retention** — a copy of the job outlives its first delivery attempt. 2. **Ownership** — a job is assigned to one worker, not broadcast to all of them. 3. **Acknowledgement** — the worker confirms completion, and the *absence* of confirmation is detectable. ## Redis Pub/Sub has none of them `PUBLISH queue "job:1234"` writes the payload into the output buffer of every currently subscribed client and returns the number of recipients. After that instant, the message does not exist anywhere in Redis. - **No retention.** A worker restarting during a deploy misses every job published in that window, permanently, with no way to notice the gap. - **No ownership.** All subscribed workers receive the same message, so ten replicas means ten executions unless you add exclusion yourself. - **No acknowledgement.** If a worker takes the job and its process is then killed, nothing anywhere records that the job was started and not finished. It is simply gone. - **Silent loss under load.** A subscriber that consumes slower than the publisher accumulates in its output buffer; on crossing `client-output-buffer-limit pubsub` (defaults: 32 MB hard, 8 MB sustained for 60 s) Redis closes the connection. The queue loses jobs precisely when it is busiest. ## What it costs to bolt the missing pieces onto Pub/Sub Teams reach for Pub/Sub because it is already there, then rebuild the queue in application code. Price each piece: - **Ownership** costs a distributed lock per job — `SET lock:<job-id> 1 NX EX <ttl>` — which means every subscriber still receives every message and N−1 of them do a wasted round trip to lose the race. You now own a TTL that must exceed the slowest job (too short and the lock expires mid-run, so two workers process concurrently — the thing the lock was for) and a lock keyspace to size and expire. - **Retention** costs a second store. Before publishing you must durably write the job somewhere — a Redis list, hash, sorted set, or a database row — because the message itself is not a record. That is two writes with no atomicity between them: the store can succeed and the publish fail, or the reverse, and you must decide which failure you prefer. - **Acknowledgement** costs a state machine you maintain: a "started at, by whom" marker, a sweeper that scans for markers older than a threshold, a per-job attempt counter so a repeatedly failing job does not loop forever, and a place to record completion. Add all three and the honest description of the result is: you have written a queue on top of a notification bus, in application code, without atomicity between its parts — and Pub/Sub has been demoted to a wake-up signal that tells workers to go look at the real store. At that point the Pub/Sub component contributes nothing you could not get by having workers block on the store directly. ## What Redis Streams provide natively A stream key is the queue; a **consumer group** is the worker pool. - **Retention**: `XADD` appends the entry to the stream key, where it lives until you trim it, persisted and replicated like any other Redis value. - **Ownership**: `XREADGROUP GROUP workers worker-3 ... STREAMS jobs >` returns entries never yet delivered *to that group*, and Redis records which consumer name received which entry. One entry, one consumer — no lock needed. (Note that `>` means "not yet delivered to this group", not "added since I connected": a worker that was down during a deploy still gets the backlog.) - **Acknowledgement**: the entry sits in the group's **pending entries list (PEL)**, tagged with its owner and an idle timer, until `XACK` removes it. Anything in the PEL is, by definition, started-but-unconfirmed work, and an unacked entry stays recoverable — another consumer can take it over rather than the job being lost. The mechanics of that takeover — idle thresholds, `XPENDING`, `XCLAIM`/`XAUTOCLAIM`, dead-lettering a poison job — are the subject of the consumer-group recovery question in this topic; the point *here* is that the primitive retains the record at all, which is exactly what Pub/Sub does not do and what no amount of bolt-on can cheaply recreate. ## Duplicates are the price of at-least-once Because a worker can finish a job and then die before `XACK`, redelivery is guaranteed to happen eventually. Handlers must be idempotent — key writes by job id, guard with `SET NX`, or make the effect naturally repeatable. Anyone claiming exactly-once from a Redis stream without describing that guard is overselling. ## The one-line answer Pub/Sub is a notification bus with no memory; a job queue needs memory, ownership and confirmation. Streams with a consumer group give those three natively, where Pub/Sub can only get them by re-implementing a queue beside it.

  • A team already runs Redis Pub/Sub in production for jobs. What is the minimum they would have to build to make it at-least-once, and why is that usually a bad trade?
    At minimum: durably store each job before publishing, take a per-job lock so only one subscriber runs it, record a started-not-finished marker, run a sweeper that re-dispatches stale markers, and count attempts so a failing job does not loop forever. That is a queue implemented in application code, with no atomicity between the publish and the store, and it demotes Pub/Sub to a wake-up ping. Redis Streams give the same three properties as primitives, so the bolt-on buys operational risk and custom code for no capability gain.
  • Does switching from Pub/Sub to Streams give you exactly-once processing?
    No — it gives at-least-once. A worker can complete the job and then crash before sending XACK, so the entry stays pending and will be handed to another consumer, which runs it again. Handlers must therefore be idempotent: key side effects by job id, guard with SET NX, or make the effect naturally repeatable. Claiming exactly-once from a Redis stream without describing that guard is a red flag in an interview.
  • A worker was down for 30 seconds during a rolling deploy. What does it see on reconnect under each primitive?
    Under Pub/Sub it sees nothing from that window — messages published while it was disconnected were copied only to the clients connected at that instant and are gone, with no gap detectable. Under a Streams consumer group, XREADGROUP with `>` returns entries never delivered to that group, so the backlog accumulated during the outage is delivered as soon as the worker comes back. The `>` id means undelivered-to-this-group, not newer-than-my-connection.

Pub/Sub is shouting a work order across the floor: if the person you meant was on break, the order was never written down. A Stream is a ticket spike — the ticket stays clipped under a worker's name until they sign it off, and it is still hanging there if that worker never comes back.

saying these in an interview costs you the question

  • Saying Pub/Sub replays or buffers messages for a subscriber that reconnects — it does not; a disconnected subscriber's messages simply never existed for it.
  • Claiming that adding a distributed lock (SET NX EX) turns Pub/Sub into an at-least-once queue — it adds exclusivity only; a crash after locking still loses the job because nothing was retained.
  • Claiming Streams give exactly-once delivery, without mentioning that a crash between finishing work and XACK forces a redelivery and requires idempotent handlers.
  • Believing XREADGROUP with `>` returns only entries added after the client connected, and therefore concluding Streams also drop the deploy-window backlog.
  • Treating XACK as optional bookkeeping — without it the entry stays pending forever and the pending list grows unbounded.

context