skip to content

A team proposes replacing a durable partitioned message log with Redis Streams. What does Redis Streams give up compared with a disk-backed replicated log, and where is it genuinely the better choice?

level: seniorimportance: should knowfreq 42%

answer

  1. memory-resident, one key = one shard
  2. async replication: acked XADD can be lost on failover
  3. MAXLEN/MINID trimming ignores consumer lag
  4. no compaction, no rebalance protocol, no per-key ordering in a group
  5. wins: sub-ms latency, no new system, minutes-to-hours history

basics

~20 s

Redis Streams keep history in RAM on one shard, replicate asynchronously, and trim by length or ID - so retention is short, an acknowledged write can be lost on failover, and one stream key cannot be partitioned across nodes. They win on latency, operational simplicity, and when Redis is already in the stack.

solid answer

~60 s

What Streams give up: - **Durability.** Entries live in memory; persistence is RDB/AOF with a loss window (about a second under `appendfsync everysec`), and replication is **asynchronous**, so an `XADD` already acknowledged to the producer can vanish in a failover. A disk-first log with synchronous replication to a quorum does not have that hole. - **Retention.** History is bounded by RAM and trimmed by `MAXLEN`/`MINID`, not by days of disk. Trimming is also unaware of consumer progress, so a lagging group can lose entries. - **Scale of one topic.** A stream is a single key, so it lives on one shard - no partitioning of one logical topic across nodes, and hash tags cannot split a key. Throughput and size ceiling are one node's. - **Ecosystem.** No log compaction, no built-in rebalancing protocol, no connector/stream-processing ecosystem. Where it wins: sub-millisecond end-to-end latency, no extra system to operate, simple consumer groups with acknowledgement, and workloads whose useful history is minutes to hours - job queues, task fan-out, short-lived event buffers.

code

text · 6 lines
text
XADD events MAXLEN ~ 1000000 * type order id 42   # retention is a length, not a duration
XADD events MINID ~ 1699996400000 * type order id 43  # or a time-derived floor

WAIT 1 100        # narrow (not close) the async-replication loss window
XINFO GROUPS events    # per-group lag and pending count
XPENDING events billing

go deeper

for a junior

Know the headline: Streams keep a bounded history in memory on one node, so they replay recent events but are not a long-retention durable log.

for a middle

Name the concrete limits - MAXLEN trimming, one key equals one shard, async replication - and the latency and simplicity upside.

for a senior

Lead with the decision criteria (consequence of loss, required replay depth, per-key ordering) and back them with the specific mechanisms and the WAIT caveat.

for a principal

Argue total cost of ownership against risk: what an acknowledged-then-lost write costs the business, whether an outbox or a durable log is warranted, and what you would build if you kept Streams.

## Frame the comparison honestly Redis Streams borrowed the shape of an append-only log: immutable ordered entries, reader cursors instead of destructive reads, consumer groups with acknowledgement. That similarity is exactly why the substitution gets proposed. The differences are not in the API shape but in the **storage and replication substrate underneath it**. ## What Streams give up **1. Durability semantics.** A Redis primary replies to `XADD` as soon as the entry is in memory. Persistence happens on its own schedule - RDB snapshots periodically, AOF with `appendfsync everysec` by default, which means a crash can lose roughly the last second of writes; `appendfsync always` closes most of that at a large throughput cost. Replication is asynchronous: replicas receive the entry after the client has already been told it succeeded. If the primary fails and a replica is promoted, entries the producer believes are committed can be missing. `WAIT numreplicas timeout` lets a producer block until N replicas have acknowledged, which narrows the window, but it is a best-effort barrier rather than a quorum commit and it costs a round trip per write. A durable log designed for the job writes to disk on multiple brokers before acknowledging, and its ack policy is a first-class producer setting. **2. Retention.** Stream entries occupy RAM on the shard that owns the key. You bound that with `XADD ... MAXLEN ~ 1000000` or `MINID`, or periodic `XTRIM`. Two consequences: your usable history is measured in what you are willing to hold in memory, and **trimming does not consider consumer progress** - if a group is down longer than your retention window, its unread entries are simply gone. A disk-first log routinely retains days or weeks because disk is cheap, and can be configured to retain by time. **3. Partitioning of a single topic.** A stream is one Redis key. In Redis Cluster a key lives entirely on one shard, so one stream cannot be spread across nodes: its write throughput, its memory and its history are bounded by one node's event loop and RAM. You can shard manually - `events:0` … `events:15`, chosen by a hash of the entity id, each read by its own consumers - but then *you* own partition assignment, rebalancing when consumers come and go, and per-partition cursors. A partitioned log makes that the built-in model, including consumer-group rebalancing and per-key ordering by construction. **4. Ordering guarantees you may assume you have.** Within one stream key, entry IDs are strictly increasing and the single-threaded server serializes appends, so total order per key is solid. But inside a consumer group, entries are handed to whichever consumer asks next: two events for the same entity can be processed concurrently by different consumers. Per-key ordering is not provided; you must partition to get it. **5. Ecosystem features.** No compaction (keeping only the latest entry per key), no schema registry, no exactly-once transactional produce-consume, no ready-made connectors or stream-processing runtime. Redis gives you the primitive and stops. ## Where Streams genuinely win **Latency.** Everything is in memory and the protocol is one round trip. `XADD` to `XREADGROUP BLOCK` delivery is sub-millisecond on a local network - materially faster than a system designed around disk batching, and often the actual product requirement for task dispatch or live coordination. **Operational cost.** If Redis is already deployed, Streams are a data type, not a new distributed system to run, monitor, upgrade and staff. That is frequently the dominant factor for a small team, and the honest form of the argument is not technical purity but total cost. **Simplicity of the consumer model.** No rebalance protocol to understand: a consumer is a name, appearing and disappearing freely, with pending entries claimable by any peer. For a job queue this is easier to reason about than partition assignment. **Right-sized problems.** Work queues, task fan-out to workers, short-lived event buffers, per-user activity feeds with capped length, real-time coordination where minutes of history is all that matters, and anything where the volume comfortably fits one shard's RAM and one core's throughput. ## How to answer the team Ask what happens when a message is lost, and how far back a consumer must be able to replay. If the answer is "lost is a bug we must never have" and "days", Redis Streams are the wrong substrate and the honest fix is either a durable log or an outbox in a database. If the answer is "a lost task is retried by an upstream timeout" and "minutes", Streams are likely the cheaper, faster, simpler choice - and the one already in your stack. Then set retention deliberately against your slowest consumer's worst outage, monitor group lag with `XINFO GROUPS` and pending backlogs with `XPENDING`, and decide up front whether you need manual sharding for per-key ordering.

  • Can a single Redis Stream be spread across cluster shards for more throughput?
    No - a stream is one key, and a key lives entirely on one shard, so its throughput and memory ceiling are one node's. You can shard manually by writing to several stream keys chosen by a hash of some entity id and assigning consumers to keys, but then partition assignment, rebalancing and per-partition cursors become your responsibility.
  • What does WAIT do for producer durability, and what does it still not guarantee?
    WAIT numreplicas timeout blocks until the given number of replicas have acknowledged the writes issued so far, or the timeout elapses, and returns how many did. It narrows the async-replication loss window but is not a quorum commit: it does not roll anything back on timeout, it does not make the write atomic across replicas, and it cannot protect against a failover that promotes a replica that had not caught up.
  • How do you size retention for a Redis Stream used as an event buffer?
    Size it against your slowest consumer group's worst tolerable outage, multiplied by peak ingest rate and average entry size, plus headroom - because trimming is blind to consumer progress and will discard unread entries. Use approximate trimming (MAXLEN ~) so trimming stays cheap, monitor group lag with XINFO GROUPS, and alert before lag approaches the retention window.

saying these in an interview costs you the question

  • Calling Redis Streams durable without mentioning async replication or the fsync window
  • Believing one stream key can be partitioned across cluster nodes
  • Assuming trimming waits until all consumer groups have read an entry
  • Expecting per-key ordering from a consumer group without manual sharding
  • Choosing on feature checklists rather than on the consequence of message loss and required replay depth

context