Your platform team owns the guidance for new event flows at a company where every team already runs its own Redis instance. Take the technical comparison as given — Redis Pub/Sub has no acknowledgement and drops messages for disconnected or slow subscribers, Redis Streams keep a bounded in-memory history on a single key with consumer groups, and a partitioned durable log keeps days of replay with per-key ordering. What organizational rule decides whether a flow is allowed to stay on Redis at all, what do you require of every flow regardless of which primitive is chosen, and what do you tell a team whose acknowledged write must never be lost?
answer
- counterparty + lifetime, not throughput
- team-internal stays; cross-team leaves
- consumers the producer cannot enumerate
- four musts: idempotent, retention+lag alert, written durability, DLQ
- acknowledged-write-can't-be-lost → outbox upstream
basics
~20 sAsk who is on the other end of the contract and how long it lives. Team-internal, short-lived, team-operated flows stay in Redis; cross-team, long-lived, unenumerable consumers move to a durable log. Always require idempotent consumers, sized retention with lag alerting, a written durability expectation, and a dead-letter path.
solid answer
~50 sI decide with one question first: **who is on the other end of the contract, and how long does it live?** - **Team-internal, short-lived, operationally owned by the team that runs the Redis instance** — it may stay in Redis. Work queues, task fan-out, cache invalidation, capped per-user feeds. - **Cross-team, long-lived, or with consumers the producer cannot enumerate** — it leaves Redis for the durable log. Not because Redis cannot move the bytes, but because that contract needs days of replay, backfill for consumers that do not exist yet, and a schema/compatibility story. A mesh of per-team Redis instances carrying cross-cutting events is unreplayable after an incident. Whatever is chosen, I require four things: idempotent consumers, retention sized against the worst outage I am willing to survive with lag alerting (`XINFO GROUPS`) firing before it is approached, a written durability expectation, and an application-owned dead-letter path. If an acknowledged write must never be lost, no broker setting fixes that — the answer is a transactional outbox upstream.
code
text · 14 lines# publish with an explicit retention bound (approximate trim is cheaper)
XADD orders MAXLEN ~ 1000000 * type completed order_id 8812
# what the alert scrapes: per-group backlog and stuck entries
XINFO GROUPS orders
# 1) 1) "name" 2) "fraud-svc"
# 3) "consumers" 4) (integer) 4
# 5) "pending" 6) (integer) 137 <- delivered, not yet ACKed
# 7) "last-delivered-id" 8) "1723641000123-0"
# 9) "entries-read" 10) (integer) 998412
# 11) "lag" 12) (integer) 1588 <- unread entries still in the stream
# oldest un-ACKed entry age -> poison-message / dead-letter trigger
XPENDING orders fraud-svc - + 1go deeper
Know the shape of the rule: events used only inside your own team can live on your team's Redis; events other teams depend on belong on the company's durable event log. Know that consumers may see a message twice, so handlers must be safe to re-run.
Be able to justify the line mechanically: Redis retention is a memory budget that trims without regard to consumer lag, so a consumer that has been down too long silently loses data — which is untenable for a consumer you do not operate. Name the four requirements you would put on any flow.
Own the operational side: pick a retention window from the worst outage you will tolerate, alert on XINFO GROUPS lag and XPENDING age before that window closes, define the dedupe key and the dead-letter path, and write down whether an acknowledged publish may be lost.
Frame it as contract governance, not technology selection: counterparty plus lifetime decides placement; cross-team flows on per-team Redis create undeclared availability dependencies and an unreplayable mesh. Then state the residual risk honestly — durability settings never close the acknowledged-write gap, a transactional outbox does — and describe how you enforce the rule by making the compliant path self-service rather than by adding a review board.
## The decision that "which primitive" hides When a team asks "Pub/Sub, Streams, or a log?", the technical comparison — loss semantics, how much replay fits in RAM, fan-out and ordering shape — is a separate question and is covered in full by the dedicated Pub/Sub vs Streams vs durable-log comparison; assume the team already holds it. What that comparison does **not** answer is the platform-level call: even when Redis Streams technically fit today, should this flow be allowed to live on a Redis instance a single team owns and operates? That is a governance decision, and it is the one a principal is actually being asked about. ## The rule: counterparty and lifetime I lead with two properties of the *contract*, not of the technology: 1. **Who is on the other end?** Consumers I can name and whose owners sit in my team, versus consumers the producer cannot enumerate — including consumers that do not exist yet. 2. **How long does the contract live?** A flow whose meaning is exhausted in seconds, versus one that will be depended on for years and whose payload shape becomes a de facto public API. A third, quieter property: **who operates the storage?** If the Redis instance is a team's own cache box, sized and restarted by that team, then every cross-team consumer has taken an undeclared dependency on another team's operational decisions. ## What stays in Redis Team-internal, short-lived, operationally owned: background work queues consumed by the same team's workers, task fan-out, cache-invalidation hints, presence and live UI ticks, capped per-user feeds. Here Redis is not a compromise — it is already deployed, latency is sub-millisecond, the consumer model is small, and if the shape of the event changes, one team changes both ends in one deploy. Insisting on a durable log for these buys nothing and adds an operational hop. ## What leaves Redis Anything crossing a team boundary with an open-ended lifetime: domain events like `order.completed` that finance, fraud, and analytics all consume; anything a future service must backfill from; anything an auditor may ask you to reconstruct. The reasons are organizational before they are technical: - **Replay for consumers who did not exist at publish time.** A cross-team contract implies someone will onboard later and need history. Redis retention is a memory budget and trimming ignores consumer lag, so the producing team would be silently rationing other teams' backfill capability. - **Schema and compatibility.** A long-lived contract needs a place where versioning and compatibility are enforced and visible. A team's Redis key namespace is not that place. - **Blast radius and ownership.** The producing team now owns an availability and capacity SLO for consumers it never agreed to serve, on an instance it also uses as a cache. - **Incident reconstruction.** Ten teams each putting cross-cutting events on their own Redis produces a mesh with no shared retention, no shared tooling, and nothing to replay after an outage. That cost lands on the platform, not the team that made the local decision. The honest framing to a team: the log is not "more correct", it is where a contract with strangers belongs. ## Non-negotiables, whatever is chosen The rule above decides *where*; these four are conditions of shipping at all. - **Idempotent consumers.** Everything on the table is at-least-once at best. Consumers dedupe on an event id or make the effect naturally repeatable. If a team cannot state its dedupe key, the design is not finished. - **Retention sized against the worst outage you are willing to survive**, not the typical one — plus lag alerting that fires well before lag approaches that window. On Redis Streams that is `XINFO GROUPS`/`XPENDING` monitoring, alerting on a group's lag and pending count; on Redis 7.0+ `XINFO GROUPS` exposes a `lag` field directly. A team that keeps one hour of history is stating that a consumer down for 61 minutes loses data — I make them say it out loud. - **A written durability expectation.** "Can an acknowledged publish be lost, and is that acceptable?" must have an answer in the design doc, not an assumption. Most flows can tolerate it; the ones that cannot need the outbox below. - **An application-owned dead-letter path.** Neither Redis primitive gives you one. Streams' `XAUTOCLAIM`/delivery counter lets you detect a poison entry, but routing it somewhere durable and alerting a human is application code that must exist before launch. ## Residual risk: the outbox When a flow's requirement is "once we acknowledge the write to the caller, the event cannot be lost", no configuration of the transport closes that gap — the write to the database and the publish are two operations that can fail between. Turning up `appendfsync` narrows the window at real throughput cost, and replication stays asynchronous, so a failover can still promote a replica missing recent entries. The pattern that actually closes it is a **transactional outbox**: the service writes the business change and an event row in one database transaction, and a relay publishes from that table to Redis or the log. The source of truth can always re-emit, which is precisely what makes at-least-once transport acceptable underneath it. ## Enforcing it without becoming a bottleneck I encode the rule as a default rather than a review queue: cross-team topics get provisioned on the log with schema registration; team-internal flows need no approval at all. The only thing I gate is the crossing — the moment a second team subscribes to a key on someone's Redis, that flow has become a contract and must move.
- A team says their flow is team-internal today, so it stays on their Redis — but you know a second team is likely to want it next quarter. How do you handle that?I treat "a second consumer is foreseeable" as already across the line, because the migration cost lands exactly when the flow is busiest and the new consumer needs history nobody kept. If the likelihood is genuinely speculative, I let it stay but make the exit explicit: the event payload is versioned from day one, and the design doc names the trigger — the first external subscriber — at which point it moves to the log. What I will not accept is the second team quietly subscribing to another team's Redis key, because that creates a contract with no owner, no retention agreement, and no schema story.
- A team argues Redis Streams are fine for a cross-team flow because they will set appendfsync always and add replicas. Does that close the durability gap?It narrows it at real cost but does not close it. `appendfsync always` makes every write wait on disk and cuts throughput substantially, and replication stays asynchronous, so a failover can promote a replica missing entries a producer already saw acknowledged; `WAIT` is a best-effort barrier, not a quorum commit. More importantly it answers the wrong objection: the reason a cross-team flow leaves Redis is replay for unknown consumers, schema ownership, and blast radius, none of which durability settings touch.
- How do you keep this rule from becoming a review bottleneck that teams route around?I make the compliant path the cheap one: team-internal flows need no approval at all, and cross-team topics are self-service on the log with schema registration built into the provisioning step. The only gate is the boundary crossing itself. I also detect violations rather than relying on goodwill — cross-team subscriptions to another team's Redis instance show up in connection metadata and in dependency reviews, and when one appears I treat it as a migration ticket, not a policy argument.
Redis is the whiteboard in your team's room: perfect for what the people in that room need today, useless as the company's filing system. The moment another department starts reading it, it stopped being a whiteboard and became a record — and records live somewhere with retention and an index.
saying these in an interview costs you the question
- Choosing by throughput or latency benchmarks when the deciding property is who owns the contract and how long it lives.
- Claiming a durable log is always the professional choice, and pushing team-internal work queues and cache invalidation off Redis for no benefit.
- Believing `appendfsync always` plus replicas makes Redis safe for a flow where an acknowledged write must never be lost — replication is asynchronous and `WAIT` is not a quorum commit.
- Sizing Redis Stream retention against typical consumer lag instead of the longest outage you are willing to survive, and having no alert before the window is consumed.
- Assuming Redis Streams provide a dead-letter queue or exactly-once delivery, so consumers need no dedupe key and no poison-message path.