skip to content

How does a broker's message retention policy differ between 'delete after acknowledgment' queues and 'keep for a fixed time/size regardless of consumption' log retention, and what does each let you do that the other doesn't?

level: middleimportance: should knowfreq 60%

answer

  1. queue deletes on ack
  2. log retains by time/size regardless of read
  3. replay only possible with log retention
  4. offset tracking is consumer's job in log model
  5. retention window too short = permanent gap for late consumers

basics

~20 s

Some brokers throw a message away as soon as it's been successfully handled once. Others keep every message around for a set amount of time no matter who's read it, so you can go back and re-read old messages later, at the cost of using more storage.

solid answer

~40 s

Classic queue retention is consumption-based: once a message is acknowledged by its consumer, it's deleted, so the queue only ever holds unconsumed backlog - simple and storage-efficient, but a message is gone forever once processed, so a new consumer added later gets nothing from before it joined. Log-based retention (as in Kafka) is time/size-based and independent of consumption: messages are retained for a configured window regardless of whether any consumer has read them, so multiple consumers can read at different paces, replay history, or a brand-new consumer can start from the beginning of the window. The trade-off is storage cost and consumer responsibility: log retention needs more disk and pushes offset tracking onto each consumer rather than the broker deleting on your behalf.

go deeper

for a junior

Can explain that some systems throw messages away once read, while others keep them for a while regardless of who's read them.

for a middle

Can match a retention model to a scenario and explain the basic storage trade-off.

for a senior

Can reason about offset-management failure modes and retention-window sizing trade-offs, and design around a consumer falling permanently out of the retention window.

for a principal

Can design a retention strategy for a system anticipating future consumers and compliance/audit needs, balancing storage cost against replay value.

## The two retention models - **The consumption-based retention model.** A message's lifecycle is tied directly to acknowledgment: the broker holds it, delivers it to a consumer, and upon receiving a positive acknowledgment, removes it from storage — the message ceases to exist after that point. The queue's storage footprint at any moment reflects only the current unconsumed backlog. - **The log-based retention model.** Published messages are appended to a durable, ordered log and kept according to a policy configured independently of consumption — typically 'keep for N days' or 'keep up to N gigabytes per partition' — and consumption never triggers deletion. Instead, each consumer tracks its own read offset into the log; the broker doesn't know or care whether a given message has been consumed by anyone, only that it hasn't yet aged out of the retention window. ## Why each one exists Consumption-based retention exists because, for pure task-distribution workloads, once a task is done there is nothing more to do with the message — keeping it forever would waste storage for no benefit. Log-based retention exists for event streaming, where the same events may need to be read by consumers that don't yet exist, by consumers at very different speeds, or by the same consumer intentionally rewinding after fixing a bug in its processing logic. None of that is possible if the broker deletes a message the instant it's first read. ## The mirrored trade-off The consumption-based model's trade-off is the mirror of the log-based model's: | Retention model | What it buys | What it charges | |---|---|---| | Consumption-based | storage stays small and predictable | you permanently lose the ability to onboard a new consumer needing anything from before it existed, and you can't replay to recover from a downstream processing bug | | Log-based | you gain replay and multi-speed flexibility | you pay for it in disk (retaining every message for days multiplies storage requirements) and in shifted responsibility — every consumer must track and persist its own offset and handle corruption or resets | With a log you must also choose a retention window carefully: too short and a slow or down consumer falls out of the window and permanently misses data; too long and storage costs balloon for data nobody reads again. ## Failure modes - **Reprocess history after the fact.** A common failure with consumption-based queues is discovering, only after the fact, that you need to reprocess history — e.g., a consumer had a bug that silently corrupted data for a week, but since the queue deletes on acknowledgment, there's no way to replay those already-consumed messages; the only backstop is an audit log kept elsewhere, if one exists. - **Offset mismanagement.** A common failure with log-based retention is a consumer's stored offset being lost or reset to the beginning, after which it can reprocess the entire retention window's worth of messages, causing downstream duplicate-processing storms if the consumer isn't idempotent. - **Falling behind the window.** Another log-retention failure is a consumer falling behind its retention window entirely — if it's down longer than the retention period, the oldest messages it needed have already been deleted by the time it returns, and it silently has a permanent gap, with no error raised by the broker. ## Two workloads, two fits - **A payroll-processing queue** (consumption-based) is the right fit for 'pay this employee, once' — once paid and acknowledged, there's no value in keeping that task around, and doing so would just waste storage; if you need an audit trail, that's a job for a separate database table. - **By contrast, a clickstream/event-tracking pipeline** (log-based, e.g. 7-day retention) is the right fit because a new analytics team six months from now might want to build a fresh aggregation job and, thanks to log retention, can replay the last week of events to backfill and validate their new consumer against real historical data before going live — something a consumption-deleted queue could never offer.

  • If a new consumer is added to a system a month after it launched, what's the difference in what it can access under each retention model?
    Under consumption-based queue retention, the new consumer gets nothing from before it joined - every prior message was already deleted once its original consumer acknowledged it. Under log-based retention with, say, a 7-day window, the new consumer can read back up to 7 days of history, but nothing older than that.
  • Why does log-based retention push offset management onto the consumer instead of the broker?
    Because the broker no longer deletes messages on read, it has no natural signal for 'this consumer is done with this message' the way an acknowledgment gives a queue; multiple independent consumers might be at completely different read positions in the same retained log at once, so each must track and persist its own offset.
  • What's a concrete risk of setting a log retention window too short for a given consumer population?
    Any consumer offline longer than the retention window permanently misses the messages that aged out during its downtime, with no error surfaced by the broker - it just silently has a gap. This is a common cause of subtle data-completeness bugs discovered much later.

A consumption-based queue is like a to-do list where you cross off and throw away each task the moment it's done - nothing left to review later. Log-based retention is like a diary you keep for 30 days: entries stay whether or not you've read them, so you can flip back and reread any day within that window, but eventually old pages get torn out to save space.

saying these in an interview costs you the question

  • thinks all brokers delete messages immediately after any consumer reads them
  • assumes log-based retention keeps messages forever with no storage cost
  • doesn't realize offset tracking becomes the consumer's responsibility in a log model
  • can't explain why replay isn't possible with consumption-based queue retention
  • assumes a downed consumer can always catch up regardless of how long it was down

context