skip to content

After a message has been successfully consumed, how does its fate typically differ between a traditional point-to-point queue and a pub-sub topic with multiple independent subscribers, and why does that difference matter for durability?

level: seniorimportance: should knowfreq 65%

answer

  1. queue: delete-on-ack, single consumer
  2. topic: per-subscriber cursor, independent copies
  3. log model decouples retention from consumption
  4. retention window = time/size policy
  5. replay only possible if retained

basics

~20 s

In a queue, once one consumer finishes a message, it's normally deleted - gone for good. In pub-sub, each subscriber has its own independent copy and cursor, so one subscriber finishing doesn't remove the message for the others, and depending on the system it may still be replayable later.

solid answer

~50 s

In a classic point-to-point queue, a message is deleted from the broker once it's acknowledged by the single consumer that received it - there's no concept of other interested parties because the model assumes exactly one logical consumer of that unit of work. In pub-sub, each subscription tracks its own independent read position, so one subscriber finishing a message has zero effect on any other subscriber's copy; the message is only removed from the topic (or ages out) once all subscriptions have consumed it or a retention window expires, whichever the system enforces. Log-based pub-sub systems like Kafka go further: they retain messages for a configurable time or size window regardless of consumption, so subscribers, and even brand-new ones, can replay from any earlier offset, decoupling storage retention entirely from consumption state. This matters for durability because it determines whether you can recover from a subscriber bug, add a late-joining consumer, or reprocess history, versus a queue where once it's gone, it's gone.

go deeper

for a junior

Should know that a queue message disappears after being consumed, while pub-sub keeps independent copies/cursors per subscriber.

for a middle

Should be able to name at least one concrete broker implementing each retention model (e.g., SQS/RabbitMQ delete-on-ack vs. Kafka log retention) and know that replay is only possible where retention exists.

for a senior

Should reason explicitly about retention-window sizing trade-offs (storage cost vs. recovery/backfill capability) and design around consumer-lag monitoring to avoid falling outside the retention window.

for a principal

Should set organization-wide retention and archival policy (e.g., short-retention live topics plus a permanent archival sink) balancing cost, compliance/audit needs, and replay requirements across many topics and teams.

## Two questions behind durability Message durability and retention describe two related but distinct questions: how long does a message persist on the broker, and what determines when it's finally removed? The answers diverge sharply between point-to-point queues and pub-sub topics, and that divergence has real consequences for what recovery and replay options you have. ## Inside a point-to-point queue **Mechanism in a point-to-point queue.** A queue models the idea that a unit of work needs to be done exactly once, by whichever consumer picks it up. There is a single logical consumption event per message. - Once that one consumer acknowledges successful processing, the broker's job is done — it deletes the message (or, in systems with a retention/DLQ policy, moves it aside only on failure, not success). - There's no notion of a second interested party for that same message within the same queue; if a second party needs to see it too, you either publish a second copy to a second queue, or you're actually looking for pub-sub instead. - RabbitMQ, SQS, and traditional JMS queues all follow this model: successful consumption is the end of that message's life on the broker. ## Inside a pub-sub topic **Mechanism in pub-sub.** A topic models multiple independent parties, each of whom needs to see every message, at their own pace. The broker must therefore track a separate read cursor/subscription state per subscriber, not just per queue. - **RabbitMQ** implements this by physically fanning a published message out into a separate queue per subscription (via a fanout or topic exchange) — so under the hood it's actually N independent per-subscriber queues, each following normal queue deletion-on-ack semantics for its own copy. - **Kafka and similar log-based systems** (Google Pub/Sub's subscription model, Amazon Kinesis) instead keep one physical copy of the message in an append-only log and let each consumer group track its own offset into that log — deletion is decoupled entirely from any subscriber's consumption and instead governed by a retention policy (keep messages for 7 days, or keep up to a size limit per partition), whichever limit is hit first. This means a message can still be sitting there, fully intact, long after every current subscriber has already read past it. ## Why the two diverge **Why this design difference exists:** it follows directly from the cardinality difference between pub-sub and queuing generally. - A queue's job is to guarantee a task is done once; keeping the message around after that single consumer is done serves no purpose and would just waste storage. - A topic's job is to let arbitrarily many, mutually unaware subscribers each see the full stream; tying deletion to whether everyone has consumed it yet would mean one slow or broken subscriber can indefinitely delay cleanup for everyone, and would make it impossible to add a new subscriber next month that needs to start from a recent point in history. Decoupling retention from consumption (the log model) solves both problems: storage is bounded by a time/size policy you control directly, not by subscriber behavior, and new subscribers can join and read backward within the retention window. ## Storage against replay The trade-offs: | | Queue-style delete-on-ack | Log-based retention | |---|---|---| | **Storage** | queue-style delete-on-ack is simple and storage-efficient — you're never storing more than the current backlog | at the direct cost of storage: retaining every message for 7 days, or forever in an event-sourcing-style setup, means paying for that storage regardless of consumption, and requires capacity planning that a delete-on-ack queue never needs | | **Replay** | offers zero replay capability; if a consumer processed a message incorrectly and only realizes the bug a day later, that message is unrecoverable from the broker unless it was captured elsewhere, such as an audit log | gives replay — enormously valuable for recovering from bugs, backfilling a new analytics pipeline, or reprocessing after fixing a parsing error | ## Failure modes in production - **On the queue side:** a common one is a team that needs just-in-case replay capability but is using a delete-on-ack queue, discovers a processing bug in production, and has no way to recover the already-consumed messages — the fix has to be waiting for the next batch of similar events or reconstructing from a separate log that often doesn't exist. - **On the pub-sub log side:** the failure mode is retention mismanagement. Setting retention too short means a subscriber that's down for maintenance longer than the retention window permanently loses messages published during the gap, despite the durable label, while setting it unnecessarily long or unbounded on a high-volume topic can produce runaway storage costs that go unnoticed until a bill spike. ## Sizing a retention window A concrete example: a clickstream analytics pipeline publishes user-interaction events to a Kafka topic with 7-day retention. The real-time dashboard consumer group reads continuously and is largely caught up. Three months later, the data science team wants to backfill a brand-new model-training pipeline using the last month of events — but because retention is only 7 days, only the most recent week is available; the rest is permanently gone unless it was separately archived, for example to object storage via a sink connector. This is a textbook illustration of why retention window sizing is a deliberate architectural decision, not an afterthought, and why many teams pair short-retention live topics with a longer-retention or permanent archival sink for anything they might need to replay far into the future.

  • If RabbitMQ fanout exchanges physically create a separate queue per subscriber, does that mean RabbitMQ pub-sub also loses messages once consumed, like a plain queue?
    Yes, for a subscription that's already registered when a message is published - once that subscription's own queue-copy is acked, that copy is gone, with no replay. This is a key contrast with log-based systems like Kafka: RabbitMQ fanout gives fan-out to currently-existing subscriptions but not historical replay for new ones or reprocessing of already-acked messages.
  • How would you decide the right retention window for a Kafka topic in a real system?
    Balance the cost of storage against the realistic recovery window needed - long enough to cover the longest expected consumer outage (maintenance windows, incident response) plus buffer, and long enough to support any planned backfill/reprocessing use cases, while recognizing that anything needed indefinitely should go to a cheaper archival sink rather than inflating the live topic's retention indefinitely.
  • Does a longer retention window fix the lost-message problem entirely?
    Not by itself - it only helps subscribers who eventually come back within the window; a subscriber that's down longer than retention still loses that data permanently, so retention sizing has to be paired with monitoring/alerting on consumer lag so operators can react before a slow or stalled subscriber falls outside the window.

A point-to-point queue is like a single physical letter handed to one recipient - once they've read it, it's gone. A log-based pub-sub topic is like a public bulletin board where the letter stays pinned up for a set number of days; anyone can walk by and read it during that window, and whether or not someone's already read it doesn't take it down.

saying these in an interview costs you the question

  • Assumes pub-sub messages are always deleted the moment any subscriber consumes them
  • Doesn't know that a delete-on-ack queue offers no replay capability
  • Thinks Kafka retention is unlimited by default with no cost implication
  • Confuses durable (survives broker restart) with retained for replay
  • Assumes a new subscriber added later can always see all historical messages regardless of retention settings

context