skip to content

questions

6

When a consumer successfully processes a message from a traditional message queue like RabbitMQ or SQS, what typically happens to that message, and how does that differ from a system like Kafka?

level: juniorimportance: must knowfreq 70%

answer

  1. destructive read vs append-only
  2. queue deletes on ack
  3. log retains + offsets
  4. replay requires a log
  5. multiple independent readers need a log

basics

~20 s

In a queue, once a message is picked up and confirmed done, it's deleted - gone for good. In an event log like Kafka, the message stays stored so other readers can still see it later, or the same reader can go back and read it again.

solid answer

~30 s

Traditional message queues (RabbitMQ, SQS) are destructive-read: once a consumer acknowledges a message, the broker removes it from the queue and it cannot be read again by anyone. Event logs (Kafka) are append-only and durable: messages are retained for a configured period regardless of whether they've been read, so multiple consumers can each read the same message independently, and any consumer can rewind and re-read history. This single difference drives almost everything else: fan-out, replay, and how 'done' is tracked.

go deeper

for a junior

Should state the core distinctive fact - queue removes on ack, log retains - and give one consequence (replay or fan-out).

for a middle

Should explain the mechanism (offsets vs ack/visibility-timeout) and connect it to at least one operational implication like reprocessing.

for a senior

Should discuss trade-offs (storage cost, consumer complexity) and pick the right tool for a stated scenario (task distribution vs event broadcast).

for a principal

Should reason about how the choice affects system evolvability - e.g., adding new consumers years later - and know when a hybrid or dual-write approach is warranted.

## Where the mechanism differs The mechanism differs at the storage layer. - **A queue broker** holds a message until a consumer acknowledges it (`ack`), at which point the broker physically deletes it from its store; while a message is 'in flight' it's typically marked invisible for a visibility-timeout window, and if the ack never arrives it's redelivered. - **An event log**, by contrast, appends every message to an immutable, ordered partition and never deletes on read - deletion happens only when a separate, time- or size-based retention policy expires a segment. Consumers of a log don't tell the broker 'I'm done'; they simply remember an **offset** (a position number) marking how far they've read, and can move that number forward, backward, or not at all without changing what's stored. ## Why the two models diverge This design split exists because the two models solve different problems. - **Queues model discrete units of work**: a task that should be picked up by exactly one worker, processed, and then be off the books - the classic 'competing consumers' pattern for load distribution. - **Logs model facts that happened**: an immutable record of history that many independent parties might want to observe, now or later, much like a database's write-ahead log exposed externally for anyone to tail. ## The trade-offs The trade-offs follow directly. | | A queue | A log | |---|---|---| | **Storage** | keeps storage small (consumed messages vanish) | costs more to operate (retention window means disk usage for however many days/GB you configure, replicated for durability) | | **Consumer** | keeps the consumer simple (no offset bookkeeping, no replay logic needed) | pushes offset management onto every consumer | | **History** | but you permanently lose history - you cannot add a second independent consumer after the fact and have it see what already happened, and you cannot recover from a downstream bug by reprocessing | but in exchange it supports replay for bug fixes and backfills, and it supports adding brand-new consumer groups years later with zero coordination with the producer | ## Failure modes in production In production, the failure modes look different too. - **A queue's classic failure** is the poison-pill message that keeps failing and getting redelivered until a dead-letter queue (DLQ) catches it after N attempts - a clean, per-message isolation mechanism. - **A log's classic failure** is a consumer that commits its offset too early and crashes, silently skipping unprocessed messages (at-most-once), or commits too late and reprocesses on restart (at-least-once, requiring idempotent handling). A slow log consumer doesn't block other consumers - it just accumulates lag on its own bookmark - whereas a queue backlog is a single shared pile that every consumer of that queue competes over. ## One event, two outcomes A concrete scenario makes the contrast vivid: an order-service publishes 'OrderPlaced.' - **Using SQS**, only the fulfillment worker that dequeues a given message ever sees it; if six months later a billing team wants to react too, they need a new fan-out subscription set up going forward, and they get nothing from before that point since nothing historical survived. - **Using Kafka**, order-service publishes to a topic; a fraud-detection team spun up half a year later can create a brand-new consumer group and read the exact same historical events from the beginning of the retention window, entirely independent of the original fulfillment consumer and without any change to the producer.

  • If a queue consumer crashes after reading but before acking a message, what happens?
    Most brokers use a visibility timeout: the message becomes invisible temporarily but not deleted, then reappears for redelivery once the timeout expires. This gives at-least-once delivery but means a slow or crash-prone consumer can cause duplicate processing. RabbitMQ's unacked-message tracking and SQS's visibility timeout both work this way.
  • Can you bolt a second, independent consumer onto an existing SQS-style queue without changing the producer?
    Not directly - a single queue delivers each message to exactly one consumer, so a second consumer would only get whatever the first one missed. You'd typically add a fan-out layer (e.g., SNS publishing to multiple SQS queues), which is effectively re-adding a broadcast/log-like layer in front of the queues.
  • Does an event log guarantee exactly-once processing?
    No - the log itself only guarantees ordered, durable delivery within a partition; exactly-once semantics require the consumer to make offset-commit and processing atomic, for example via idempotent writes keyed by a stable ID, or a transactional producer/consumer API. Otherwise you get at-least-once with possible duplicates.

A queue is like a physical inbox tray - once you take a letter, it's gone from the tray. A log is like a shared bulletin board where notices stay pinned for weeks - everyone walking by can read the same notice, and you can re-read one you already saw.

saying these in an interview costs you the question

  • Says Kafka messages disappear once read like a queue
  • Doesn't know that queue messages are deleted after ack
  • Thinks queues support replay out of the box
  • Confuses 'multiple consumers' with 'multiple independent consumer groups'
  • Can't explain why storage grows differently between the two

context

open as a page

Concretely, how does a Kafka-style consumer track 'what have I already processed' using offsets, and how is that different from how an SQS or RabbitMQ consumer acknowledges a message?

level: middleimportance: must knowfreq 75%

basics

~20 s

A log consumer remembers a number (an offset) marking its place in an ordered list of messages, and moves that number forward as it reads. A queue consumer instead tells the broker 'I'm done with this one specific message,' and the broker deletes just that message.

open as a page

A consumer service has a bug that silently corrupts 3 days' worth of derived data before anyone notices. If the upstream system is a Kafka-style event log with 14-day retention, how would you recover, and why would the same recovery be much harder (or impossible) if the upstream were an SQS queue instead?

level: seniorimportance: must knowfreq 65%

basics

~20 s

With a log, you can rewind your reader to before the bug started and reprocess those 3 days from the original messages, since they're all still stored. With a queue, those messages are already deleted once the first, buggy pass consumed them, so there's nothing left to replay - you'd need another way to get that data back.

open as a page

If three different services (billing, shipping, and analytics) all need to react to the same 'OrderPlaced' notification, how does supporting that differ structurally between a destructive-read queue like RabbitMQ and an append-only log like Kafka?

level: middleimportance: should knowfreq 60%

basics

~20 s

With a queue, you need three separate queues (or a fan-out router) so each service gets its own copy, because one message can only be taken once. With a log, all three can just read the same shared stream independently, each keeping its own place.

open as a page

You're designing a system to send password-reset emails: an API call enqueues a 'send reset email' task, and a worker pool sends it exactly once soon after. Would you reach for a destructive-read queue like SQS or an event log like Kafka, and why?

level: seniorimportance: should knowfreq 55%

basics

~20 s

A queue - you want the task done once by one worker and then forgotten, with automatic retry if it fails. A log is overkill here because nobody else needs to replay or re-read 'send this email' later; you'd actually want to avoid accidentally resending it.

open as a page

A single malformed message lands in position 500 of a Kafka partition that a consumer group is processing sequentially, and the consumer's deserializer throws every time it hits that offset. How does the resulting outage differ from what would happen if the same bad message had instead been delivered via an SQS queue, and why?

level: principalimportance: nice to knowfreq 35%

basics

~20 s

In Kafka, that one bad message can jam up the whole line behind it, because the consumer must go through messages in order - everything after gets stuck waiting. In SQS, a bad message just gets set aside (retried, then dead-lettered) while all the other messages keep moving normally.

open as a page