skip to content

What is the fundamental difference between Kafka's log-based model and a traditional message queue like RabbitMQ or ActiveMQ?

level: juniorimportance: must knowfreq 85%

answer

  1. append-only log vs delete-on-ack
  2. offset = consumer's own position
  3. non-destructive read, retention deletes
  4. queue = destructive consume
  5. replay + fan-out for free

basics

~10 s

Kafka stores messages as an append-only log that consumers read by position (offset); messages stay for a retention period. A traditional queue deletes each message once a consumer acknowledges it, so reading is destructive.

solid answer

~40 s

Kafka is a distributed, append-only commit log. Producers append records to a partition; each record gets a monotonically increasing offset. Consumers track their own offset and read forward, so reading is non-destructive — the same record can be re-read or read by many independent consumer groups, and stays until a retention policy (time/size) deletes it. A traditional broker (RabbitMQ, ActiveMQ, SQS) models a queue: the broker pushes/dispenses a message, the consumer acknowledges it, and the broker then removes it. Consumption is destructive and the broker tracks per-message state. Consequence: Kafka excels at replay, multi-subscriber fan-out, and high-throughput streaming; classic queues excel at per-message work dispatch with rich routing and per-message acks. Kafka offloads delivery-state tracking to the consumer (the offset), which is why it scales to very high throughput.

go deeper

for a junior

Know the one-liner: log you read by position vs queue that deletes on ack.

for a middle

Explain offsets, retention, and that reading is non-destructive; tie to replay and fan-out.

for a senior

Discuss why offloading state to the consumer enables throughput, and the per-partition ordering trade-off.

for a principal

Frame as a storage/dataflow architecture decision: stream-as-source-of-truth vs transient work dispatch, and the downstream coupling implications.

## The two models **Message broker / queue model (RabbitMQ, ActiveMQ, AWS SQS):** A *queue* is a holding area. A producer sends a message; the broker stores it and is responsible for delivering it to a consumer. When a consumer processes the message and sends an **acknowledgement (ack)**, the broker **deletes** the message. This is *destructive consumption* — once acked and removed, the message is gone. The broker maintains per-message delivery state (delivered? acked? redelivered?). RabbitMQ/ActiveMQ also add **exchanges/topics and routing** (direct, topic, fanout, headers) so one message can be copied into multiple queues at publish time. **Log model (Kafka):** A Kafka **topic** is split into **partitions**, and each partition is an **append-only commit log** — an ordered, immutable sequence of records on disk. Every record in a partition gets an **offset**: a monotonically increasing integer position (0, 1, 2, …). Producers only ever *append* to the end. Consumers *read forward* from a chosen offset and **track their own position**. Reading does NOT delete anything. Records are removed only by a **retention policy**: time-based (`retention.ms`, default 7 days) or size-based (`retention.bytes`), or compacted (`cleanup.policy=compact`, which keeps the latest value per key). ## Why this matters 1. **Replay:** Because records persist and consumers control their offset, you can rewind (`auto.offset.reset=earliest`, or seek to a specific offset/timestamp) and reprocess history — e.g., to fix a bug or feed a new model. A queue can't replay; acked messages are gone. 2. **Fan-out without copies:** Multiple **consumer groups** each maintain independent offsets over the *same* partitions, so N teams read the full stream without the broker duplicating data. In RabbitMQ you'd bind N queues to a fanout exchange (N physical copies). 3. **Throughput:** Kafka does sequential disk I/O (append + sequential reads) and pushes delivery-state tracking to the consumer (one offset per partition), instead of tracking state for every individual message. That's why Kafka sustains very high throughput. 4. **Ordering:** Order is guaranteed only *within a partition*. Queues with competing consumers generally don't guarantee global order either, but the unit differs. ## Edge cases / nuance - Kafka offsets are committed in `__consumer_offsets`; if a consumer doesn't commit, it re-reads on restart (at-least-once). - 'Destructive' queues can still support redelivery via nacks/dead-letter queues, but that is per-message machinery, not replay of arbitrary history. - Kafka retention is independent of consumption: a slow consumer can fall off the back of the log and lose data if retention expires before it catches up.

  • If reading doesn't delete the message, when does Kafka actually remove data?
    Only via retention: time (retention.ms), size (retention.bytes), or log compaction (cleanup.policy=compact keeps the latest record per key). Consumption never deletes.
  • Who tracks delivery state in Kafka vs RabbitMQ?
    Kafka pushes it to the consumer as a committed offset (stored in __consumer_offsets). RabbitMQ tracks per-message state on the broker and removes the message on ack.

saying these in an interview costs you the question

  • Saying Kafka deletes a message after a consumer reads/acks it (it does not — only retention removes data).
  • Claiming the broker tracks each consumer's position in Kafka — consumers track their own offset.
  • Saying a traditional queue supports replay of arbitrary history the way Kafka does.

context