skip to content

A team asks whether to use Kafka or a traditional message queue (RabbitMQ/SQS) for a new system. How do you decide, and what workloads favor each?

level: seniorimportance: should knowfreq 60%

answer

  1. replay + multi-consumer + throughput → Kafka
  2. per-message acks/priority/TTL/DLQ → queue
  3. task dispatch vs event log
  4. partition count caps Kafka parallelism
  5. managed SQS = low ops; Kafka = heavy ops
  6. hybrid: Kafka spine + queue for tasks

basics

~20 s

Choose Kafka for high-throughput event streams that need replay, multi-consumer fan-out, and ordered-per-key data. Choose a queue for task/work dispatch needing per-message acks, elastic workers, priorities, and rich retry/DLQ, with no need to replay history.

solid answer

~50 s

Frame it by workload, not hype. Kafka fits an **event log / streaming backbone**: high and sustained throughput, multiple independent consumers reading the same data, the need to **replay** history (reprocessing, new consumers backfilling, event sourcing), ordered-per-key streams, and integration with stream processing (Kafka Streams, Flink, ksqlDB) and connectors. Traditional queues fit **command/task dispatch**: each message is a unit of work consumed once and acked, you want an elastic pool of competing workers, fine-grained **per-message** features (priority, TTL, delayed delivery, selective retry, dead-letter queues), and request/response or RPC-style patterns. Also weigh operational cost: Kafka (brokers, partitions, retention, lag monitoring; ZooKeeper or KRaft) is heavier than managed SQS or a single RabbitMQ. Pitfalls: don't pick Kafka just for 'a queue' if you really need per-message redelivery and priorities; don't pick a queue if you'll later need replay or many independent subscribers. Many architectures use both: Kafka as the event spine, a queue for targeted task processing.

go deeper

for a junior

Know the rough split: Kafka for streams/replay/many consumers, queue for one-off tasks with acks.

for a middle

List concrete deciding factors (replay, fan-out, per-message features, throughput) and matching examples.

for a senior

Run a structured decision and call out anti-patterns and operational cost; recognize hybrid designs.

for a principal

Set org-wide guidance: when Kafka is the event backbone vs when queues own task processing, factoring team ops capacity and migration cost.

## A decision checklist Ask what the system actually needs: 1. **Do multiple, independent consumers need the same data?** → Kafka (consumer groups, one stored copy). A queue would need physical fan-out per subscriber. 2. **Do you need to replay / reprocess history** (bug fixes, new consumers backfilling, event sourcing, ML feature recompute)? → Kafka (retention + offsets). Queues delete on ack. 3. **Is it high, sustained throughput** (hundreds of thousands+ msgs/sec, ordered per key)? → Kafka's sequential-log design and partition parallelism shine. 4. **Is each message a discrete task** consumed once, where you need **per-message** control: priority queues, delayed/scheduled delivery, message TTL, selective retry, dead-letter routing? → Traditional queue (RabbitMQ is especially rich here; SQS gives DLQ + visibility timeout + delay). 5. **Do you need elastic, fine-grained worker scaling** beyond a fixed partition count? → Queues let you add competing consumers freely; Kafka caps useful parallelism per group at partition count. 6. **Request/response or RPC semantics, point-to-point routing?** → RabbitMQ (exchanges, routing keys, reply-to) or SQS+SNS; not Kafka's sweet spot. 7. **Operational appetite?** Managed SQS ≈ near-zero ops. RabbitMQ = moderate. Kafka = significant (partitions, retention, rebalances, lag, KRaft/ZooKeeper, schema registry) — though managed offerings (MSK, Confluent Cloud) reduce this. ## Workloads that favor Kafka - Event-driven backbone / event sourcing (the log is the source of truth). - Stream processing and real-time analytics (Kafka Streams, Flink, Spark, ksqlDB). - Log/metrics/CDC pipelines (Debezium → Kafka → sinks via Kafka Connect). - Multi-team data distribution where many consumers read the same stream and may backfill. ## Workloads that favor a queue - Background job / task processing (email, image resize, billing jobs) with competing workers. - Workflows needing priorities, delays, per-message TTL, and targeted retries/DLQs. - RPC / request-reply, point-to-point command dispatch. - Simple, low-volume decoupling where Kafka's operational weight isn't justified. ## Common anti-patterns - **Kafka as a plain task queue:** you reimplement per-message retry/priority/delay on top of offsets — painful. Use a queue. - **Queue when you'll need replay/fan-out:** you later discover you can't backfill a new consumer or reprocess. Migrating is costly. - **Ignoring partition count up front:** under-partitioning caps parallelism; over-partitioning adds overhead and rebalance cost. Plan to peak throughput. ## The hybrid reality Mature systems often run **both**: Kafka as the durable, replayable event spine that fans out to many consumers, and queues (SQS/RabbitMQ) for specific task-processing stages that need per-message delivery semantics. The decision is per workload, not a religion.

  • What goes wrong if you use Kafka as a plain background-job queue?
    You lose built-in per-message features (priority, delayed delivery, per-message retry/DLQ) and must reimplement them over offsets; partition count caps worker parallelism; and head-of-line blocking on one slow message can stall a partition. A queue (RabbitMQ/SQS) handles these natively.
  • Name a concrete signal that pushes you toward Kafka over a queue.
    Multiple independent teams need to consume the same event stream and may need to backfill/replay history — Kafka gives that with one stored copy and per-group offsets; a queue would require per-subscriber copies and can't replay.

saying these in an interview costs you the question

  • Recommending Kafka purely because it's higher throughput, ignoring per-message delivery needs.
  • Treating Kafka and RabbitMQ as interchangeable 'message brokers.'
  • Ignoring operational cost (partitions, retention, lag, KRaft/ZooKeeper) when recommending Kafka.
  • Claiming you can't use both — hybrid architectures are common and often correct.

context