Compare the ordering and retention guarantees of Kafka versus a traditional queue, and explain the trade-offs each model forces.
answer
- order only within a partition / per key
- retention decoupled from consumption = replay
- slow consumer falls off the log
- competing consumers break queue order
- ack-delete = backlog-bounded, no replay
- idempotence to keep producer order on retry
basics
~20 sKafka guarantees order only within a partition and keeps data for a retention window, enabling replay. Queues with competing consumers generally don't preserve order and delete on ack, so there's no replay but messages can't pile up indefinitely.
solid answer
~60 sKafka orders records strictly *within a partition*; across partitions there's no global order. To keep related events ordered you partition by key (same key → same partition), which trades global ordering for parallelism. Retention is time/size-based and decoupled from consumption: data lives for `retention.ms` whether or not anyone read it, which is what makes replay possible — but a slow consumer can fall off the back of the log and lose data. Traditional queues (RabbitMQ, SQS standard) deliver to competing consumers and generally give up strict ordering once you parallelize; SQS FIFO and RabbitMQ single-consumer/consistent-hash setups recover ordering at a throughput cost. Queues delete on ack, so storage is bounded by the backlog, not a clock, and there's no historical replay; unprocessed messages either sit in the queue or land in a DLQ. The core trade-off: Kafka buys replay + high-throughput ordered-per-key streams at the cost of partition-count rigidity and retention management; queues buy simple per-message delivery + bounded storage at the cost of no replay and weaker ordering under parallelism.
go deeper
Know order is per partition and data has a retention window in Kafka.
Explain keyed partitioning and that retention is independent of consumption (enables replay).
Reason about lag-vs-retention data loss, idempotence for retry ordering, and queue ordering trade-offs (FIFO/competing consumers).
Drive partition-count and retention sizing as capacity decisions; choose Kafka vs queue per workload's ordering/replay/backpressure needs.
## Ordering **Kafka:** A partition is a totally ordered log. Kafka guarantees order **only within a single partition**. A topic with P partitions has P independent ordered streams; there is **no cross-partition ordering**. The standard technique is **keyed partitioning**: the producer hashes the record key to choose a partition (`partition = hash(key) % numPartitions` with the default partitioner), so all records with the same key (e.g., same `customerId`) land in the same partition and are processed in order. This means ordering is *per key*, and parallelism is bounded by partition count. Caveats: changing the partition count rehashes keys (breaks the same-key→same-partition invariant for future records); and to truly preserve producer order on retries you need `enable.idempotence=true` (and historically `max.in.flight.requests.per.connection<=5` with idempotence, or `=1` without) so a retried batch can't be reordered. **Traditional queue:** A single queue is FIFO in principle, but the moment you add **competing consumers** (multiple workers pulling from one queue for throughput), strict end-to-end order is lost because workers process at different speeds and redeliveries reshuffle. To get ordering you either use one consumer (no parallelism) or a partitioned scheme: **AWS SQS FIFO queues** preserve order *per MessageGroupId* (analogous to Kafka's per-key ordering) but cap throughput; **RabbitMQ** can use a consistent-hash exchange or single active consumer for per-key/serial ordering. So both worlds ultimately trade global order for parallelism — Kafka just makes the partition the explicit, durable unit. ## Retention **Kafka:** Retention is **decoupled from consumption**. Records persist for `retention.ms` (default 7 days) or until `retention.bytes` is exceeded, regardless of whether consumers have read them. Compacted topics (`cleanup.policy=compact`) instead keep the latest value per key indefinitely (good for changelog/state). Consequences: - **Replay** is possible — seek to an earlier offset or timestamp and reprocess. - **Data loss risk for slow consumers:** if a consumer lags beyond retention, the oldest unread records are deleted before it reads them. Monitor **consumer lag** (e.g., via `kafka-consumer-groups.sh --describe`, Burrow, or lag metrics). - **Storage scales with retention × throughput**, not backlog — you pay to keep history even if everyone's caught up. **Traditional queue:** Storage holds the **backlog of unacked messages**. Once acked, gone. Messages can have a **TTL**; expired or repeatedly-failed messages go to a **dead-letter queue (DLQ)**. There's no replay of acked history. SQS has a max retention (up to 14 days) for *undelivered* messages, but that's a backlog cap, not a replay log. So storage is naturally bounded by how far consumers fall behind, and you reason about backpressure and DLQs rather than retention windows. ## Trade-off matrix | | Kafka | Queue (RabbitMQ/SQS) | |---|---|---| | Ordering | Per partition (per key) | FIFO degrades with competing consumers; per-group via FIFO/hashing | | Retention | Time/size, replayable, consumption-independent | Until ack/TTL; backlog-bounded; no replay | | Risk | Slow consumer drops off the log | Backlog growth / DLQ overflow | | Parallelism unit | Partition (fixed-ish) | Consumers on a queue (elastic) | | Reprocessing | Seek + replay | Re-enqueue / DLQ redrive | ## When each wins - **Kafka:** event sourcing, stream processing, multi-consumer analytics, anything needing replay or ordered-per-key high throughput. - **Queue:** task/work dispatch with per-message acks, elastic worker pools, rich retry/DLQ semantics, and no need to keep or replay history.
- A consumer is lagging badly. Why is that more dangerous in Kafka than in a standard queue?Kafka retention is time/size-based and independent of consumption, so if lag exceeds retention the oldest unread records are deleted before the consumer reads them — silent data loss. A queue just grows its backlog (until limits/TTL) without deleting unread messages.
- How do you preserve per-customer ordering across millions of customers while still parallelizing?Partition by customerId (Kafka) or MessageGroupId (SQS FIFO). Same key → same partition/group → ordered, while different keys spread across partitions for parallelism.
saying these in an interview costs you the question
- Claiming Kafka guarantees global ordering across a topic (it's only per partition).
- Saying Kafka keeps data forever by default (default retention.ms is 7 days; only compaction keeps latest-per-key).
- Asserting a queue with many competing consumers preserves strict FIFO order.
- Forgetting that lag beyond retention causes data loss in Kafka.