skip to content

How does dead-letter queue routing actually get implemented differently in a broker like Amazon SQS or RabbitMQ versus a log-based system like Apache Kafka, where there is no built-in 'move failed message elsewhere' primitive?

level: middleimportance: should knowfreq 55%

answer

  1. SQS redrive policy (maxReceiveCount)
  2. RabbitMQ dead-letter-exchange (DLX)
  3. Kafka has no native DLQ — app/framework-level DLQ topic
  4. must commit offset past failed Kafka record
  5. Kafka Connect errors.deadletterqueue.topic.name

basics

~20 s

Queue systems like SQS or RabbitMQ can automatically move a failed message to a separate DLQ for you. Kafka has no such built-in feature — the consumer's own code has to catch the failure and manually publish the message to a separate 'DLQ topic' itself.

solid answer

~50 s

SQS and RabbitMQ are consumer-acknowledgment-based queues: the broker tracks per-message delivery/receive counts and, via a redrive policy (SQS) or a dead-letter-exchange binding (RabbitMQ), automatically routes a message to a configured DLQ once it crosses a threshold — the broker does the work. Kafka is fundamentally different: it's an append-only log with consumer-tracked offsets, not a per-message queue, and has no native concept of 'this message failed, move it.' So DLQ behavior in Kafka has to be implemented at the application or framework layer — the consumer catches the processing exception, publishes the raw message (plus failure metadata) to a separate DLQ topic it owns, and only then commits/advances its own offset past the failing record, so the pipeline doesn't stall. Kafka Connect and Spring Kafka provide built-in error-handling configs (e.g., Connect's errors.deadletterqueue.topic.name) that implement this pattern for you.

go deeper

for a junior

Should know that some systems move failed messages to a DLQ automatically while others require writing code to do it, without needing the exact configuration names.

for a middle

Should name at least one broker-native mechanism (SQS redrive policy or RabbitMQ DLX) and know that Kafka requires an application/framework-level DLQ topic instead.

for a senior

Should explain the offset-commit ordering hazard in Kafka's app-level DLQ pattern and be able to name a framework feature (Kafka Connect or Spring Kafka) that implements it.

for a principal

Should reason about the architectural trade-off between broker-native simplicity and app-level flexibility, and design a DLQ strategy that differentiates transient vs permanent failures even within a single Kafka consumer.

## Broker-native routing: the broker does the work In broker-native systems like SQS and RabbitMQ, the DECISION that a message has failed enough times and the ACT of moving it are both broker responsibilities. - **SQS** attaches a **'redrive policy'** to a queue specifying `maxReceiveCount` and the ARN of a DLQ; SQS tracks `ApproximateReceiveCount` server-side per message, and crossing the threshold automatically moves it. - **RabbitMQ** instead uses a **dead-letter-exchange** (an `x-dead-letter-exchange` argument on a queue): a message is dead-lettered when it's explicitly nacked/rejected without requeue, when its TTL expires, or when the queue's max-length is exceeded — RabbitMQ then republishes it to the configured exchange, which is bound to a DLQ. In both cases consumer code doesn't have to build any queue-routing logic itself. ## Kafka has no server-side redrive Kafka is fundamentally different: topics are partitioned, ordered, immutable logs, and consumers don't 'pop' messages — they read at an offset and advance it independently. There's no broker-side concept of per-record state (delivered/failed/dead), so there's no server-side redrive. The application must: 1. Catch the exception around record processing. 2. Produce a copy of the failed record — often wrapped with headers like original topic/partition/offset, exception class, and stack trace — onto a dedicated DLQ topic via a producer. 3. Commit the consumer offset past that record so the main topic's consumption keeps advancing instead of reprocessing the poison record on every poll. Frameworks formalize this: - **Kafka Connect** exposes `errors.tolerance=all` plus `errors.deadletterqueue.topic.name` for sink connectors. - **Spring Kafka**'s `DeadLetterPublishingRecoverer` does the same for `@KafkaListener` consumers. - **Kafka Streams** exposes exception handlers for a similar purpose. ## The trade-off: simplicity versus flexibility The trade-off is simplicity versus flexibility. - **Broker-native DLQ** gives correct-by-default behavior with minimal code, but constrains you to what the broker's redrive policy supports — SQS's threshold, for example, is a single number, not conditional on error type. - **Kafka's app-level approach** is more code and more moving pieces — you must wire it into every consumer, keep offset-commit ordering correct, and avoid infinite loops where a DLQ-producer failure itself needs handling — but it's more flexible: you can inspect the exception type and route only permanent failures to the DLQ topic while relying on Kafka's normal offset-based redelivery (simply not committing the offset) for transient ones, or attach arbitrarily rich metadata as headers. ## The bugs this pattern produces 1. **Forgetting to commit the offset** past the failed record after successfully writing it to the DLQ topic is the most common production bug in Kafka's pattern, and it causes the same record to be reprocessed and re-published to the DLQ topic repeatedly on every consumer restart or rebalance, effectively duplicating it forever. 2. **The inverse bug** — committing the offset before successfully publishing to the DLQ topic — silently loses the record if the DLQ publish then fails. 3. In **RabbitMQ**, a common misunderstanding is what counts as a 'failed delivery': if a consumer crashes without explicitly nacking (e.g., its connection drops), the message is typically requeued rather than dead-lettered, so a crash-looping consumer might never trigger the DLX path the way an explicit reject-without-requeue would. ## A worked example A worked example: a Kafka Connect JDBC sink connector configured with `errors.tolerance=all` and `errors.deadletterqueue.topic.name=dlq-orders`. When a record can't be written to the database (e.g., a schema mismatch), the connector automatically writes the offending record, along with the exception message and stack trace as headers, to `dlq-orders`, and continues processing subsequent records from the source topic rather than stalling the whole connector task — mirroring, at the application/framework layer, exactly what SQS's redrive policy or RabbitMQ's DLX gives you natively.

  • In RabbitMQ, does simply letting a message time out via TTL also trigger dead-lettering?
    Yes — RabbitMQ dead-letters a message for several reasons besides an explicit reject: TTL expiry, the queue hitting a max-length limit, or an explicit nack/reject with requeue=false; all of these route through the queue's configured dead-letter-exchange if one is set, not just repeated processing failures.
  • Why does a Kafka consumer need to be careful about offset-commit ordering when implementing a DLQ topic?
    Because Kafka tracks progress purely via committed offsets, not per-message state, the consumer must publish the failing record to the DLQ topic and only then commit past that offset — committing first risks losing the record if the DLQ publish fails, while never committing causes the same record to be reprocessed and re-published to the DLQ topic on every restart.
  • Could you build DLQ-like behavior in Kafka without a separate DLQ topic at all?
    You could skip past bad records without publishing them anywhere and just log the error, but that loses the message entirely with no way to inspect or replay it later, which defeats the main purpose of a DLQ; a dedicated DLQ topic is what makes the failure recoverable rather than just observable.

Broker-native DLQ (SQS/RabbitMQ) is like a postal sorting facility that automatically pulls undeliverable mail into a separate bin for you; Kafka is more like a shared logbook everyone reads from their own bookmark, so if an entry can't be handled, the reader has to personally copy it into a separate logbook themselves.

saying these in an interview costs you the question

  • thinks Kafka has a built-in DLQ feature identical to SQS/RabbitMQ
  • doesn't mention that Kafka DLQ requires publishing to a separate topic plus advancing the consumer offset
  • unaware that RabbitMQ dead-lettering can be triggered by TTL/queue-length, not only failed processing
  • assumes any consumer crash in RabbitMQ automatically dead-letters the message

context