skip to content

Error Handling & Dead-Letter Recovery

The default error handler retries with a backoff and then publishes to a dead-letter topic, while @RetryableTopic moves retries off the main topic so the partition keeps flowing. Distinguishing recoverable from fatal exceptions is the design point.

part ofSpring Frameworkoverview, primer and where to startread it →
on this pageshow

explore

questions

5

What is Spring Kafka's DefaultErrorHandler and how does it retry a failed record?

level: juniorimportance: must knowfreq 70%

answer

  1. seek back to failed offset = blocking retry
  2. FixedBackOff(0,9) = 10 deliveries
  3. pauses partitions during back-off (no rebalance)
  4. exhausted → recoverer (default logs/skips)
  5. replaced SeekToCurrentErrorHandler in 2.8

basics

~10 s

DefaultErrorHandler is the container's default handler when a @KafkaListener throws. It re-delivers the failed record several times (using a BackOff for delays), and if all attempts fail it hands the record to a recoverer.

solid answer

~40 s

DefaultErrorHandler is the error handler the Spring Kafka message listener container uses when your @KafkaListener throws an exception. Instead of losing or committing the bad record, it uses the consumer's seek to reposition the partition back to the failed offset so the same record is polled and re-delivered. A configurable BackOff controls how many attempts and how long between them (default FixedBackOff(0, 9) = 10 total deliveries, no delay). During the back-off the container pauses the partitions and keeps polling so the broker doesn't trigger a rebalance. When retries are exhausted it calls a recoverer (default just logs and moves on); you usually plug in a DeadLetterPublishingRecoverer to send the record to a dead-letter topic. It replaced SeekToCurrentErrorHandler in Spring Kafka 2.8.

code

java · 22 lines
java
@Bean
public DefaultErrorHandler errorHandler(KafkaTemplate<Object, Object> template) {
    // Republish to <topic>.DLT after retries are exhausted
    var recoverer = new DeadLetterPublishingRecoverer(template);
    // 3 retries with exponential back-off (1s, 2s, 4s)
    var backOff = new ExponentialBackOffWithMaxRetries(3);
    backOff.setInitialInterval(1_000L);
    backOff.setMultiplier(2.0);
    backOff.setMaxInterval(10_000L);
    var handler = new DefaultErrorHandler(recoverer, backOff);
    handler.addNotRetryableExceptions(IllegalArgumentException.class);
    return handler;
}

@Bean
public ConcurrentKafkaListenerContainerFactory<String, String> kafkaListenerContainerFactory(
        ConsumerFactory<String, String> cf, DefaultErrorHandler errorHandler) {
    var factory = new ConcurrentKafkaListenerContainerFactory<String, String>();
    factory.setConsumerFactory(cf);
    factory.setCommonErrorHandler(errorHandler);
    return factory;
}

go deeper

for a junior

Know it retries the failed record and can send it to a dead-letter topic after giving up.

for a middle

Explain seek-based redelivery, the BackOff (attempts + delay), and wiring a DeadLetterPublishingRecoverer.

for a senior

Discuss blocking semantics, partition pausing to avoid rebalance, fatal-exception classification, and default-recoverer data loss.

for a principal

Frame the ordering-vs-throughput trade-off of blocking seek retries versus non-blocking retry topics and when each fits an SLA.

**The problem.** A `@KafkaListener` method processes one record (or a batch) per poll. If it throws, the container must decide: skip the record, retry it, or stop. Kafka has no per-message ack/redelivery like a broker queue does — the consumer just tracks an *offset*. So 'retry' means: don't advance the committed offset past the bad record, and poll it again. **What DefaultErrorHandler is.** It is the default `CommonErrorHandler` used by the `MessageListenerContainer` since Spring Kafka 2.8 (it unified and replaced `SeekToCurrentErrorHandler` for record listeners and `RecoveringBatchErrorHandler`/`SeekToCurrentBatchErrorHandler` for batch listeners). You rarely see it explicitly unless you customize it — it is wired automatically. **Seek-based redelivery.** When the listener throws, DefaultErrorHandler calls `consumer.seek(partition, failedOffset)` — it rewinds the consumer's *position* (in-memory, not the committed offset) so the very next `poll()` returns the same record again. This is the 'seek-based' mechanism: no separate retry topic, no in-memory queue of records — the source of truth stays the log itself. This is a **blocking retry**: the consumer thread is busy redelivering this one record and does not make progress on later records in that partition until it succeeds or is recovered. Ordering within the partition is therefore preserved. **BackOff.** How many times and how fast to retry is controlled by a `org.springframework.util.backoff.BackOff`: - `FixedBackOff(interval, maxAttempts)` — default is `FixedBackOff(0L, 9L)`, i.e. no delay and 9 *retries* → **10 total deliveries**. - `ExponentialBackOff` / `ExponentialBackOffWithMaxRetries` — growing delays (e.g. 1s, 2s, 4s… capped), bounded attempts. You pass the BackOff to the constructor: `new DefaultErrorHandler(recoverer, backOff)`. **Avoiding rebalance during back-off.** If a back-off interval is long, the consumer might exceed `max.poll.interval.ms` and get kicked out of the group, causing a rebalance. To prevent this, DefaultErrorHandler (2.8+) **pauses** the assigned partitions and keeps calling `poll()` (which returns nothing while paused) so the consumer stays alive and heartbeats, then resumes when the back-off elapses. **The recoverer.** When attempts are exhausted, DefaultErrorHandler invokes a `ConsumerRecordRecoverer`. The default recoverer just logs the record and lets the container commit past it (so the poison record is skipped). To keep the record, supply a `DeadLetterPublishingRecoverer` which republishes it to a `.DLT` topic. **Fatal exceptions.** Some exceptions can never succeed on retry (e.g. `DeserializationException`, `MessageConversionException`, `ClassCastException`). DefaultErrorHandler treats these as **non-retryable/fatal** by default and recovers them immediately without wasting retries. You can tune the classification with `addRetryableExceptions(...)` / `addNotRetryableExceptions(...)`. **When to use.** DefaultErrorHandler + DeadLetterPublishingRecoverer is the standard, simplest resilient setup for a record listener when you need **ordering** and can tolerate the partition being blocked while one record retries. If blocking the partition is unacceptable, use non-blocking retries via `@RetryableTopic` instead. **Gotchas.** (1) Zero-delay default means 10 near-instant retries — often you want a real BackOff. (2) Long fixed delays block the whole partition; a stuck record starves everything behind it. (3) The default recoverer *skips* (commits past) the record — data loss unless you add a DLT recoverer. (4) Manual `AckMode.MANUAL_IMMEDIATE` and error handling interact; DefaultErrorHandler handles the seek/commit for you in the common auto-ack modes.

  • What is the default BackOff and why is it often a bad default in production?
    FixedBackOff(0L, 9L): 10 deliveries with zero delay. It hammers the same failing record 10 times in a burst, giving transient issues (e.g. a downstream blip) no time to recover, and for a permanent failure it just spins. Production usually wants an ExponentialBackOff with real intervals.
  • How does DefaultErrorHandler avoid a consumer-group rebalance when a back-off interval is long?
    It pauses the assigned partitions and keeps polling (paused polls return nothing) so the consumer keeps heartbeating and stays under max.poll.interval.ms, then resumes when the delay elapses.

saying these in an interview costs you the question

  • Thinking DefaultErrorHandler uses a separate retry topic (that's @RetryableTopic; DefaultErrorHandler seeks in place)
  • Believing the default recoverer preserves the record (it logs and skips, committing past it — data loss)
  • Saying retries don't block the partition (blocking retry blocks progress within the partition)
  • Assuming there's a delay between default retries (default interval is 0)

context

open as a page

What does DeadLetterPublishingRecoverer do, and what does the target topic and message look like?

level: middleimportance: must knowfreq 65%

basics

~20 s

It's a recoverer that, after retries are exhausted, republishes the failed record to a dead-letter topic (default: original topic name + ".DLT") using a KafkaTemplate, adding headers describing the failure so the message isn't lost.

open as a page

What is @RetryableTopic non-blocking retry, and how does it differ from DefaultErrorHandler's blocking seek retries?

level: seniorimportance: must knowfreq 60%

basics

~20 s

@RetryableTopic makes retries non-blocking: a failed record is forwarded to separate retry topics (with delays), so the main consumer keeps processing new records instead of being stuck. After the configured attempts, it goes to a DLT. DefaultErrorHandler instead retries in place by seeking, which blocks the partition.

open as a page

How does DefaultErrorHandler distinguish recoverable from fatal (non-retryable) exceptions, and how do you customize the classification?

level: seniorimportance: should knowfreq 50%

basics

~20 s

It uses an exception classifier. Some exceptions (like deserialization or conversion errors) are fatal by default — retrying can't help — so they skip retries and go straight to the recoverer. You adjust the lists with addRetryableExceptions / addNotRetryableExceptions.

open as a page

How would you choose and design a Kafka error/retry/DLT strategy for a service with mixed ordering and throughput requirements?

level: principalimportance: should knowfreq 40%

basics

~20 s

Match the mechanism to the constraint: use blocking seek retries (DefaultErrorHandler) where per-partition ordering matters and outages are short; use non-blocking @RetryableTopic where throughput and long back-offs matter and ordering can be relaxed. Always classify transient vs fatal, and route the unrecoverable to a DLT with replay tooling.

open as a page