skip to content

Error Handling, Retry Topics and DLT

Retry strategies in Spring Kafka: blocking backoff versus non-blocking retry topics, and publishing to a dead-letter topic. Interviewers ask because blocking retries stall the whole partition behind one bad record.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

6

What is DefaultErrorHandler in Spring for Apache Kafka, and what does it do when a listener throws an exception?

level: juniorimportance: must knowfreq 70%

answer

  1. Default = FixedBackOff(0,9) = 10 attempts
  2. Blocking, in-memory, seek-back retries
  3. Recoverer runs when retries exhausted
  4. Replaced SeekToCurrentErrorHandler in 2.8
  5. Fatal exceptions skip straight to recovery

basics

~20 s

DefaultErrorHandler is the standard component that catches exceptions thrown by your @KafkaListener. It retries the record a configurable number of times with a backoff delay, and if all retries fail it calls a recoverer (by default just logs the failed record).

solid answer

~30 s

DefaultErrorHandler is the default container-level error handler in Spring Kafka (since 2.8, replacing SeekToCurrentErrorHandler). When a @KafkaListener throws, the container hands the exception to it. It performs in-memory, blocking retries by re-seeking the consumer back to the failed offset and re-polling, using a BackOff (default FixedBackOff of 9 retries with no delay = 10 attempts). When retries are exhausted it invokes a ConsumerRecordRecoverer — by default a logging recoverer, but commonly a DeadLetterPublishingRecoverer to send the record to a DLT. It also distinguishes 'fatal' (non-retryable) exceptions, which skip retries and go straight to recovery, from retryable ones.

go deeper

for a junior

Know it catches listener exceptions, retries with a backoff, then recovers (logs) the record so it doesn't block the partition.

for a middle

Explain the seek-based blocking retry mechanism, the FixedBackOff(0,9) default, and wiring a DeadLetterPublishingRecoverer as recoverer.

for a senior

Discuss fatal vs retryable classification, max.poll.interval.ms risk from long backoffs, and when blocking retries are the wrong tool.

for a principal

Reason about delivery guarantees, recoverer idempotency/failure handling, and choosing blocking vs non-blocking retry strategies across a fleet of consumers.

## What problem it solves When you consume from Kafka with a `@KafkaListener` method and that method throws an exception, something has to decide what happens next: retry? skip? stop? In Spring for Apache Kafka, that decision is made by a **CommonErrorHandler**, and the default implementation is **`DefaultErrorHandler`** (introduced in Spring Kafka 2.8, superseding the older `SeekToCurrentErrorHandler` and `RecoveringBatchErrorHandler`). ## How it works mechanically Kafka consumers don't have per-message acknowledgement the way some brokers do — a consumer polls a **batch** of records and tracks a single **offset** (position) per partition. So to 'retry one record', Spring uses **seek**: it tells the consumer to reposition (`seek`) back to the offset of the failed record, so the next `poll()` returns that same record again. This is why these retries are called **blocking** and **in-memory** — the same consumer thread re-delivers the record, and the partition makes no forward progress until the record either succeeds or is recovered. The number of attempts and the delay between them are controlled by a **`BackOff`** (from Spring Core). The default is `new FixedBackOff(0L, 9L)` — 9 retries with 0ms delay, i.e. **10 total attempts**. You can supply a `FixedBackOff(interval, maxAttempts)` for constant spacing or an `ExponentialBackOff`/`ExponentialBackOffWithMaxRetries` for growing delays. ## What happens after retries are exhausted When the `BackOff` says 'stop', `DefaultErrorHandler` calls its **recoverer**, a `ConsumerRecordRecoverer`. By default this just **logs** the failed record and the consumer commits past it (so the poison record doesn't block the partition forever). The most common production choice is to inject a **`DeadLetterPublishingRecoverer`**, which publishes the failed record to a **dead-letter topic (DLT)** so it can be inspected or reprocessed later. ## Fatal vs retryable exceptions Not every exception should be retried. A deserialization failure or a validation error will never succeed on retry. `DefaultErrorHandler` keeps a list of **'fatal' (not-retryable) exception types** (e.g. `DeserializationException`, `MessageConversionException`, `ConversionException`, `MethodArgumentResolutionException`, `NoSuchMethodException`, `ClassCastException`). For those it **skips retries** and goes straight to recovery. You tune this via `addNotRetryableExceptions(...)` / `addRetryableExceptions(...)` or `setClassifications(...)`, and `defaultFalse()` to make the classifier deny-by-default. ## Edge cases - Retries are **blocking**, so a long backoff can exceed `max.poll.interval.ms` and trigger a consumer **rebalance**; Spring mitigates this by pausing the consumer between deliveries within the same poll where possible. - If the recoverer itself fails (e.g. DLT publish fails), the offset is **not** committed and the record is retried again later — recovery is meant to be reliable. - For batch listeners there's special handling (`BatchListenerFailedException` to indicate which record in the batch failed).

  • How do you change the number of retries and the delay?
    Construct DefaultErrorHandler with a BackOff: e.g. new FixedBackOff(1000L, 3) for 3 retries spaced 1s apart, or an ExponentialBackOffWithMaxRetries for growing delays. Pass it to the handler constructor (optionally with a recoverer).
  • Why are these retries called 'blocking'?
    Because the same consumer thread re-seeks and re-polls the failed record, so that partition makes no forward progress and processes no other records until the record succeeds or is recovered. The thread is effectively blocked on that record.

saying these in an interview costs you the question

  • Saying retries happen on a separate thread or via separate topics — DefaultErrorHandler retries are blocking/in-memory by re-seeking the same consumer.
  • Claiming the default is infinite retries — it's 9 retries (10 attempts) by default.
  • Confusing it with @RetryableTopic (non-blocking) — DefaultErrorHandler is blocking.
  • Saying the default recoverer sends to a DLT — by default it only logs; you must wire a DeadLetterPublishingRecoverer.

context

open as a page

A non-deserializable 'poison-pill' record keeps crashing your consumer in an infinite loop. Why does this happen with naive handling, and how do you fix it?

level: middleimportance: must knowfreq 55%

basics

~20 s

A poison pill is a record that can't be deserialized, so it throws before your listener even runs. The consumer keeps re-reading the same offset and looping forever. Fix it with ErrorHandlingDeserializer, which catches the failure and lets the error handler skip the record to a DLT.

open as a page

Contrast blocking retries (DefaultErrorHandler) with non-blocking retries (@RetryableTopic). When would you choose one over the other?

level: seniorimportance: must knowfreq 65%

basics

~20 s

Blocking retries re-process the failed record on the same consumer thread, holding up the partition until it succeeds. Non-blocking retries (@RetryableTopic) forward the record to separate retry topics with delays, so the main topic keeps flowing while failed records are retried independently.

open as a page

Compare FixedBackOff and ExponentialBackOff for retry spacing. How do you cap attempts with ExponentialBackOff, and why add jitter?

level: middleimportance: should knowfreq 40%

basics

~20 s

FixedBackOff retries at a constant interval for a set number of attempts. ExponentialBackOff grows the delay each time (×multiplier up to a max). To bound retries with exponential delays, use ExponentialBackOffWithMaxRetries. Jitter (randomness) spreads retries out to avoid all consumers hammering a recovering dependency at once.

open as a page

Walk through configuring @RetryableTopic: attempts, backoff topics, exception classification, and where the DLT fits.

level: middleimportance: should knowfreq 50%

basics

~20 s

@RetryableTopic on a listener tells Spring to auto-create retry topics with increasing delays and a final DLT. You set attempts, the backoff (delay/multiplier), which exceptions to include/exclude, and a @DltHandler method to process records that exhausted all retries.

open as a page

Explain the difference between a 'record-level' fatal exception and a container-fatal exception in Spring Kafka error handling. How does each affect the consumer?

level: seniorimportance: should knowfreq 30%

basics

~20 s

A record-level fatal exception means 'don't retry THIS record' — it's classified non-retryable so the error handler skips straight to recovery (DLT) and the consumer keeps going. A container-fatal exception is more severe: it can stop the whole listener container, halting consumption for all partitions it owns.

open as a page