What is Spring Kafka's DefaultErrorHandler and how does it retry a failed record?
answer
- seek back to failed offset = blocking retry
- FixedBackOff(0,9) = 10 deliveries
- pauses partitions during back-off (no rebalance)
- exhausted → recoverer (default logs/skips)
- replaced SeekToCurrentErrorHandler in 2.8
basics
~10 sDefaultErrorHandler is the container's default handler when a @KafkaListener throws. It re-delivers the failed record several times (using a BackOff for delays), and if all attempts fail it hands the record to a recoverer.
solid answer
~40 sDefaultErrorHandler is the error handler the Spring Kafka message listener container uses when your @KafkaListener throws an exception. Instead of losing or committing the bad record, it uses the consumer's seek to reposition the partition back to the failed offset so the same record is polled and re-delivered. A configurable BackOff controls how many attempts and how long between them (default FixedBackOff(0, 9) = 10 total deliveries, no delay). During the back-off the container pauses the partitions and keeps polling so the broker doesn't trigger a rebalance. When retries are exhausted it calls a recoverer (default just logs and moves on); you usually plug in a DeadLetterPublishingRecoverer to send the record to a dead-letter topic. It replaced SeekToCurrentErrorHandler in Spring Kafka 2.8.
code
java · 22 lines@Bean
public DefaultErrorHandler errorHandler(KafkaTemplate<Object, Object> template) {
// Republish to <topic>.DLT after retries are exhausted
var recoverer = new DeadLetterPublishingRecoverer(template);
// 3 retries with exponential back-off (1s, 2s, 4s)
var backOff = new ExponentialBackOffWithMaxRetries(3);
backOff.setInitialInterval(1_000L);
backOff.setMultiplier(2.0);
backOff.setMaxInterval(10_000L);
var handler = new DefaultErrorHandler(recoverer, backOff);
handler.addNotRetryableExceptions(IllegalArgumentException.class);
return handler;
}
@Bean
public ConcurrentKafkaListenerContainerFactory<String, String> kafkaListenerContainerFactory(
ConsumerFactory<String, String> cf, DefaultErrorHandler errorHandler) {
var factory = new ConcurrentKafkaListenerContainerFactory<String, String>();
factory.setConsumerFactory(cf);
factory.setCommonErrorHandler(errorHandler);
return factory;
}go deeper
Know it retries the failed record and can send it to a dead-letter topic after giving up.
Explain seek-based redelivery, the BackOff (attempts + delay), and wiring a DeadLetterPublishingRecoverer.
Discuss blocking semantics, partition pausing to avoid rebalance, fatal-exception classification, and default-recoverer data loss.
Frame the ordering-vs-throughput trade-off of blocking seek retries versus non-blocking retry topics and when each fits an SLA.
**The problem.** A `@KafkaListener` method processes one record (or a batch) per poll. If it throws, the container must decide: skip the record, retry it, or stop. Kafka has no per-message ack/redelivery like a broker queue does — the consumer just tracks an *offset*. So 'retry' means: don't advance the committed offset past the bad record, and poll it again. **What DefaultErrorHandler is.** It is the default `CommonErrorHandler` used by the `MessageListenerContainer` since Spring Kafka 2.8 (it unified and replaced `SeekToCurrentErrorHandler` for record listeners and `RecoveringBatchErrorHandler`/`SeekToCurrentBatchErrorHandler` for batch listeners). You rarely see it explicitly unless you customize it — it is wired automatically. **Seek-based redelivery.** When the listener throws, DefaultErrorHandler calls `consumer.seek(partition, failedOffset)` — it rewinds the consumer's *position* (in-memory, not the committed offset) so the very next `poll()` returns the same record again. This is the 'seek-based' mechanism: no separate retry topic, no in-memory queue of records — the source of truth stays the log itself. This is a **blocking retry**: the consumer thread is busy redelivering this one record and does not make progress on later records in that partition until it succeeds or is recovered. Ordering within the partition is therefore preserved. **BackOff.** How many times and how fast to retry is controlled by a `org.springframework.util.backoff.BackOff`: - `FixedBackOff(interval, maxAttempts)` — default is `FixedBackOff(0L, 9L)`, i.e. no delay and 9 *retries* → **10 total deliveries**. - `ExponentialBackOff` / `ExponentialBackOffWithMaxRetries` — growing delays (e.g. 1s, 2s, 4s… capped), bounded attempts. You pass the BackOff to the constructor: `new DefaultErrorHandler(recoverer, backOff)`. **Avoiding rebalance during back-off.** If a back-off interval is long, the consumer might exceed `max.poll.interval.ms` and get kicked out of the group, causing a rebalance. To prevent this, DefaultErrorHandler (2.8+) **pauses** the assigned partitions and keeps calling `poll()` (which returns nothing while paused) so the consumer stays alive and heartbeats, then resumes when the back-off elapses. **The recoverer.** When attempts are exhausted, DefaultErrorHandler invokes a `ConsumerRecordRecoverer`. The default recoverer just logs the record and lets the container commit past it (so the poison record is skipped). To keep the record, supply a `DeadLetterPublishingRecoverer` which republishes it to a `.DLT` topic. **Fatal exceptions.** Some exceptions can never succeed on retry (e.g. `DeserializationException`, `MessageConversionException`, `ClassCastException`). DefaultErrorHandler treats these as **non-retryable/fatal** by default and recovers them immediately without wasting retries. You can tune the classification with `addRetryableExceptions(...)` / `addNotRetryableExceptions(...)`. **When to use.** DefaultErrorHandler + DeadLetterPublishingRecoverer is the standard, simplest resilient setup for a record listener when you need **ordering** and can tolerate the partition being blocked while one record retries. If blocking the partition is unacceptable, use non-blocking retries via `@RetryableTopic` instead. **Gotchas.** (1) Zero-delay default means 10 near-instant retries — often you want a real BackOff. (2) Long fixed delays block the whole partition; a stuck record starves everything behind it. (3) The default recoverer *skips* (commits past) the record — data loss unless you add a DLT recoverer. (4) Manual `AckMode.MANUAL_IMMEDIATE` and error handling interact; DefaultErrorHandler handles the seek/commit for you in the common auto-ack modes.
- What is the default BackOff and why is it often a bad default in production?FixedBackOff(0L, 9L): 10 deliveries with zero delay. It hammers the same failing record 10 times in a burst, giving transient issues (e.g. a downstream blip) no time to recover, and for a permanent failure it just spins. Production usually wants an ExponentialBackOff with real intervals.
- How does DefaultErrorHandler avoid a consumer-group rebalance when a back-off interval is long?It pauses the assigned partitions and keeps polling (paused polls return nothing) so the consumer keeps heartbeating and stays under max.poll.interval.ms, then resumes when the delay elapses.
saying these in an interview costs you the question
- Thinking DefaultErrorHandler uses a separate retry topic (that's @RetryableTopic; DefaultErrorHandler seeks in place)
- Believing the default recoverer preserves the record (it logs and skips, committing past it — data loss)
- Saying retries don't block the partition (blocking retry blocks progress within the partition)
- Assuming there's a delay between default retries (default interval is 0)