What is DefaultErrorHandler in Spring for Apache Kafka, and what does it do when a listener throws an exception?
answer
- Default = FixedBackOff(0,9) = 10 attempts
- Blocking, in-memory, seek-back retries
- Recoverer runs when retries exhausted
- Replaced SeekToCurrentErrorHandler in 2.8
- Fatal exceptions skip straight to recovery
basics
~20 sDefaultErrorHandler is the standard component that catches exceptions thrown by your @KafkaListener. It retries the record a configurable number of times with a backoff delay, and if all retries fail it calls a recoverer (by default just logs the failed record).
solid answer
~30 sDefaultErrorHandler is the default container-level error handler in Spring Kafka (since 2.8, replacing SeekToCurrentErrorHandler). When a @KafkaListener throws, the container hands the exception to it. It performs in-memory, blocking retries by re-seeking the consumer back to the failed offset and re-polling, using a BackOff (default FixedBackOff of 9 retries with no delay = 10 attempts). When retries are exhausted it invokes a ConsumerRecordRecoverer — by default a logging recoverer, but commonly a DeadLetterPublishingRecoverer to send the record to a DLT. It also distinguishes 'fatal' (non-retryable) exceptions, which skip retries and go straight to recovery, from retryable ones.
go deeper
Know it catches listener exceptions, retries with a backoff, then recovers (logs) the record so it doesn't block the partition.
Explain the seek-based blocking retry mechanism, the FixedBackOff(0,9) default, and wiring a DeadLetterPublishingRecoverer as recoverer.
Discuss fatal vs retryable classification, max.poll.interval.ms risk from long backoffs, and when blocking retries are the wrong tool.
Reason about delivery guarantees, recoverer idempotency/failure handling, and choosing blocking vs non-blocking retry strategies across a fleet of consumers.
## What problem it solves When you consume from Kafka with a `@KafkaListener` method and that method throws an exception, something has to decide what happens next: retry? skip? stop? In Spring for Apache Kafka, that decision is made by a **CommonErrorHandler**, and the default implementation is **`DefaultErrorHandler`** (introduced in Spring Kafka 2.8, superseding the older `SeekToCurrentErrorHandler` and `RecoveringBatchErrorHandler`). ## How it works mechanically Kafka consumers don't have per-message acknowledgement the way some brokers do — a consumer polls a **batch** of records and tracks a single **offset** (position) per partition. So to 'retry one record', Spring uses **seek**: it tells the consumer to reposition (`seek`) back to the offset of the failed record, so the next `poll()` returns that same record again. This is why these retries are called **blocking** and **in-memory** — the same consumer thread re-delivers the record, and the partition makes no forward progress until the record either succeeds or is recovered. The number of attempts and the delay between them are controlled by a **`BackOff`** (from Spring Core). The default is `new FixedBackOff(0L, 9L)` — 9 retries with 0ms delay, i.e. **10 total attempts**. You can supply a `FixedBackOff(interval, maxAttempts)` for constant spacing or an `ExponentialBackOff`/`ExponentialBackOffWithMaxRetries` for growing delays. ## What happens after retries are exhausted When the `BackOff` says 'stop', `DefaultErrorHandler` calls its **recoverer**, a `ConsumerRecordRecoverer`. By default this just **logs** the failed record and the consumer commits past it (so the poison record doesn't block the partition forever). The most common production choice is to inject a **`DeadLetterPublishingRecoverer`**, which publishes the failed record to a **dead-letter topic (DLT)** so it can be inspected or reprocessed later. ## Fatal vs retryable exceptions Not every exception should be retried. A deserialization failure or a validation error will never succeed on retry. `DefaultErrorHandler` keeps a list of **'fatal' (not-retryable) exception types** (e.g. `DeserializationException`, `MessageConversionException`, `ConversionException`, `MethodArgumentResolutionException`, `NoSuchMethodException`, `ClassCastException`). For those it **skips retries** and goes straight to recovery. You tune this via `addNotRetryableExceptions(...)` / `addRetryableExceptions(...)` or `setClassifications(...)`, and `defaultFalse()` to make the classifier deny-by-default. ## Edge cases - Retries are **blocking**, so a long backoff can exceed `max.poll.interval.ms` and trigger a consumer **rebalance**; Spring mitigates this by pausing the consumer between deliveries within the same poll where possible. - If the recoverer itself fails (e.g. DLT publish fails), the offset is **not** committed and the record is retried again later — recovery is meant to be reliable. - For batch listeners there's special handling (`BatchListenerFailedException` to indicate which record in the batch failed).
- How do you change the number of retries and the delay?Construct DefaultErrorHandler with a BackOff: e.g. new FixedBackOff(1000L, 3) for 3 retries spaced 1s apart, or an ExponentialBackOffWithMaxRetries for growing delays. Pass it to the handler constructor (optionally with a recoverer).
- Why are these retries called 'blocking'?Because the same consumer thread re-seeks and re-polls the failed record, so that partition makes no forward progress and processes no other records until the record succeeds or is recovered. The thread is effectively blocked on that record.
saying these in an interview costs you the question
- Saying retries happen on a separate thread or via separate topics — DefaultErrorHandler retries are blocking/in-memory by re-seeking the same consumer.
- Claiming the default is infinite retries — it's 9 retries (10 attempts) by default.
- Confusing it with @RetryableTopic (non-blocking) — DefaultErrorHandler is blocking.
- Saying the default recoverer sends to a DLT — by default it only logs; you must wire a DeadLetterPublishingRecoverer.