skip to content

How does DefaultErrorHandler distinguish recoverable from fatal (non-retryable) exceptions, and how do you customize the classification?

level: seniorimportance: should knowfreq 50%

answer

  1. fatal by default: Deserialization/MessageConversion/ClassCast/MethodArgumentResolution
  2. addRetryableExceptions / addNotRetryableExceptions
  3. setClassifications(map, defaultFalse) = whitelist retries
  4. fatal → skip BackOff, recover now
  5. classifier traverses cause chain

basics

~20 s

It uses an exception classifier. Some exceptions (like deserialization or conversion errors) are fatal by default — retrying can't help — so they skip retries and go straight to the recoverer. You adjust the lists with addRetryableExceptions / addNotRetryableExceptions.

solid answer

~40 s

DefaultErrorHandler holds a `BinaryExceptionClassifier` that decides, per thrown exception, whether to retry (recoverable) or recover immediately (fatal). By default a fixed set of exceptions is non-retryable because retrying is pointless: `DeserializationException`, `MessageConversionException`, `ConversionException`, `MethodArgumentResolutionException`, `NoSuchMethodException`, and `ClassCastException`. For those it skips the BackOff and calls the recoverer (e.g. DLT) right away. You tune it with `addRetryableExceptions(...)` to force-retry something, or `addNotRetryableExceptions(...)` to mark something fatal. The classifier traverses the cause chain by default. You can also flip the default to 'retry nothing unless listed' via `setClassifications(map, defaultRetryable=false)`. This matters because retrying a permanently-broken record (bad payload, programming bug) just wastes 10 attempts and delays every record behind it.

code

java · 18 lines
java
var handler = new DefaultErrorHandler(recoverer, new FixedBackOff(2000L, 3));

// Deterministic bad input: don't waste retries, send to DLT immediately
handler.addNotRetryableExceptions(
    IllegalArgumentException.class,
    ValidationException.class);

// A normally-fatal type we actually want to retry (e.g. transient remote parse)
handler.addRetryableExceptions(RemoteTimeoutException.class);

// Different back-off depending on the exception
handler.setBackOffFunction((record, ex) ->
    ex instanceof RateLimitedException
        ? new FixedBackOff(30_000L, 5)   // back off hard on rate limits
        : new FixedBackOff(1_000L, 3));

handler.setRetryListeners((record, ex, deliveryAttempt) ->
    meterRegistry.counter("kafka.retry", "topic", record.topic()).increment());

go deeper

for a junior

Know some errors (like bad/unparseable messages) shouldn't be retried.

for a middle

Name the default fatal exceptions and the add(Not)RetryableExceptions methods.

for a senior

Explain the classifier, cause-chain traversal, per-exception BackOff, and the whitelist-inversion pattern.

for a principal

Design a domain error taxonomy (transient vs deterministic), map it to retry/DLT routing, and instrument it with RetryListeners/metrics.

**Why classify at all.** Not every failure is worth retrying. A transient failure (downstream DB timeout, a 503 from a service, an optimistic-lock clash) will likely succeed if tried again. A *deterministic* failure (payload can't be deserialized, a `ClassCastException`, a validation `IllegalArgumentException`, a null-pointer bug) will fail identically every time — retrying wastes attempts, and because retries are **blocking**, it also delays every record queued behind it in the partition. So DefaultErrorHandler classifies each exception as **retryable (recoverable)** or **not-retryable (fatal)**. **The default fatal set.** Out of the box these are treated as non-retryable and go straight to the recoverer: - `DeserializationException` (Spring Kafka) — the bytes can't be turned into the object; will never parse. - `MessageConversionException` / `ConversionException` — payload → method-argument conversion fails. - `MethodArgumentResolutionException` — can't bind arguments. - `NoSuchMethodException` — listener wiring problem. - `ClassCastException` — type mismatch, deterministic. All other exceptions are **retryable by default**. **Customizing.** - `handler.addRetryableExceptions(SomeTransient.class)` — force an exception (even a normally-fatal one) to be retried. - `handler.addNotRetryableExceptions(IllegalArgumentException.class)` — mark your own deterministic exceptions fatal so you don't burn retries on bad input. - `handler.setClassifications(Map<Class<? extends Throwable>, Boolean>, boolean defaultValue)` — replace the whole map and set the default. Passing `false` as the default inverts the policy to **'nothing retries unless explicitly listed'**, useful when most failures in your domain are deterministic. - `handler.removeClassification(...)` / and the classifier traverses the **cause chain** (`setClassifications`'s underlying `BinaryExceptionClassifier` checks causes) so a wrapped transient cause is still matched — you can control that with the classifier's `traverseCauses` behavior. **Interaction with BackOff.** Fatal → recover immediately (no BackOff). Retryable → apply the BackOff, then recover when exhausted. There's also per-exception BackOff support: `setBackOffFunction((record, ex) -> backOff)` lets you choose a different BackOff depending on the exception type (e.g. long back-off for rate-limit errors, short for lock contention). **Retry listeners & observability.** `handler.setRetryListeners(RetryListener...)` gives callbacks on each failed delivery and on recovery — useful for metrics ('how many records are retrying', 'what's going to DLT'). **Non-blocking parallel.** `@RetryableTopic` has the analogous concept: `@RetryableTopic(exclude = {...}, include = {...})` and `traversingCauses` control which exceptions get retried vs sent straight to the DLT. **Gotchas.** (1) If you throw a broad `RuntimeException` for genuinely-bad input, it's retryable by default — wrap or classify it as fatal, or you'll retry poison records 10× each and stall the partition. (2) The classifier checks the *actual* thrown exception and its causes — a `ListenerExecutionFailedException` wrapper is unwrapped for you. (3) Marking something fatal without a DLT recoverer means immediate skip = data loss. (4) `DeserializationException` being fatal is why you should pair with a DLT that can hold raw bytes.

  • You throw IllegalArgumentException for invalid payloads. With default settings, what happens and why is that bad?
    IllegalArgumentException is retryable by default, so the record is retried the full BackOff count even though it will fail identically every time — wasting attempts and, because retries block, delaying every record behind it. Fix: addNotRetryableExceptions(IllegalArgumentException.class) so it goes straight to the DLT.
  • How do you invert the policy so exceptions are NOT retried unless explicitly whitelisted?
    Call setClassifications(map, false) — the boolean is the default classification; false means non-retryable by default, and only the exceptions you list with value true are retried.

saying these in an interview costs you the question

  • Assuming all exceptions are retried the same number of times (deserialization/conversion/ClassCast are fatal by default)
  • Thinking you can only set one global BackOff (setBackOffFunction allows per-exception BackOff)
  • Believing the classifier ignores the cause chain (it traverses causes and unwraps the listener wrapper)

context