Explain the difference between a 'record-level' fatal exception and a container-fatal exception in Spring Kafka error handling. How does each affect the consumer?
answer
- Record-fatal = non-retryable → skip to DLT, keep going
- Container-fatal = stop the whole listener
- Classifier: addNotRetryable / setClassifications / defaultFalse
- Deser/Conversion/ClassCast = record-fatal defaults
- Container down = alert + restart, not graceful skip
basics
~20 sA record-level fatal exception means 'don't retry THIS record' — it's classified non-retryable so the error handler skips straight to recovery (DLT) and the consumer keeps going. A container-fatal exception is more severe: it can stop the whole listener container, halting consumption for all partitions it owns.
solid answer
~40 sThese are two different notions of 'fatal'. A record-level fatal (a.k.a. non-retryable) exception is about retry classification: DefaultErrorHandler keeps a list (DeserializationException, MessageConversionException, ClassCastException, etc.) that it won't retry; it skips the BackOff and goes straight to the recoverer/DLT, then continues consuming. You tune this with addNotRetryableExceptions/addRetryableExceptions and setClassifications. A container-fatal error is one the container can't recover from per-record — historically thrown as a Throwable the handler can't handle, or surfaced via setCommitRecovered/ack issues, or an authorization/config error. By default certain exceptions cause the container to stop (so it doesn't spin). You influence this with isAckAfterHandle, the container's CommonErrorHandler.handleOtherException path, and properties governing whether the container stops or pauses. The practical contrast: record-fatal = skip one record and continue; container-fatal = stop/pause the listener so a human or restart intervenes.
go deeper
Know there are two kinds of fatal: one skips a single record to the DLT, the other can stop the whole consumer.
Identify the default non-retryable exceptions and customize the classifier with addNotRetryable/setClassifications.
Explain both scopes precisely, the operational implications, and monitoring (DLT volume vs container running state).
Design failure taxonomy, alerting, and remediation runbooks distinguishing quarantine-and-continue from halt-and-page across services.
## Two scopes of 'fatal' Spring Kafka's error handling reasons about failures at two scopes, and conflating them is a common interview slip. ### Record-level fatal (a.k.a. **not-retryable**) This is purely a **retry-classification** decision for a *single record*. `DefaultErrorHandler` holds a **classifier** (`BinaryExceptionClassifier`) seeded with exceptions that should **never be retried**, because retrying can't help: - `DeserializationException`, `MessageConversionException`, `ConversionException` - `MethodArgumentResolutionException`, `NoSuchMethodException`, `ClassCastException` - `org.springframework.kafka.support.serializer.DeserializationException` (from `ErrorHandlingDeserializer`) When one of these is thrown, the handler **skips the BackOff entirely** and goes **straight to the recoverer** (typically `DeadLetterPublishingRecoverer` → DLT). The consumer then **commits past** the record and **keeps consuming**. You customize the list: ``` handler.addNotRetryableExceptions(IllegalArgumentException.class); handler.addRetryableExceptions(MyTransientException.class); handler.setClassifications(map, defaultRetryable); // full control + defaultFalse() ``` `defaultFalse()` makes the classifier **deny-by-default** (only listed exceptions retry). ### Container-fatal This is about the **whole listener container** (`MessageListenerContainer`), which owns one or more partitions on a consumer thread. Some failures aren't about one record at all: - The error handler itself **can't handle** the throwable (`CommonErrorHandler.handleOtherException`). - **Authorization** failures, fatal **configuration**/serialization-setup errors, or unrecoverable consumer errors. - A recoverer that **keeps failing** (e.g. can't publish to the DLT) — the offset isn't committed and the situation may escalate. When a container-fatal condition occurs, the container may **stop** (cease polling, so consumption halts for all its partitions) rather than spin in a tight failure loop. This is intentional: it surfaces a systemic problem (bad credentials, misconfig) loudly instead of silently dropping data. Recovery generally needs human intervention or a restart. ## Why the distinction matters operationally - **Record-fatal** → graceful degradation: one bad record is quarantined to the DLT, throughput continues. Alert on DLT volume. - **Container-fatal** → the listener is **down**; you need health checks / alerts on container state (`KafkaListenerEndpointRegistry`, container `isRunning()`) and a remediation path, because no records are being processed. ## Edge cases - A record-fatal exception still uses the recoverer; if the recoverer (DLT publish) fails, that failure can escalate toward container-level handling. - `setCommitRecovered(true)` ensures the offset is committed after recovery so a recovered record isn't reprocessed after a rebalance. - For batch listeners, throwing the wrong exception type can turn a single-record problem into reprocessing the whole batch — use `BatchListenerFailedException` to pinpoint the offending record. - You can register `setRetryListeners(...)` to observe each retry and the final recovery, useful for metrics distinguishing the two scopes.
- How do you make IllegalArgumentException non-retryable but keep retrying a custom transient exception?On DefaultErrorHandler call addNotRetryableExceptions(IllegalArgumentException.class) and addRetryableExceptions(MyTransientException.class), or use setClassifications(map, defaultRetryable) for full control including defaultFalse() to deny-by-default.
- Why might Spring deliberately stop a container instead of skipping on a fatal error?To avoid silently dropping data and to surface systemic problems (bad credentials, misconfiguration, persistently failing recoverer) loudly. Stopping forces an alert/restart rather than spinning in a tight failure loop or quietly losing records.
saying these in an interview costs you the question
- Treating 'fatal' as one concept — record-fatal (skip one record) and container-fatal (stop the listener) are different scopes.
- Saying record-fatal exceptions are retried — they skip the backoff and go straight to recovery.
- Assuming a stopped container auto-recovers — it typically needs a restart or intervention.
- Forgetting to monitor container running state, so a container-fatal stop goes unnoticed.