skip to content

What is the difference between a retriable and a fatal (non-retriable) exception in the Kafka producer, and how does each affect a send?

level: juniorimportance: must knowfreq 70%

answer

  1. RetriableException base class
  2. transient (leader change/network) vs invalid request
  3. fatal = serialize/too-large/authz
  4. producer auto-resends retriable only
  5. exhausted retries -> TimeoutException

basics

~20 s

Retriable errors (like a leader change or timeout) are transient, so the producer automatically resends. Fatal errors (like an unknown topic, message too large, or auth failure) cannot be fixed by resending, so the send fails immediately.

solid answer

~30 s

Kafka classifies broker/network errors as retriable or fatal. Retriable exceptions extend RetriableException (e.g. NotLeaderOrFollowerException, NetworkException, request timeouts, NotEnoughReplicasException) and represent transient conditions the producer resolves by resending the record — controlled by retries / delivery.timeout.ms and spaced by retry.backoff.ms. Fatal exceptions (e.g. RecordTooLargeException, SerializationException, UnknownTopicOrPartitionException when auto-create is off, AuthorizationException, InvalidConfigurationException) cannot be cured by retrying, so the future completes exceptionally at once and your callback/Future.get sees the error. The key practical point: the producer retries transient failures for you, so application code should focus on handling the fatal ones and on the case where retries are eventually exhausted.

go deeper

for a junior

Know the two buckets: transient -> auto-retried; invalid/permanent -> fails fast.

for a middle

Name concrete examples per bucket and the RetriableException base class; know retries are bounded by delivery.timeout.ms.

for a senior

Explain the duplicate risk when a retriable timeout follows a lost ack, and how idempotence addresses it.

for a principal

Frame error classification as the contract that lets you build at-least-once vs exactly-once semantics on top, and reason about which fatal errors should page vs drop.

## Background When you call `producer.send(record)`, the Kafka producer batches the record, picks the partition leader, and sends a `ProduceRequest` to that broker. Many things can go wrong: the broker may be momentarily unavailable, a leader election may be in progress, the network may hiccup, or the request itself may be fundamentally invalid. Kafka divides these failures into two families. ## Retriable exceptions These extend `org.apache.kafka.common.errors.RetriableException`. They describe a **transient** condition that is likely to disappear on its own: - `NotLeaderOrFollowerException` / `LeaderNotAvailableException` — a leader election is happening; metadata will refresh and a new leader will appear. - `NetworkException` — a connection dropped. - `TimeoutException` — the request or batch exceeded a timeout but might succeed if resent. - `NotEnoughReplicasException` — fewer than `min.insync.replicas` are in-sync right now (with `acks=all`); replicas may catch up. - `CoordinatorNotAvailableException`, `UnknownTopicOrPartitionException` (transiently, right after topic creation). For these, the producer **automatically resends** the batch. How long it keeps trying is bounded by `delivery.timeout.ms` (and historically `retries`), and each attempt waits `retry.backoff.ms`. ## Fatal / non-retriable exceptions These do **not** extend `RetriableException`. Resending cannot help because the problem is in the request itself or the cluster state: - `RecordTooLargeException` — record exceeds `max.request.size` or the broker's `message.max.bytes`. - `SerializationException` — the key/value serializer threw (often thrown synchronously from `send`). - `UnknownTopicOrPartitionException` (persistent) — the topic does not exist and auto-create is disabled. - `TopicAuthorizationException` / `ClusterAuthorizationException` — ACL denial. - `InvalidConfigurationException`, `UnsupportedVersionException`. For these the record's `Future`/callback **completes exceptionally immediately**; no retry occurs. ## Why it matters Because the producer transparently retries the transient family, application code mostly needs to handle (a) fatal exceptions surfaced in the send callback and (b) the situation where retries are exhausted within `delivery.timeout.ms`, which surfaces as a `TimeoutException`. A subtle edge case: a `TimeoutException` after the request was actually written but the ack was lost can cause a *duplicate* on retry — which is exactly why idempotence (`enable.idempotence=true`) exists.

  • Name two exceptions the producer will NOT retry.
    RecordTooLargeException and any AuthorizationException (e.g. TopicAuthorizationException); also SerializationException, which is usually thrown synchronously from send().
  • If retries succeed automatically, why does application code still need a send callback?
    To handle fatal exceptions and the case where retries are exhausted (surfacing as TimeoutException), and to observe successful metadata (offset/partition). The callback is the only place to react to a permanently failed send.

saying these in an interview costs you the question

  • Saying the producer retries every exception including serialization or authorization errors.
  • Claiming you must manually catch and resend retriable errors — the producer does it for you.
  • Treating TimeoutException as always meaning the message was not written (it may have been written but the ack was lost).

context