skip to content

Error Handling, Retries and Ordering

Which producer errors are retriable, how delivery.timeout.ms bounds the whole attempt, and when retries can reorder records. Interviewers ask because retries combined with several in-flight requests is a subtle correctness bug.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What is the difference between a retriable and a fatal (non-retriable) exception in the Kafka producer, and how does each affect a send?

level: juniorimportance: must knowfreq 70%

answer

  1. RetriableException base class
  2. transient (leader change/network) vs invalid request
  3. fatal = serialize/too-large/authz
  4. producer auto-resends retriable only
  5. exhausted retries -> TimeoutException

basics

~20 s

Retriable errors (like a leader change or timeout) are transient, so the producer automatically resends. Fatal errors (like an unknown topic, message too large, or auth failure) cannot be fixed by resending, so the send fails immediately.

solid answer

~30 s

Kafka classifies broker/network errors as retriable or fatal. Retriable exceptions extend RetriableException (e.g. NotLeaderOrFollowerException, NetworkException, request timeouts, NotEnoughReplicasException) and represent transient conditions the producer resolves by resending the record — controlled by retries / delivery.timeout.ms and spaced by retry.backoff.ms. Fatal exceptions (e.g. RecordTooLargeException, SerializationException, UnknownTopicOrPartitionException when auto-create is off, AuthorizationException, InvalidConfigurationException) cannot be cured by retrying, so the future completes exceptionally at once and your callback/Future.get sees the error. The key practical point: the producer retries transient failures for you, so application code should focus on handling the fatal ones and on the case where retries are eventually exhausted.

go deeper

for a junior

Know the two buckets: transient -> auto-retried; invalid/permanent -> fails fast.

for a middle

Name concrete examples per bucket and the RetriableException base class; know retries are bounded by delivery.timeout.ms.

for a senior

Explain the duplicate risk when a retriable timeout follows a lost ack, and how idempotence addresses it.

for a principal

Frame error classification as the contract that lets you build at-least-once vs exactly-once semantics on top, and reason about which fatal errors should page vs drop.

## Background When you call `producer.send(record)`, the Kafka producer batches the record, picks the partition leader, and sends a `ProduceRequest` to that broker. Many things can go wrong: the broker may be momentarily unavailable, a leader election may be in progress, the network may hiccup, or the request itself may be fundamentally invalid. Kafka divides these failures into two families. ## Retriable exceptions These extend `org.apache.kafka.common.errors.RetriableException`. They describe a **transient** condition that is likely to disappear on its own: - `NotLeaderOrFollowerException` / `LeaderNotAvailableException` — a leader election is happening; metadata will refresh and a new leader will appear. - `NetworkException` — a connection dropped. - `TimeoutException` — the request or batch exceeded a timeout but might succeed if resent. - `NotEnoughReplicasException` — fewer than `min.insync.replicas` are in-sync right now (with `acks=all`); replicas may catch up. - `CoordinatorNotAvailableException`, `UnknownTopicOrPartitionException` (transiently, right after topic creation). For these, the producer **automatically resends** the batch. How long it keeps trying is bounded by `delivery.timeout.ms` (and historically `retries`), and each attempt waits `retry.backoff.ms`. ## Fatal / non-retriable exceptions These do **not** extend `RetriableException`. Resending cannot help because the problem is in the request itself or the cluster state: - `RecordTooLargeException` — record exceeds `max.request.size` or the broker's `message.max.bytes`. - `SerializationException` — the key/value serializer threw (often thrown synchronously from `send`). - `UnknownTopicOrPartitionException` (persistent) — the topic does not exist and auto-create is disabled. - `TopicAuthorizationException` / `ClusterAuthorizationException` — ACL denial. - `InvalidConfigurationException`, `UnsupportedVersionException`. For these the record's `Future`/callback **completes exceptionally immediately**; no retry occurs. ## Why it matters Because the producer transparently retries the transient family, application code mostly needs to handle (a) fatal exceptions surfaced in the send callback and (b) the situation where retries are exhausted within `delivery.timeout.ms`, which surfaces as a `TimeoutException`. A subtle edge case: a `TimeoutException` after the request was actually written but the ack was lost can cause a *duplicate* on retry — which is exactly why idempotence (`enable.idempotence=true`) exists.

  • Name two exceptions the producer will NOT retry.
    RecordTooLargeException and any AuthorizationException (e.g. TopicAuthorizationException); also SerializationException, which is usually thrown synchronously from send().
  • If retries succeed automatically, why does application code still need a send callback?
    To handle fatal exceptions and the case where retries are exhausted (surfacing as TimeoutException), and to observe successful metadata (offset/partition). The callback is the only place to react to a permanently failed send.

saying these in an interview costs you the question

  • Saying the producer retries every exception including serialization or authorization errors.
  • Claiming you must manually catch and resend retriable errors — the producer does it for you.
  • Treating TimeoutException as always meaning the message was not written (it may have been written but the ack was lost).

context

open as a page

Explain delivery.timeout.ms (KIP-91) and how it relates to retries, request.timeout.ms, and linger.ms in bounding the total time a send can take.

level: seniorimportance: must knowfreq 60%

basics

~20 s

delivery.timeout.ms is the single upper bound on the total time from send() to success or failure, covering batching, all retries, and inflight requests. It must be >= linger.ms + request.timeout.ms. When it expires, the record fails with TimeoutException regardless of remaining retries.

open as a page

How can max.in.flight.requests.per.connection cause message reordering with retries, and how do you prevent it while keeping throughput?

level: seniorimportance: must knowfreq 65%

basics

~20 s

If more than one request is in flight per connection and an earlier batch is retried while a later one already succeeded, the retried batch lands after it, reordering messages within a partition. Enabling idempotence (enable.idempotence=true) prevents reordering even with up to 5 in-flight requests.

open as a page

What do retry.backoff.ms and request.timeout.ms control, and how do they interact during a failed send?

level: middleimportance: should knowfreq 45%

basics

~20 s

request.timeout.ms is how long the producer waits for a single request's ack before giving up on that attempt. retry.backoff.ms is the pause before retrying a failed attempt, so the producer doesn't hammer a struggling broker. Both repeat until delivery.timeout.ms expires.

open as a page

When does an idempotent producer throw OutOfOrderSequenceException, and what does it imply about delivery guarantees?

level: principalimportance: should knowfreq 35%

basics

~20 s

It means the broker received a producer's batch with a sequence number that doesn't follow the last one it accepted, so a gap exists — usually because an earlier batch was permanently lost or its state expired. It signals the producer can no longer guarantee ordered, gap-free delivery for that session.

open as a page