skip to content

Explain delivery.timeout.ms (KIP-91) and how it relates to retries, request.timeout.ms, and linger.ms in bounding the total time a send can take.

level: seniorimportance: must knowfreq 60%

answer

  1. KIP-91, default 120000ms
  2. single end-to-end cap: batch + retries + inflight
  3. delivery.timeout.ms >= linger.ms + request.timeout.ms
  4. retries default Integer.MAX_VALUE since 2.1
  5. expiry -> TimeoutException regardless of retries left

basics

~20 s

delivery.timeout.ms is the single upper bound on the total time from send() to success or failure, covering batching, all retries, and inflight requests. It must be >= linger.ms + request.timeout.ms. When it expires, the record fails with TimeoutException regardless of remaining retries.

solid answer

~40 s

KIP-91 introduced delivery.timeout.ms (default 120000) as the authoritative cap on how long the producer will spend on a record after send() returns. The clock covers the whole lifecycle: time waiting in the accumulator (bounded by linger.ms and batch.size), each ProduceRequest attempt (bounded by request.timeout.ms), and the retry.backoff.ms pauses between attempts. retries then becomes effectively secondary — the producer keeps retrying retriable errors until delivery.timeout.ms is hit, then completes the future with a TimeoutException. The constraint delivery.timeout.ms >= linger.ms + request.timeout.ms is enforced at construction. This gives one tunable that maps directly to an SLA ('a message must succeed or fail within N ms') instead of the old, hard-to-reason-about combination of retries x (request.timeout.ms + backoff).

go deeper

for a junior

Know there's one config (delivery.timeout.ms) that caps how long a send can take.

for a middle

Know its default (120s) and that it covers batching + retries + the inflight request, superseding raw retries.

for a senior

Articulate the enforced inequality and how to tune it for latency vs durability SLAs.

for a principal

Reason about deadline-expiry semantics (late acks, no duplicate with idempotence) and how delivery.timeout.ms anchors producer behavior to system-level latency budgets.

## The problem KIP-91 solved Before Kafka 2.1, the maximum time a `send()` could take was an obscure function of several configs: `retries`, `request.timeout.ms`, `retry.backoff.ms`, plus metadata-fetch timeouts. There was no single knob that said 'fail this record after N milliseconds.' Operators could not map producer behavior to an end-to-end latency SLA. KIP-91 added **`delivery.timeout.ms`** (default 120000 ms = 2 minutes) as that single bound. ## What the clock covers `delivery.timeout.ms` starts when `send()` returns (the record is appended to the accumulator) and covers, in order: 1. **Batching delay** — time the record waits in the `RecordAccumulator` until the batch is ready, governed by `linger.ms` (max wait) and `batch.size` (fills early). 2. **Each in-flight attempt** — a `ProduceRequest` is sent and awaited up to `request.timeout.ms`. If the broker doesn't ack in time, that attempt fails with a (retriable) timeout. 3. **Backoff between retries** — `retry.backoff.ms` (with jitter / exponential growth in newer clients, capped by `retry.backoff.max.ms`). The producer loops over steps 2–3 for retriable errors **until the total elapsed time reaches `delivery.timeout.ms`**. At that point the record's callback/Future completes with `org.apache.kafka.common.errors.TimeoutException`, even if `retries` has not been exhausted. ## Relationship to `retries` Since Kafka 2.1, `retries` defaults to `Integer.MAX_VALUE` precisely because `delivery.timeout.ms` is the real bound. Setting a small `retries` can still cut things short, but the recommended pattern is to leave `retries` high and tune `delivery.timeout.ms` to your SLA. Whichever limit is reached first ends the attempts. ## The enforced inequality The client validates at construction: ``` delivery.timeout.ms >= linger.ms + request.timeout.ms ``` It must be at least enough for one batching window plus one request attempt; otherwise the config is rejected. If you raise `request.timeout.ms` or `linger.ms`, you may need to raise `delivery.timeout.ms` too. ## Practical tuning - Latency-sensitive: lower `delivery.timeout.ms` so failures surface fast and the app can react (e.g. shed load, route elsewhere). - Durability-sensitive: raise it so transient broker unavailability (rolling restart, leader election storm) is ridden out without dropping records. - A `TimeoutException` here does **not** guarantee the record was not written — late acks can still arrive after the deadline; with idempotence on, a real retry won't duplicate, but a deadline-expiry simply abandons the record.

  • Why does retries default to Integer.MAX_VALUE in modern Kafka clients?
    Because delivery.timeout.ms is now the real bound on send duration; the producer retries until that wall-clock deadline, so an explicit retry count is usually unnecessary and just risks cutting things short early.
  • Your producer rejects its config at startup complaining about delivery.timeout.ms. What's the likely cause?
    delivery.timeout.ms is set below linger.ms + request.timeout.ms. Raise delivery.timeout.ms (or lower linger/request.timeout) to satisfy the enforced inequality.

saying these in an interview costs you the question

  • Saying retries alone bounds total send time — it doesn't; delivery.timeout.ms does.
  • Claiming delivery.timeout.ms only covers the network request and not the batching/backoff time.
  • Thinking request.timeout.ms is the total send budget; it's per-attempt.
  • Assuming a delivery TimeoutException proves the record was never persisted.

context