skip to content

What do retry.backoff.ms and request.timeout.ms control, and how do they interact during a failed send?

level: middleimportance: should knowfreq 45%

answer

  1. request.timeout.ms = per-attempt ack wait (30s)
  2. retry.backoff.ms = pause before retry (100ms)
  3. backoff is exponential + jittered (KIP-580), max 1000ms
  4. loop until delivery.timeout.ms
  5. set request.timeout > replica.lag for acks=all

basics

~20 s

request.timeout.ms is how long the producer waits for a single request's ack before giving up on that attempt. retry.backoff.ms is the pause before retrying a failed attempt, so the producer doesn't hammer a struggling broker. Both repeat until delivery.timeout.ms expires.

solid answer

~40 s

request.timeout.ms (default 30000) bounds a single in-flight ProduceRequest: if the broker doesn't acknowledge within it, that attempt fails with a retriable timeout. retry.backoff.ms (default 100) is the wait inserted before re-attempting a failed batch, preventing a tight retry loop against a broker that's mid-election or overloaded; modern clients add jitter and grow it exponentially up to retry.backoff.max.ms (default 1000). The cycle is: send attempt -> wait up to request.timeout.ms -> on retriable failure, sleep retry.backoff.ms -> retry. This loop continues until delivery.timeout.ms (the overall cap) is reached. So request.timeout.ms is per-attempt, retry.backoff.ms is the inter-attempt gap, and delivery.timeout.ms is the total budget enclosing both.

go deeper

for a junior

Know request.timeout.ms waits for one ack and retry.backoff.ms is the pause before retrying.

for a middle

Know defaults (30s, 100ms) and that the attempt/backoff loop runs until delivery.timeout.ms.

for a senior

Explain exponential backoff + jitter (KIP-580) and the request.timeout vs replica.lag relationship under acks=all.

for a principal

Reason about incident-time tuning trade-offs (fast failure detection vs. broker protection) across a fleet of producers.

## request.timeout.ms Default **30000 ms**. Once a `ProduceRequest` is on the wire, the producer waits this long for the broker's acknowledgement. If the ack doesn't arrive (slow broker, network stall, replication lag with `acks=all`), the attempt is failed with a **retriable** `TimeoutException` and the batch becomes eligible for retry. It is a **per-attempt** timeout — it does not bound the whole send. Note it should be set comfortably larger than the broker's `replica.lag.time.max.ms` when using `acks=all`, so that normal replication delays don't trip spurious timeouts. ## retry.backoff.ms Default **100 ms**. After a retriable failure, the producer waits this long before resending the batch. Its purpose is to avoid a hot loop hammering a broker that is temporarily unable to serve — e.g. during a leader election the producer would otherwise spin, refetch metadata, and retry instantly hundreds of times. Modern clients (KIP-580) apply **exponential backoff with jitter**, growing the wait up to **retry.backoff.max.ms** (default **1000 ms**), which smooths thundering-herd retries across many producers. ## How they interact The per-record lifecycle on a retriable error looks like: ``` attempt -> wait <= request.timeout.ms for ack -> retriable failure -> sleep retry.backoff.ms (growing, jittered) -> attempt again ... ``` This repeats until **`delivery.timeout.ms`** (default 120000) — the overall wall-clock budget — is exhausted, at which point the send fails permanently with a `TimeoutException`. ## Tuning interplay - Lowering `request.timeout.ms` surfaces stuck attempts faster but risks declaring healthy-but-slow brokers failed. - Raising `retry.backoff.ms` (or its max) reduces broker load during incidents at the cost of slower recovery. - All three must coexist: `delivery.timeout.ms` must be at least `linger.ms + request.timeout.ms`, and in practice large enough to fit several `request.timeout.ms + retry.backoff.ms` cycles if you want resilience to transient outages.

  • Why add jitter and exponential growth to retry.backoff.ms?
    To prevent a thundering herd: without jitter, many producers retry in lockstep and re-overload a recovering broker. Exponential growth backs off harder during sustained trouble; jitter spreads retries out in time.

saying these in an interview costs you the question

  • Confusing request.timeout.ms (per-attempt) with delivery.timeout.ms (total).
  • Saying retry.backoff.ms is the total retry budget — it's just the gap between attempts.
  • Setting request.timeout.ms below typical replication time with acks=all, causing spurious timeouts.

context