What do retry.backoff.ms and request.timeout.ms control, and how do they interact during a failed send?
answer
- request.timeout.ms = per-attempt ack wait (30s)
- retry.backoff.ms = pause before retry (100ms)
- backoff is exponential + jittered (KIP-580), max 1000ms
- loop until delivery.timeout.ms
- set request.timeout > replica.lag for acks=all
basics
~20 srequest.timeout.ms is how long the producer waits for a single request's ack before giving up on that attempt. retry.backoff.ms is the pause before retrying a failed attempt, so the producer doesn't hammer a struggling broker. Both repeat until delivery.timeout.ms expires.
solid answer
~40 srequest.timeout.ms (default 30000) bounds a single in-flight ProduceRequest: if the broker doesn't acknowledge within it, that attempt fails with a retriable timeout. retry.backoff.ms (default 100) is the wait inserted before re-attempting a failed batch, preventing a tight retry loop against a broker that's mid-election or overloaded; modern clients add jitter and grow it exponentially up to retry.backoff.max.ms (default 1000). The cycle is: send attempt -> wait up to request.timeout.ms -> on retriable failure, sleep retry.backoff.ms -> retry. This loop continues until delivery.timeout.ms (the overall cap) is reached. So request.timeout.ms is per-attempt, retry.backoff.ms is the inter-attempt gap, and delivery.timeout.ms is the total budget enclosing both.
go deeper
Know request.timeout.ms waits for one ack and retry.backoff.ms is the pause before retrying.
Know defaults (30s, 100ms) and that the attempt/backoff loop runs until delivery.timeout.ms.
Explain exponential backoff + jitter (KIP-580) and the request.timeout vs replica.lag relationship under acks=all.
Reason about incident-time tuning trade-offs (fast failure detection vs. broker protection) across a fleet of producers.
## request.timeout.ms Default **30000 ms**. Once a `ProduceRequest` is on the wire, the producer waits this long for the broker's acknowledgement. If the ack doesn't arrive (slow broker, network stall, replication lag with `acks=all`), the attempt is failed with a **retriable** `TimeoutException` and the batch becomes eligible for retry. It is a **per-attempt** timeout — it does not bound the whole send. Note it should be set comfortably larger than the broker's `replica.lag.time.max.ms` when using `acks=all`, so that normal replication delays don't trip spurious timeouts. ## retry.backoff.ms Default **100 ms**. After a retriable failure, the producer waits this long before resending the batch. Its purpose is to avoid a hot loop hammering a broker that is temporarily unable to serve — e.g. during a leader election the producer would otherwise spin, refetch metadata, and retry instantly hundreds of times. Modern clients (KIP-580) apply **exponential backoff with jitter**, growing the wait up to **retry.backoff.max.ms** (default **1000 ms**), which smooths thundering-herd retries across many producers. ## How they interact The per-record lifecycle on a retriable error looks like: ``` attempt -> wait <= request.timeout.ms for ack -> retriable failure -> sleep retry.backoff.ms (growing, jittered) -> attempt again ... ``` This repeats until **`delivery.timeout.ms`** (default 120000) — the overall wall-clock budget — is exhausted, at which point the send fails permanently with a `TimeoutException`. ## Tuning interplay - Lowering `request.timeout.ms` surfaces stuck attempts faster but risks declaring healthy-but-slow brokers failed. - Raising `retry.backoff.ms` (or its max) reduces broker load during incidents at the cost of slower recovery. - All three must coexist: `delivery.timeout.ms` must be at least `linger.ms + request.timeout.ms`, and in practice large enough to fit several `request.timeout.ms + retry.backoff.ms` cycles if you want resilience to transient outages.
- Why add jitter and exponential growth to retry.backoff.ms?To prevent a thundering herd: without jitter, many producers retry in lockstep and re-overload a recovering broker. Exponential growth backs off harder during sustained trouble; jitter spreads retries out in time.
saying these in an interview costs you the question
- Confusing request.timeout.ms (per-attempt) with delivery.timeout.ms (total).
- Saying retry.backoff.ms is the total retry budget — it's just the gap between attempts.
- Setting request.timeout.ms below typical replication time with acks=all, causing spurious timeouts.