skip to content

Explain errors.retry.timeout and errors.retry.delay.max.ms. How do retries interact with errors.tolerance?

level: seniorimportance: should knowfreq 55%

answer

  1. retry.timeout: 0=off, -1=forever
  2. delay.max.ms default 60000, exp backoff + jitter
  3. retries FIRST, then tolerance
  4. retries for transient (RetriableException)
  5. useless for deterministic bad records

basics

~20 s

errors.retry.timeout is how long Connect keeps retrying a failed operation (default 0 = no retries; -1 = forever). errors.retry.delay.max.ms caps the backoff between attempts (default 60000). Retries run first; only after they're exhausted does errors.tolerance decide skip vs fail.

solid answer

~50 s

These two configs govern automatic retry of operations that fail in a potentially transient way (e.g. a sink put() to a downstream system that's temporarily unavailable). errors.retry.timeout sets the total time budget for retrying a single failed operation: 0 (default) disables retries, a positive value is the milliseconds to keep trying, and -1 retries indefinitely. errors.retry.delay.max.ms (default 60000) caps the exponential backoff between successive attempts — Connect doubles the delay each retry up to this ceiling, with jitter. The ordering is the key insight: when an operation fails, Connect retries it within the timeout budget first; only if it still fails after retries are exhausted does errors.tolerance take effect. So tolerance=none + exhausted retries → task fails; tolerance=all + exhausted retries → record is logged/sent to DLQ and skipped. Retries help transient failures (network blips, downstream restarts); they do nothing for deterministic failures like malformed records, where every retry fails identically and just wastes the timeout budget.

go deeper

for a junior

Know retries exist for transient failures and that retry.timeout=0 by default means no retries.

for a middle

Distinguish retry.timeout (total budget; 0/-1 specials) from delay.max.ms (per-attempt backoff cap, default 60s).

for a senior

Articulate the retries-then-tolerance ordering and that retries only help RetriableException/transient failures, not deterministic bad records.

for a principal

Tune retry budgets against downstream SLAs, reason about partition-stall risk from long/-1 timeouts, and pair retries with DLQ for post-budget quarantine.

## The problem retries solve Not all failures are equal. A **transient** failure — the downstream database is briefly unreachable, a connection reset, a momentary timeout — will likely succeed if you simply try again a moment later. A **deterministic** failure — malformed JSON, a schema violation — fails identically every time. Kafka Connect's retry mechanism exists for the transient class. ## errors.retry.timeout This is the **total time budget** (in ms) Connect will spend retrying a *single* failed operation before giving up: - **0** (default): no retries — fail (or tolerate) on the first error. - **positive N**: keep retrying for up to N milliseconds. - **-1**: retry indefinitely (use with care — a permanently broken downstream stalls the task forever). ## errors.retry.delay.max.ms Connect uses **exponential backoff with jitter** between retry attempts: the delay grows (roughly doubling) each attempt to avoid hammering a struggling downstream. `errors.retry.delay.max.ms` (default **60000**, i.e. 60s) is the **ceiling** on that per-attempt delay. So delays might go 300ms → 600ms → 1.2s → … capped at 60s. Jitter randomizes the exact value to prevent thundering-herd retries across tasks. ## Ordering: retries THEN tolerance This is the most-tested nuance. The pipeline for a failing operation is: 1. Operation fails. 2. Connect retries within the `errors.retry.timeout` budget, backing off up to `errors.retry.delay.max.ms`. 3. If a retry succeeds → processing continues normally, no error recorded. 4. If the budget is exhausted and it still fails → **now** `errors.tolerance` decides: - `none` → task goes to FAILED and stops. - `all` → the record is logged (if `errors.log.enable`) and/or sent to the DLQ, then skipped. So retries and tolerance are complementary, not alternatives: retries handle 'maybe transient', tolerance handles 'definitely give up now'. ## Which failures are retryable Retries apply primarily to operations that can throw a **RetriableException** — typically the sink connector's `put()` against the external system, or producer/consumer interactions. Pure conversion/SMT failures on a bad record are deterministic and effectively non-retriable; configuring a large retry timeout there just delays the inevitable skip/fail by the whole budget. ## Practical guidance - For sinks writing to flaky downstreams: set a modest `errors.retry.timeout` (e.g. 30000–300000) so brief outages self-heal without failing the task. - Keep `errors.retry.delay.max.ms` reasonable so backoff doesn't grow to many minutes and stall the partition. - Avoid `-1` unless you have external monitoring/alerting on a stuck task. - Combine with `errors.tolerance=all` + DLQ so that records still failing after the retry budget are quarantined rather than crashing the task.

  • If errors.tolerance=none and errors.retry.timeout=30000, when exactly does the task fail?
    Only after the operation has been retried for the full 30 seconds and still fails. Retries run first; the task fails when the retry budget is exhausted, not on the first error.
  • Why is a large errors.retry.timeout pointless for a malformed-record conversion error?
    Conversion of a deterministically bad record fails identically on every attempt, so retrying just burns the whole timeout budget before the inevitable skip/fail. Retries only help transient (RetriableException) failures like a temporarily unavailable downstream.
  • What is the risk of errors.retry.timeout=-1?
    If the downstream is permanently broken, the task retries forever and never makes progress (or fails), silently stalling the partition unless you have external monitoring/alerting on stuck tasks.

saying these in an interview costs you the question

  • Saying retries and tolerance are mutually exclusive — they compose, retries run first
  • Claiming retries help malformed records
  • Stating the default retry.timeout is non-zero (it's 0 = no retries)
  • Thinking delay.max.ms is the total retry duration (it's the per-attempt backoff cap)

context