Explain errors.retry.timeout and errors.retry.delay.max.ms. How do retries interact with errors.tolerance?
answer
- retry.timeout: 0=off, -1=forever
- delay.max.ms default 60000, exp backoff + jitter
- retries FIRST, then tolerance
- retries for transient (RetriableException)
- useless for deterministic bad records
basics
~20 serrors.retry.timeout is how long Connect keeps retrying a failed operation (default 0 = no retries; -1 = forever). errors.retry.delay.max.ms caps the backoff between attempts (default 60000). Retries run first; only after they're exhausted does errors.tolerance decide skip vs fail.
solid answer
~50 sThese two configs govern automatic retry of operations that fail in a potentially transient way (e.g. a sink put() to a downstream system that's temporarily unavailable). errors.retry.timeout sets the total time budget for retrying a single failed operation: 0 (default) disables retries, a positive value is the milliseconds to keep trying, and -1 retries indefinitely. errors.retry.delay.max.ms (default 60000) caps the exponential backoff between successive attempts — Connect doubles the delay each retry up to this ceiling, with jitter. The ordering is the key insight: when an operation fails, Connect retries it within the timeout budget first; only if it still fails after retries are exhausted does errors.tolerance take effect. So tolerance=none + exhausted retries → task fails; tolerance=all + exhausted retries → record is logged/sent to DLQ and skipped. Retries help transient failures (network blips, downstream restarts); they do nothing for deterministic failures like malformed records, where every retry fails identically and just wastes the timeout budget.
go deeper
Know retries exist for transient failures and that retry.timeout=0 by default means no retries.
Distinguish retry.timeout (total budget; 0/-1 specials) from delay.max.ms (per-attempt backoff cap, default 60s).
Articulate the retries-then-tolerance ordering and that retries only help RetriableException/transient failures, not deterministic bad records.
Tune retry budgets against downstream SLAs, reason about partition-stall risk from long/-1 timeouts, and pair retries with DLQ for post-budget quarantine.
## The problem retries solve Not all failures are equal. A **transient** failure — the downstream database is briefly unreachable, a connection reset, a momentary timeout — will likely succeed if you simply try again a moment later. A **deterministic** failure — malformed JSON, a schema violation — fails identically every time. Kafka Connect's retry mechanism exists for the transient class. ## errors.retry.timeout This is the **total time budget** (in ms) Connect will spend retrying a *single* failed operation before giving up: - **0** (default): no retries — fail (or tolerate) on the first error. - **positive N**: keep retrying for up to N milliseconds. - **-1**: retry indefinitely (use with care — a permanently broken downstream stalls the task forever). ## errors.retry.delay.max.ms Connect uses **exponential backoff with jitter** between retry attempts: the delay grows (roughly doubling) each attempt to avoid hammering a struggling downstream. `errors.retry.delay.max.ms` (default **60000**, i.e. 60s) is the **ceiling** on that per-attempt delay. So delays might go 300ms → 600ms → 1.2s → … capped at 60s. Jitter randomizes the exact value to prevent thundering-herd retries across tasks. ## Ordering: retries THEN tolerance This is the most-tested nuance. The pipeline for a failing operation is: 1. Operation fails. 2. Connect retries within the `errors.retry.timeout` budget, backing off up to `errors.retry.delay.max.ms`. 3. If a retry succeeds → processing continues normally, no error recorded. 4. If the budget is exhausted and it still fails → **now** `errors.tolerance` decides: - `none` → task goes to FAILED and stops. - `all` → the record is logged (if `errors.log.enable`) and/or sent to the DLQ, then skipped. So retries and tolerance are complementary, not alternatives: retries handle 'maybe transient', tolerance handles 'definitely give up now'. ## Which failures are retryable Retries apply primarily to operations that can throw a **RetriableException** — typically the sink connector's `put()` against the external system, or producer/consumer interactions. Pure conversion/SMT failures on a bad record are deterministic and effectively non-retriable; configuring a large retry timeout there just delays the inevitable skip/fail by the whole budget. ## Practical guidance - For sinks writing to flaky downstreams: set a modest `errors.retry.timeout` (e.g. 30000–300000) so brief outages self-heal without failing the task. - Keep `errors.retry.delay.max.ms` reasonable so backoff doesn't grow to many minutes and stall the partition. - Avoid `-1` unless you have external monitoring/alerting on a stuck task. - Combine with `errors.tolerance=all` + DLQ so that records still failing after the retry budget are quarantined rather than crashing the task.
- If errors.tolerance=none and errors.retry.timeout=30000, when exactly does the task fail?Only after the operation has been retried for the full 30 seconds and still fails. Retries run first; the task fails when the retry budget is exhausted, not on the first error.
- Why is a large errors.retry.timeout pointless for a malformed-record conversion error?Conversion of a deterministically bad record fails identically on every attempt, so retrying just burns the whole timeout budget before the inevitable skip/fail. Retries only help transient (RetriableException) failures like a temporarily unavailable downstream.
- What is the risk of errors.retry.timeout=-1?If the downstream is permanently broken, the task retries forever and never makes progress (or fails), silently stalling the partition unless you have external monitoring/alerting on stuck tasks.
saying these in an interview costs you the question
- Saying retries and tolerance are mutually exclusive — they compose, retries run first
- Claiming retries help malformed records
- Stating the default retry.timeout is non-zero (it's 0 = no retries)
- Thinking delay.max.ms is the total retry duration (it's the per-attempt backoff cap)