Which categories of failures does Kafka Connect's error-handling framework tolerate, and which does it NOT? Distinguish conversion/transform/put failures from other error types.
answer
- controlled boundary: converter, SMT, put()/poll()
- KIP-298 error handling, KIP-610 ErrantRecordReporter
- NOT tolerated: consumer/producer, offset commit, rebalance, config
- conversion = clean per-record DLQ
- batch put() needs ErrantRecordReporter for per-record DLQ
basics
~20 sThe framework tolerates errors in the parts it controls: converters (deserialization/serialization), SMTs, and the sink connector's put() handoff. It does NOT tolerate failures outside that boundary — like consumer/producer/offset-commit errors or framework/configuration errors, which still fail the task.
solid answer
~50 serrors.tolerance and the DLQ apply only to the **operation-scoped** stages Connect mediates per record: converter failures (e.g. deserializing bad JSON/Avro, serializing on a source), Single Message Transform failures, and the sink connector's put() (and source poll()) handoff where a RetriableException or a tolerated exception can be caught. Within that boundary, retries run first and then tolerance decides skip-vs-fail. Outside it, the framework does NOT apply tolerance: errors from the Kafka consumer/producer themselves, offset commit failures, worker/rebalance issues, connector startup/validation, or fatal configuration errors are not 'per-record' tolerable — they fail the task or worker regardless of errors.tolerance=all. A subtle case: a sink connector that throws a *non-retriable* exception from put() for the whole batch (rather than reporting a per-record failure) can still take down the task; clean per-record DLQ behavior depends on the connector reporting individual record errors (via the ErrantRecordReporter) rather than failing the batch. So tolerance is bounded to deterministic per-record processing inside Connect's controlled pipeline, not arbitrary infrastructure failures.
go deeper
Know tolerance covers bad-record processing (conversion, transforms, sink writes), not general infrastructure failures.
List the controlled stages (converter, SMT, put/poll) and name examples that are NOT tolerated (consumer/producer, offset commit, config).
Explain the per-record attribution boundary, the conversion-vs-batch-put distinction, and the ErrantRecordReporter requirement for clean put() DLQ.
Reason about end-to-end resilience: which failures need tolerance/DLQ vs retries vs alerting/failover, and how connector capabilities (KIP-610 support) shape the achievable guarantees.
## The 'controlled boundary' concept Kafka Connect's error-handling framework (KIP-298) wraps the stages where it processes records **one at a time** and can therefore isolate a single bad record. Tolerance/DLQ/logging/retries apply **only inside this boundary**: 1. **Converters** — deserializing Kafka bytes into Connect records (sink) or serializing Connect records into bytes (source). Bad JSON, schema-registry misses, type mismatches surface here. 2. **Single Message Transforms (SMTs)** — per-record transformation logic; an SMT that throws on a record is tolerable. 3. **The connector handoff** — sink `put()` / source `poll()`. Errors writing a record downstream (or a `RetriableException`) are caught here; retries apply, then tolerance. Within this boundary the flow is: operation fails → retry within `errors.retry.timeout` → if still failing, `errors.tolerance` decides (none=fail task, all=log/DLQ/skip). ## What is NOT tolerated Errors **outside** the per-record controlled stages are not subject to `errors.tolerance`: - **Kafka client errors**: the underlying consumer (sink) or producer (source) failing — authentication failures, broker unavailability beyond client retries, fatal serialization at the client level. - **Offset commit failures**: committing consumer offsets is not a per-record tolerable operation. - **Rebalance / worker / framework errors**: task assignment, group coordination, worker crashes. - **Connector startup, configuration, and validation errors**: a misconfigured connector fails fast regardless of tolerance. - **Out-of-memory / JVM-level failures**. These fail the task (or worker) even with `errors.tolerance=all`, because they aren't 'this one record is bad' situations. ## Conversion/transform vs put() — the nuance - **Conversion and SMT failures** are inherently **per-record and deterministic** — Connect knows exactly which record failed and can route just that one to the DLQ. These are the cleanest DLQ cases. - **put() failures** are trickier. A sink connector receives a **batch**. If it throws a single exception for the batch, Connect can't always tell which record(s) caused it, and a non-retriable batch exception can fail the task. To get clean per-record DLQ routing from put(), the connector must use the **ErrantRecordReporter** (KIP-610) to report individual failed records to the framework, which then logs/DLQs just those. Without that, a batch-level throw may bypass nice per-record tolerance. ## Why this boundary exists The framework can only quarantine what it can attribute to a single record. Infrastructure failures (a dead broker, a bad offset commit) aren't attributable to one record and have no meaningful 'skip and continue' semantics — skipping wouldn't fix them. So Connect deliberately scopes tolerance to deterministic per-record processing. ## Practical implications - Don't expect `errors.tolerance=all` to keep a task alive through broker outages or auth failures — those still kill the task. - For robust sink DLQ behavior, verify the connector supports ErrantRecordReporter; otherwise put() failures may fail the task instead of DLQing cleanly. - Use `errors.retry.timeout` for the genuinely transient slice of put() failures, and DLQ for the deterministic remainder.
- Why doesn't errors.tolerance=all keep a sink task running through a broker outage?Broker/consumer/offset-commit failures are outside the per-record controlled boundary; they aren't attributable to a single bad record and have no skip-and-continue semantics, so the framework fails the task regardless of tolerance.
- What is the ErrantRecordReporter and why does it matter for put() failures?It's the KIP-610 API letting a sink connector report individual failed records to the framework instead of throwing for the whole batch. With it, put() failures get clean per-record DLQ/log/skip handling; without it, a batch-level throw can fail the task instead of quarantining one record.
- Which failure type gives the cleanest DLQ behavior: a converter error or a batch put() error, and why?Converter errors — they're inherently per-record and deterministic, so Connect knows exactly which record failed and routes just that one to the DLQ. Batch put() errors need ErrantRecordReporter to achieve the same granularity.
saying these in an interview costs you the question
- Claiming errors.tolerance=all survives any failure including broker/auth outages
- Saying offset-commit or consumer errors are DLQ-able
- Assuming every sink connector cleanly DLQs put() failures without ErrantRecordReporter
- Thinking tolerance applies to non-per-record/infrastructure errors