skip to content

What does errors.tolerance control in Kafka Connect, and what is the difference between the values none and all?

level: juniorimportance: must knowfreq 78%

answer

  1. none = fail fast (default)
  2. all = skip and continue
  3. only two values
  4. covers converter + SMT + put()
  5. pair with DLQ + log.enable

basics

~10 s

errors.tolerance decides what Connect does when a record fails processing. none (the default) stops the connector task on the first error; all skips the bad record and keeps going.

solid answer

~40 s

errors.tolerance is a connector-level config governing fault tolerance during record processing. The default, none, means zero tolerance: any error in converting, transforming, or handing a record to the connector fails the task immediately, so the task moves to FAILED state and stops. Setting it to all means the task tolerates every such error: the failing record is skipped (optionally logged and/or sent to a dead-letter queue) and processing continues. Only these two values exist. It applies to errors raised during the framework-controlled parts of the pipeline — converters (deserialization/serialization), Single Message Transforms (SMTs), and the sink connector's put() / source poll() handoff. Choosing all trades data completeness for availability, so it is normally paired with errors.deadletterqueue.topic.name and errors.log.enable so skipped records are not silently lost.

go deeper

for a junior

Know the two values: none fails the task, all skips bad records and keeps going; none is the default.

for a middle

Explain which stages it covers (converter, SMT, sink put) and why you pair all with logging and a DLQ.

for a senior

Discuss the availability-vs-completeness tradeoff and the interaction with retry configs (tolerance applies only after retries are exhausted).

for a principal

Frame it as a data-contract decision: when fail-fast protects correctness vs when tolerate-and-quarantine protects pipeline availability, and how to make skipped records auditable and replayable.

## What is Kafka Connect Kafka Connect is a framework for moving data between Apache Kafka and external systems. A **source connector** reads from an external system and writes to Kafka; a **sink connector** reads from Kafka and writes to an external system. Each connector runs as one or more **tasks** (threads/processes doing the actual work). ## The record pipeline For a sink connector, every record goes through framework-controlled stages: (1) the **converter** deserializes the raw bytes from Kafka into a Connect record (e.g. JSON or Avro → an internal `SchemaAndValue`); (2) zero or more **Single Message Transforms (SMTs)** modify the record; (3) the connector's **put()** method hands the batch to the sink task to write downstream. Source connectors are symmetric: poll() produces records, SMTs run, then the converter serializes. Any of these stages can throw. Bad JSON, a schema mismatch, an Avro record whose schema isn't in the registry — all surface as exceptions. ## errors.tolerance This config tells Connect how to react to those errors. It has exactly two legal values: - **none** (default): the first error fails the **task**. The task transitions to the FAILED state and stops processing. This is fail-fast — good when every record matters and you'd rather halt than skip data. - **all**: the task **tolerates** the error. The offending record is skipped and the task continues with the next record. Tolerance here applies to errors raised by converters, SMTs, and the sink `put()` handoff. ## Why pair it with other configs By itself, `errors.tolerance=all` silently drops bad records — dangerous in production. So it is normally combined with: - `errors.log.enable=true` — log each tolerated error. - `errors.deadletterqueue.topic.name=<topic>` — route the failed *original* record to a DLQ topic (sink connectors only). ## Edge cases - `errors.tolerance` only covers the framework-controlled stages. It does NOT make a flaky downstream system magically succeed forever — that's what `errors.retry.timeout` is for, and only after retries are exhausted does tolerance decide skip-vs-fail. - DLQ is available for **sink** connectors only; source connectors can tolerate/log but have no DLQ. - There is no partial / numeric tolerance — it is strictly all-or-nothing.

  • What is the default value of errors.tolerance and why does that default make sense?
    none. It is fail-fast so you never silently lose data; you must consciously opt into tolerating/skipping errors with all.
  • If you set errors.tolerance=all but configure nothing else, what's the operational risk?
    Bad records are silently dropped with no record of them. You should add errors.log.enable=true and, for sinks, a dead-letter-queue topic so failures are observable and recoverable.

saying these in an interview costs you the question

  • Saying errors.tolerance can be set to a number or percentage (only none/all exist)
  • Claiming all retries the record forever (retries are a separate config; tolerance acts after retries are exhausted)
  • Thinking none skips one record then continues — it fails the whole task

context