skip to content

Error Handling and Dead-Letter Queue

Tolerating bad records with errors.tolerance, a dead-letter topic carrying context headers, and bounded retries. Interviewers ask because the default behaviour is to fail the task on the first bad record.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What does errors.tolerance control in Kafka Connect, and what is the difference between the values none and all?

level: juniorimportance: must knowfreq 78%

answer

  1. none = fail fast (default)
  2. all = skip and continue
  3. only two values
  4. covers converter + SMT + put()
  5. pair with DLQ + log.enable

basics

~10 s

errors.tolerance decides what Connect does when a record fails processing. none (the default) stops the connector task on the first error; all skips the bad record and keeps going.

solid answer

~40 s

errors.tolerance is a connector-level config governing fault tolerance during record processing. The default, none, means zero tolerance: any error in converting, transforming, or handing a record to the connector fails the task immediately, so the task moves to FAILED state and stops. Setting it to all means the task tolerates every such error: the failing record is skipped (optionally logged and/or sent to a dead-letter queue) and processing continues. Only these two values exist. It applies to errors raised during the framework-controlled parts of the pipeline — converters (deserialization/serialization), Single Message Transforms (SMTs), and the sink connector's put() / source poll() handoff. Choosing all trades data completeness for availability, so it is normally paired with errors.deadletterqueue.topic.name and errors.log.enable so skipped records are not silently lost.

go deeper

for a junior

Know the two values: none fails the task, all skips bad records and keeps going; none is the default.

for a middle

Explain which stages it covers (converter, SMT, sink put) and why you pair all with logging and a DLQ.

for a senior

Discuss the availability-vs-completeness tradeoff and the interaction with retry configs (tolerance applies only after retries are exhausted).

for a principal

Frame it as a data-contract decision: when fail-fast protects correctness vs when tolerate-and-quarantine protects pipeline availability, and how to make skipped records auditable and replayable.

## What is Kafka Connect Kafka Connect is a framework for moving data between Apache Kafka and external systems. A **source connector** reads from an external system and writes to Kafka; a **sink connector** reads from Kafka and writes to an external system. Each connector runs as one or more **tasks** (threads/processes doing the actual work). ## The record pipeline For a sink connector, every record goes through framework-controlled stages: (1) the **converter** deserializes the raw bytes from Kafka into a Connect record (e.g. JSON or Avro → an internal `SchemaAndValue`); (2) zero or more **Single Message Transforms (SMTs)** modify the record; (3) the connector's **put()** method hands the batch to the sink task to write downstream. Source connectors are symmetric: poll() produces records, SMTs run, then the converter serializes. Any of these stages can throw. Bad JSON, a schema mismatch, an Avro record whose schema isn't in the registry — all surface as exceptions. ## errors.tolerance This config tells Connect how to react to those errors. It has exactly two legal values: - **none** (default): the first error fails the **task**. The task transitions to the FAILED state and stops processing. This is fail-fast — good when every record matters and you'd rather halt than skip data. - **all**: the task **tolerates** the error. The offending record is skipped and the task continues with the next record. Tolerance here applies to errors raised by converters, SMTs, and the sink `put()` handoff. ## Why pair it with other configs By itself, `errors.tolerance=all` silently drops bad records — dangerous in production. So it is normally combined with: - `errors.log.enable=true` — log each tolerated error. - `errors.deadletterqueue.topic.name=<topic>` — route the failed *original* record to a DLQ topic (sink connectors only). ## Edge cases - `errors.tolerance` only covers the framework-controlled stages. It does NOT make a flaky downstream system magically succeed forever — that's what `errors.retry.timeout` is for, and only after retries are exhausted does tolerance decide skip-vs-fail. - DLQ is available for **sink** connectors only; source connectors can tolerate/log but have no DLQ. - There is no partial / numeric tolerance — it is strictly all-or-nothing.

  • What is the default value of errors.tolerance and why does that default make sense?
    none. It is fail-fast so you never silently lose data; you must consciously opt into tolerating/skipping errors with all.
  • If you set errors.tolerance=all but configure nothing else, what's the operational risk?
    Bad records are silently dropped with no record of them. You should add errors.log.enable=true and, for sinks, a dead-letter-queue topic so failures are observable and recoverable.

saying these in an interview costs you the question

  • Saying errors.tolerance can be set to a number or percentage (only none/all exist)
  • Claiming all retries the record forever (retries are a separate config; tolerance acts after retries are exhausted)
  • Thinking none skips one record then continues — it fails the whole task

context

open as a page

How do you configure a dead-letter queue for a Kafka Connect sink connector, and what does context.headers.enable add?

level: middleimportance: must knowfreq 70%

basics

~10 s

Set errors.tolerance=all plus errors.deadletterqueue.topic.name=<topic>. Failed records are written to that topic. errors.deadletterqueue.context.headers.enable=true adds Kafka headers describing why and where the record failed.

open as a page

What do errors.log.enable and errors.log.include.messages do, and what is the security consideration with the latter?

level: middleimportance: should knowfreq 45%

basics

~10 s

errors.log.enable=true logs each failed record's error context to the Connect worker log. errors.log.include.messages=true additionally logs the record's key/value contents. Including messages can leak sensitive payload data into logs.

open as a page

Which categories of failures does Kafka Connect's error-handling framework tolerate, and which does it NOT? Distinguish conversion/transform/put failures from other error types.

level: seniorimportance: should knowfreq 50%

basics

~20 s

The framework tolerates errors in the parts it controls: converters (deserialization/serialization), SMTs, and the sink connector's put() handoff. It does NOT tolerate failures outside that boundary — like consumer/producer/offset-commit errors or framework/configuration errors, which still fail the task.

open as a page

Explain errors.retry.timeout and errors.retry.delay.max.ms. How do retries interact with errors.tolerance?

level: seniorimportance: should knowfreq 55%

basics

~20 s

errors.retry.timeout is how long Connect keeps retrying a failed operation (default 0 = no retries; -1 = forever). errors.retry.delay.max.ms caps the backoff between attempts (default 60000). Retries run first; only after they're exhausted does errors.tolerance decide skip vs fail.

open as a page