How do you configure a dead-letter queue for a Kafka Connect sink connector, and what does context.headers.enable add?
answer
- tolerance=all is mandatory for DLQ
- deadletterqueue.topic.name
- raw original bytes written
- context.headers.enable → __connect.errors.* headers
- sink-only, replication.factor default 3
basics
~10 sSet errors.tolerance=all plus errors.deadletterqueue.topic.name=<topic>. Failed records are written to that topic. errors.deadletterqueue.context.headers.enable=true adds Kafka headers describing why and where the record failed.
solid answer
~50 sA DLQ in Connect is a regular Kafka topic that receives records a sink connector could not process. To enable it you must set errors.tolerance=all (the DLQ is ignored under none, which just fails the task) and errors.deadletterqueue.topic.name to the target topic. The connector writes the *original raw* key/value bytes of the failed record there — so DLQ records are not re-deserialized by the failing converter. Because it auto-create may not exist, you also set errors.deadletterqueue.topic.replication.factor (default 3) to match your cluster. Turning on errors.deadletterqueue.context.headers.enable=true enriches each DLQ record with Kafka headers (prefixed __connect.errors.) carrying diagnostic context: the original topic, partition, offset, the connector/task that failed, the stage and exception class, and the stack trace. This metadata is what makes a DLQ actionable — you can locate the source record and understand the failure without it. DLQ is sink-only; source connectors have no DLQ.
go deeper
Know that a DLQ is a Kafka topic for failed records, enabled by tolerance=all + deadletterqueue.topic.name.
Configure it fully including replication.factor and explain that context.headers.enable adds diagnostic headers.
Explain raw-bytes preservation enabling replay, the __connect.errors.* header set, and DLQ-write failure modes.
Design an end-to-end quarantine-and-replay strategy: DLQ provisioning/ACLs, monitoring DLQ lag, automated reprocessing, and alerting on DLQ growth.
## What a dead-letter queue is A **dead-letter queue (DLQ)** is a holding area for messages that could not be processed. In Kafka Connect the DLQ is itself just an ordinary Kafka **topic**. Instead of throwing away (or failing on) a record a **sink** connector can't handle, Connect republishes that record to the DLQ topic so it can be inspected and reprocessed later. ## Minimum configuration The DLQ only activates when the connector is allowed to tolerate errors: ``` errors.tolerance=all errors.deadletterqueue.topic.name=my-connector-dlq ``` With `errors.tolerance=none` (the default), a failure fails the task and the DLQ name is ignored entirely. You typically also set: ``` errors.deadletterqueue.topic.replication.factor=3 ``` Default is 3; lower it (e.g. to 1) for single-broker/dev clusters or the DLQ topic creation fails. ## What gets written Connect writes the **original raw bytes** of the failed record's key and value to the DLQ — the pre-conversion payload. This matters: if the converter is what failed, re-serializing isn't possible, so the unmodified source bytes are preserved. That lets you fix the converter/schema and replay the DLQ topic through a corrected connector. ## context.headers.enable By default the DLQ record carries the payload but no explanation. Setting: ``` errors.deadletterqueue.context.headers.enable=true ``` adds **Kafka record headers** (key prefix `__connect.errors.`) to every DLQ message. These include: the original `topic`, `partition`, and `offset`; the `connector.name`, `task.id`; the processing `stage` (e.g. VALUE_CONVERTER, TRANSFORMATION, PUT) and `class.name` of the component; and the exception's `exception.class.name`, `exception.message`, and `exception.stacktrace`. Without these headers a DLQ is a graveyard of mystery bytes; with them, each record is self-describing — you know exactly which source offset failed, in which stage, and why. ## Edge cases and gotchas - **Sink-only**: there is no DLQ for source connectors. - **Single DLQ topic per connector**, shared by all tasks of that connector. - A record that fails because the **sink put()** rejected it (not just conversion) is also DLQ-eligible under tolerance=all. - DLQ writes themselves can fail (e.g. topic missing, ACL denied) — provision and authorize the topic deliberately. - The DLQ producer uses the worker's producer config; secure clusters need appropriate write ACLs.
- Why must errors.tolerance be all for the DLQ to work?Under none the task fails on the first error and never reaches the routing logic; the DLQ name is simply ignored. The DLQ is the 'where do skipped records go' destination, and records are only skipped when tolerance=all.
- What header information does context.headers.enable add and why is it valuable?Headers (prefixed __connect.errors.) carrying original topic/partition/offset, connector/task, failing stage, and exception class/message/stacktrace — making each DLQ record self-describing so you can locate the source record and diagnose the failure.
saying these in an interview costs you the question
- Saying the DLQ works with errors.tolerance=none
- Claiming the DLQ stores the *transformed* record (it stores the original raw bytes)
- Asserting source connectors support DLQ (they don't)
- Forgetting replication.factor must fit the cluster or topic creation fails