skip to content

Converters and Schema Handling

Converters that map Connect's internal Schema and Struct model to Avro, JSON, Protobuf or raw bytes. Interviewers ask because a converter mismatch is the number-one Connect pipeline failure.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

6

What is a converter in Kafka Connect, and what is the difference between key.converter and value.converter?

level: juniorimportance: must knowfreq 80%

answer

  1. Bridge between byte[] and Connect Schema/Struct
  2. key.converter vs value.converter independent
  3. worker default, connector override
  4. fromConnectData / toConnectData
  5. not a connector, not an SMT

basics

~10 s

A converter serializes/deserializes Connect data to/from bytes on the Kafka topic. key.converter handles the record key; value.converter handles the record value. They are configured independently and can differ.

solid answer

~40 s

A converter is a pluggable component (implementing org.apache.kafka.connect.storage.Converter) that translates between Connect's internal in-memory representation (Schema + Struct/primitive) and the raw byte[] stored in Kafka. For a source connector it serializes Connect data to bytes before producing; for a sink connector it deserializes bytes from Kafka into Connect data before the connector processes them. key.converter is applied to the record key and value.converter to the value, so you can, for example, use StringConverter for keys and AvroConverter for values. Converters are set at the worker level (defaults for all connectors) and can be overridden per connector. Common implementations: AvroConverter, JsonConverter, ProtobufConverter, StringConverter, ByteArrayConverter. The converter is distinct from a connector and from transforms (SMTs).

go deeper

for a junior

Know that a converter turns bytes into typed data and back, and that key/value have separate converter settings.

for a middle

Explain the source vs sink direction, worker-vs-connector override, and that converter sits at the byte boundary before SMTs.

for a senior

Discuss the Converter interface, schemas.enable / schema-registry configs, and how a converter mismatch causes deserialization failures.

for a principal

Reason about org-wide converter standardization, schema-registry strategy, and migration paths between serialization formats.

## The problem converters solve Kafka stores everything as raw bytes (`byte[]`) — it has no concept of types or schemas. Kafka Connect, however, works with a typed in-memory model: a `Schema` plus either a `Struct` (for structured records) or a Java primitive. A **converter** is the bridge between these two worlds. - **Sink flow (Kafka -> connector):** bytes are read from the topic, the converter *deserializes* them into a Connect `SchemaAndValue`, and the sink connector writes that to the external system. - **Source flow (connector -> Kafka):** the source connector produces Connect records, the converter *serializes* the Connect data into bytes, and those bytes are produced to the topic. ## Key vs value converter A Kafka record has a key and a value, each independently serialized. Connect exposes two settings: - `key.converter` — applied to the record key - `value.converter` — applied to the record value They are fully independent. A very common pattern is `key.converter=org.apache.kafka.connect.storage.StringConverter` (keys are often simple IDs or strings) and `value.converter=io.confluent.connect.avro.AvroConverter` (values are rich structured payloads). There is also a third, `header.converter`, for record headers (default `SimpleHeaderConverter`). ## Where converters are configured - **Worker level** (in the worker properties / distributed worker config): sets the *defaults* for every connector on that worker. - **Connector level** (in the connector's JSON config): overrides the worker default for just that connector. ## The interface Converters implement `org.apache.kafka.connect.storage.Converter` with two core methods: `fromConnectData(topic, schema, value) -> byte[]` and `toConnectData(topic, byte[]) -> SchemaAndValue`. Schema-aware converters also implement `ConverterConfig` handling (e.g. `schemas.enable` for JsonConverter, `schema.registry.url` for Avro/Protobuf). ## What a converter is NOT - It is **not** a connector (the connector talks to the external system). - It is **not** a Single Message Transform (SMT) — SMTs operate on the already-deserialized Connect record, *between* the converter and the connector. Getting key/value converter mismatches right is one of the most common operational issues in Connect: if a sink's `value.converter` does not match how the data was actually serialized, deserialization fails.

  • Can the key and value use different converters?
    Yes. They are configured separately, so e.g. StringConverter for the key and AvroConverter for the value is a common and valid combination.
  • Where in the pipeline does the converter sit relative to SMTs?
    On the sink side: bytes -> converter (deserialize) -> SMTs -> connector. On the source side: connector -> SMTs -> converter (serialize) -> bytes. The converter is always at the byte boundary; SMTs operate on the typed Connect record.

saying these in an interview costs you the question

  • Saying the converter talks to the external system (that's the connector)
  • Claiming key and value must use the same converter
  • Confusing converters with SMTs/transforms
  • Thinking converters are configured only at the worker level and cannot be overridden per connector

context

open as a page

Explain JsonConverter and the schemas.enable setting. What changes in the on-wire payload when it is true vs false?

level: middleimportance: must knowfreq 75%

basics

~10 s

JsonConverter serializes Connect data as JSON. With schemas.enable=true it wraps the payload as {"schema":...,"payload":...} so the schema travels inline. With false it emits just the raw JSON value with no schema.

open as a page

Compare AvroConverter, ProtobufConverter, and JsonConverter, and explain how Schema Registry integration works for the registry-backed converters.

level: seniorimportance: must knowfreq 70%

basics

~20 s

AvroConverter and ProtobufConverter store schemas in Schema Registry and put only a schema ID + binary payload on the topic; JsonConverter embeds or omits the schema inline with no registry. The registry-backed ones give compact messages and enforced compatibility.

open as a page

Describe Connect's internal data model — Schema and Struct — and how converters relate to it.

level: middleimportance: should knowfreq 55%

basics

~10 s

Connect represents data internally with a Schema (the type/shape) and a Struct (the values), independent of any wire format. Converters translate between this Schema/Struct model and the raw bytes on the topic.

open as a page

When would you use StringConverter, ByteArrayConverter, and a header.converter? What are their constraints?

level: middleimportance: should knowfreq 45%

basics

~10 s

StringConverter treats data as plain text (UTF-8 by default). ByteArrayConverter passes raw bytes through untouched (no schema). header.converter serializes record headers separately, defaulting to SimpleHeaderConverter.

open as a page

A sink connector throws DataException / SerializationException on deserialize. How do you diagnose and resolve converter and schema mismatches?

level: seniorimportance: should knowfreq 55%

basics

~20 s

The sink's converter usually doesn't match how the data was actually written. Identify the real on-topic format, align the sink's key/value converter and settings (e.g. schemas.enable, schema.registry.url) to it, and use error tolerance / a DLQ for poison records.

open as a page