Compare AvroConverter, ProtobufConverter, and JsonConverter, and explain how Schema Registry integration works for the registry-backed converters.
answer
- registry-backed: magic byte + 4-byte schema ID + binary
- schema stored once in registry, not per message
- compatibility check (BACKWARD/FORWARD/FULL) at register
- subject = topic-key / topic-value (TopicNameStrategy)
- JsonConverter != JsonSchemaConverter
basics
~20 sAvroConverter and ProtobufConverter store schemas in Schema Registry and put only a schema ID + binary payload on the topic; JsonConverter embeds or omits the schema inline with no registry. The registry-backed ones give compact messages and enforced compatibility.
solid answer
~50 sAvroConverter (io.confluent.connect.avro.AvroConverter) and ProtobufConverter (io.confluent.connect.protobuf.ProtobufConverter) are Confluent converters that integrate with Schema Registry: they register the Connect schema (translated to Avro/Protobuf) under a subject, and each message on the wire is a compact binary payload prefixed with a magic byte and a 4-byte schema ID — the schema itself is stored once in the registry, not per message. This gives small messages, strong typing, and registry-enforced compatibility (BACKWARD/FORWARD/FULL). They require schema.registry.url. JsonConverter is core Apache Kafka, uses no registry, and either embeds the whole schema per record (schemas.enable=true) or omits it (false) — verbose or lossy. Protobuf differs from Avro in supporting nested message definitions and being widely used with gRPC; Avro is the long-standing Connect default in Confluent stacks. Choosing among them is mostly about message size, type fidelity, compatibility governance, and ecosystem fit. Confluent also offers JsonSchemaConverter (registry-backed JSON Schema), distinct from core JsonConverter.
code
properties · 5 linesvalue.converter=io.confluent.connect.avro.AvroConverter
value.converter.schema.registry.url=http://schema-registry:8081
value.converter.auto.register.schemas=false
value.converter.use.latest.version=true
key.converter=org.apache.kafka.connect.storage.StringConvertergo deeper
Know Avro/Protobuf use a schema registry and are compact, while JSON carries the schema inline or not at all.
Describe the magic-byte + schema-ID wire format and that the registry stores schemas keyed by subject.
Explain registration, compatibility enforcement, deserialization-by-ID, and configs like schema.registry.url / auto.register.schemas.
Set org-wide serialization and compatibility policy, subject naming strategy, and migration between formats with governance.
## The three converters | Converter | Class | Registry? | Wire form | |---|---|---|---| | Avro | `io.confluent.connect.avro.AvroConverter` | Yes | magic byte + 4-byte schema ID + Avro binary | | Protobuf | `io.confluent.connect.protobuf.ProtobufConverter` | Yes | magic byte + schema ID + message indexes + Protobuf binary | | JSON | `org.apache.kafka.connect.json.JsonConverter` | No | JSON (envelope or bare) | (Confluent also ships `io.confluent.connect.json.JsonSchemaConverter`, a *registry-backed* JSON Schema converter — not the same as core JsonConverter.) ## How Schema Registry integration works Schema Registry is a separate service storing versioned schemas, each keyed by a **subject** (by default `<topic>-key` and `<topic>-value` under the TopicNameStrategy) and assigned a globally unique integer **schema ID**. **Serialization (source / produce side):** 1. The converter translates the Connect `Schema` into an Avro/Protobuf schema. 2. It registers that schema under the subject (or looks up the existing ID). The registry runs a **compatibility check** against prior versions; if the new schema violates the configured rule (e.g. BACKWARD), registration is rejected and serialization fails. 3. The registry returns the schema **ID**. 4. The converter writes the wire bytes: a `0x00` magic byte, the 4-byte big-endian schema ID, then the binary-encoded payload. The schema itself is **not** in the message. **Deserialization (sink / consume side):** 1. The converter reads the magic byte + schema ID from the message. 2. It fetches that schema from the registry (cached after first fetch). 3. It decodes the binary payload using that schema and (for Avro) the reader schema, then builds a Connect `Schema`/`Struct`. Configuration: `value.converter.schema.registry.url=http://sr:8081` (and the same with `key.converter.` prefix). Auto-registration can be disabled with `auto.register.schemas=false`, and `use.latest.version=true` forces a fixed reader schema. ## Why registry-backed beats inline JSON for many use cases - **Size:** the schema is stored once, not repeated per record — a few bytes of ID instead of a full schema envelope. - **Type fidelity:** Avro/Protobuf preserve exact types and logical types; bare JSON loses int width and logical types. - **Governance:** the registry enforces compatibility so a producer cannot ship a breaking schema that would crash consumers. ## Avro vs Protobuf nuances - **Avro** has been the de-facto Connect format in Confluent stacks; mature SMT/logical-type support. - **Protobuf** supports nested/`oneof` message structures and meshes with gRPC ecosystems; its wire format also includes message-index bytes to identify which message type within a `.proto` file. ProtobufConverter maps these to Connect Schema/Struct. ## Common pitfalls - Forgetting `schema.registry.url` on a registry-backed converter -> startup/serialization error. - A **sink** using AvroConverter against a topic that was written with plain JSON -> deserialization fails on the magic byte. - Relying on auto-registration in production where a governed, pre-registered schema with strict compatibility is preferred. - Confusing core `JsonConverter` with `JsonSchemaConverter`.
- What exactly is on the wire for an Avro-serialized value?A 1-byte magic byte (0x00), a 4-byte big-endian schema ID, then the Avro-binary-encoded payload. The schema itself lives in the registry, referenced by that ID.
- How does the registry prevent a breaking schema change from being deployed?On registration the registry runs the subject's compatibility check (e.g. BACKWARD) against existing versions; if the new schema is incompatible, registration is rejected and the producing converter fails rather than emitting unreadable data.
- What does the magic byte and schema ID let a sink do?It lets the sink converter look up the writer schema by ID from the registry to correctly decode the binary payload, without the schema being embedded in every message.
saying these in an interview costs you the question
- Saying AvroConverter embeds the full schema in every message (it stores an ID; schema is in the registry)
- Claiming JsonConverter uses Schema Registry
- Forgetting the magic byte / 4-byte schema ID wire prefix
- Asserting Protobuf cannot do nested messages or that Avro/Protobuf skip compatibility checks
- Thinking the schema ID is per-topic rather than globally unique in the registry