skip to content

Confluent Schema Registry supports Avro, Protobuf, and JSON Schema. How do you produce Protobuf or JSON Schema messages, and how does the registry know which format a schema is?

level: juniorimportance: must knowfreq 60%

answer

  1. schemaType: AVRO | PROTOBUF | JSON
  2. Pick format = pick serializer class
  3. magic byte 0x00 + 4-byte schema ID header
  4. Multi-format since CP 5.5
  5. AVRO is the default schemaType

basics

~10 s

Use the format-specific serializer: KafkaProtobufSerializer or KafkaJsonSchemaSerializer instead of KafkaAvroSerializer. Each subject in the registry carries a schemaType (AVRO, PROTOBUF, or JSON) so the registry knows which format the stored schema is.

solid answer

~40 s

Confluent Schema Registry is multi-format: each registered schema record has a schemaType field that is AVRO (default), PROTOBUF, or JSON. You pick the format on the client by choosing the matching (de)serializer pair: KafkaProtobufSerializer/Deserializer or KafkaJsonSchemaSerializer/Deserializer, configured with schema.registry.url just like Avro. The serializer auto-registers (or looks up) the schema under the subject derived by the subject-name strategy (default TopicNameStrategy -> <topic>-value), tagging it with the right schemaType. All three formats share the same 5-byte wire-format header (magic byte 0x00 + 4-byte schema ID), then format-specific payload bytes follow. A registry can hold different formats on different subjects simultaneously; compatibility checking is enforced per-format.

go deeper

for a junior

Know the three serializer classes and that schemaType marks the format.

for a middle

Explain the shared magic-byte + schema-ID header and per-subject schemaType.

for a senior

Discuss subject-name strategy choices and per-format compatibility enforcement.

for a principal

Reason about governance across a multi-format registry: standardizing formats per domain, compatibility policy, and migration constraints between formats.

## Background Confluent **Schema Registry** is a service that stores schemas and assigns each one a globally unique integer **schema ID**. Producers serialize records against a schema and prepend the schema's ID to the bytes; consumers read the ID, fetch the schema, and deserialize. Originally Avro-only, since Confluent Platform 5.5 the registry is **multi-format**, supporting Avro, **Protobuf**, and **JSON Schema**. ## How the registry distinguishes formats Every stored schema has a **`schemaType`** attribute with one of three values: `AVRO` (the default if omitted, for backward compatibility), `PROTOBUF`, or `JSON`. This is part of the registration request (`POST /subjects/<subject>/versions` with a JSON body containing `schema`, `schemaType`, and optional `references`). So the *registry* knows a schema's format from this stored field — not from inspecting the bytes. ## How the client picks the format You choose the format on the **client side** by selecting the matching serializer/deserializer: - Avro: `KafkaAvroSerializer` / `KafkaAvroDeserializer` - Protobuf: `KafkaProtobufSerializer` / `KafkaProtobufDeserializer` - JSON Schema: `KafkaJsonSchemaSerializer` / `KafkaJsonSchemaDeserializer` All are configured the same way — at minimum `schema.registry.url`, and typically `auto.register.schemas` (default true) and the subject-name strategy. The serializer either **auto-registers** the schema (in dev) or looks it up, then stamps the message with the resulting schema ID and the correct `schemaType`. ## Shared wire format Regardless of format, the on-wire layout starts with the same 5-byte header: a **magic byte `0x00`** followed by a **4-byte big-endian schema ID**. After that the payload differs: - Avro: binary Avro body. - Protobuf: a **message-index** varint array, then the Protobuf binary body (covered in the dedicated question). - JSON Schema: UTF-8 JSON text body. ## Subject-name strategy The **subject** under which a schema is registered is computed by the configured strategy. Default is `TopicNameStrategy` (`<topic>-value` / `<topic>-key`). `RecordNameStrategy` and `TopicRecordNameStrategy` are also available and are commonly used with Protobuf/JSON when one topic carries multiple record types. ## Edge cases - A single registry instance happily holds different formats across different subjects. - **Compatibility checking is per-format** and per-subject; you cannot evolve an Avro subject into a Protobuf subject. - The default global compatibility level is `BACKWARD`.

  • What is the default schemaType if you register a schema without specifying one?
    AVRO — the field defaults to AVRO for backward compatibility with the original Avro-only registry.
  • Can one topic carry both Protobuf and JSON Schema messages on the same value subject?
    No — a single subject has one schemaType and an evolution history within that format. Different formats need different subjects, and consumers must use the matching deserializer. You could mix record types within one Protobuf subject using RecordNameStrategy, but not mix formats.

saying these in an interview costs you the question

  • Saying the registry inspects the bytes to detect the format — it relies on the stored schemaType field, and the client picks via the serializer class.
  • Claiming Protobuf/JSON need a different registry than Avro — it is the same multi-format registry.
  • Thinking the wire header differs by format — the 5-byte magic+ID header is identical; only the payload after it differs.

context