skip to content

Schemas and Serialization

How bytes on a topic get meaning: serdes and the wire format, the Schema Registry, Avro and Protobuf, and compatible schema evolution. Interviewers care because topics outlive the services that write to them.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

explore

questions

page 2 of 2

Mechanically, when and how does Schema Registry enforce a compatibility check, and how do you configure or test it via the API/CLI?

level: seniorimportance: should knowfreq 45%

basics

~20 s

When a new schema version is registered for a subject, the registry checks it against prior version(s) per the subject's mode and returns HTTP 409 if incompatible. You set the mode with PUT /config/<subject> and can dry-run a check with the /compatibility endpoint.

open as a page

How do you secure access to a Schema Registry, and why does registry authentication matter for deserialization safety?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Protect the registry with TLS plus authentication — HTTP Basic auth or mutual TLS — and restrict who can register schemas. It matters because the schema the consumer fetches is itself input; if an attacker can spoof or alter it, they can manipulate how payloads are decoded.

open as a page

After quarantining poison pills to a DLQ, how do you keep offset progress and replay/reprocessing safe and idempotent?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Only commit the offset after the DLQ publish succeeds, so a crash mid-quarantine re-tries instead of skipping silently. Make handling deterministic per record and DLQ writes idempotent/keyed, so replaying the same bad record from an earlier offset produces the same quarantine, not duplicates or new gaps.

open as a page

How is Protobuf oneof represented, and what are the compatibility rules when evolving a oneof or adding/removing fields in Protobuf and JSON Schema?

level: seniorimportance: should knowfreq 30%

basics

~20 s

A Protobuf oneof groups fields where at most one is set; setting one clears the others. Adding a new field to a oneof is generally backward/forward compatible, but moving an existing field into or out of a oneof, or removing a field, can break compatibility. JSON Schema models oneof via oneOf/anyOf with $ref.

open as a page

How do .proto imports and JSON Schema $ref translate to Schema Registry references? How do you register a schema that depends on another?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Each imported .proto or $ref'd JSON Schema is registered as its own subject/version, then the dependent schema declares schema references: a list of {name, subject, version} so the registry can resolve the import. Register the dependencies first, then the parent with its references array.

open as a page

How do you govern access to Schema Registry — what controls exist for subjects, contexts, and import/export operations?

level: seniorimportance: should knowfreq 22%

basics

~20 s

Schema Registry access is governed by ACLs/RBAC scoped to subjects (and contexts). You grant operations like READ, WRITE, and global SUBJECT_COMPATIBILITY/MODE on specific subject (or context) resource patterns, so producers can register schemas, consumers can only read, and only privileged roles can change compatibility, mode, or run import/export.

open as a page

What is a schema context in Confluent Schema Registry, and how do qualified subjects work?

level: seniorimportance: should knowfreq 30%

basics

~20 s

A schema context is an independent namespace inside one Schema Registry, with its own set of subjects and schema IDs. You address a schema in a non-default context using a qualified subject like :.contextname:subjectname, letting separate environments or imported schemas coexist without ID/subject collisions.

open as a page

How does schema linking work between two Schema Registries, and what role do exporters and IMPORT mode play?

level: seniorimportance: should knowfreq 25%

basics

~20 s

Schema linking continuously replicates schemas from a source registry to a destination by configuring an exporter (a schema link). The exported schemas land in a context on the destination, which runs in IMPORT mode so it can register them with the source's exact IDs and versions, keeping IDs stable for replicated data.

open as a page

How does Schema Registry store its data durably, and what is the role of the _schemas topic?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Schema Registry stores all schemas in a special Kafka topic called _schemas. It is a single-partition, log-compacted topic that acts as a commit log; each registry node reads it into an in-memory cache, so the registry itself holds no separate database.

open as a page

How would you implement a custom Serializer/Deserializer, and what production concerns must it handle?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Implement Serializer<T>/Deserializer<T>, putting your encoding in serialize()/deserialize(). Handle null in and out, make it thread-safe and stateless, honor the isKey flag in configure(), version your format for evolution, and fail with SerializationException — never crash the whole consumer.

open as a page

In Kafka Streams, what is a Serde and how does the Serdes factory class fit in?

level: seniorimportance: should knowfreq 55%

basics

~20 s

A Serde<T> bundles a Serializer<T> and a Deserializer<T> into one object, because Streams both reads and writes data. The Serdes factory class provides ready-made ones, e.g. Serdes.String(), Serdes.Long(), and you set defaults with default.key.serde / default.value.serde.

open as a page

You need to publish several related event types (OrderCreated, OrderShipped, OrderCancelled) to a single topic to preserve their ordering. How do you configure subject naming and serialization to make this work?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Switch off the default per-topic subject by setting value.subject.name.strategy to RecordNameStrategy or TopicRecordNameStrategy, so each event type gets its own subject. Also disable auto-union validation issues by using a schema that allows the multiple types (e.g. an Avro union).

open as a page

Design question: across many teams with long-retention and compacted topics, how would you choose between non-transitive and transitive compatibility modes, and what failure does the wrong choice cause?

level: principalimportance: should knowfreq 35%

basics

~20 s

Use a _TRANSITIVE mode when old-schema data persists in topics (long retention or compaction), so every schema is checked against all prior versions. Non-transitive only guarantees adjacent versions, so version 3 might fail to read version 1's data even though each step passed.

open as a page

Design a hardened consumer-side deserialization strategy for a service that processes events from a partially-trusted set of producers. What controls would you put in place?

level: principalimportance: should knowfreq 35%

basics

~20 s

Use a schema-validated, data-only format; allowlist only the types/subjects you expect; authenticate producers and the registry; bound payload size and depth; and wrap deserialization so a bad record fails safely (poison-pill handling) instead of crashing the consumer.

open as a page

How would you design a poison-pill strategy that distinguishes truly corrupt records from transient/recoverable deserialization failures, and routes each correctly?

level: principalimportance: should knowfreq 25%

basics

~20 s

Classify the failure: genuinely corrupt or wrong-format bytes are permanent — skip to DLQ immediately. Failures from a transiently unavailable Schema Registry or an unregistered-but-fixable schema are recoverable — retry with backoff and do NOT discard, because retrying corrupt bytes wastes effort and discarding recoverable ones causes false data loss.

open as a page

As a platform architect standardizing subject naming across many teams, what are the trade-offs of RecordNameStrategy (sharing) versus TopicRecordNameStrategy (isolation), and how would you govern them?

level: principalimportance: should knowfreq 30%

basics

~20 s

Sharing (RecordNameStrategy) gives one canonical type definition reused everywhere but a large blast radius for changes and unclear ownership. Isolation (TopicRecordNameStrategy) limits blast radius per topic but allows drift and many subjects. Govern with ownership, compatibility policy, and CI registration.

open as a page

What is an Avro schema fingerprint, and how does it differ from the Schema Registry's schema ID?

level: principalimportance: nice to knowfreq 25%

basics

~20 s

A schema fingerprint is a deterministic hash of a schema's canonical form, the same everywhere for identical schemas. A registry schema ID is a small integer the Schema Registry assigns per schema; it's registry-local, not a hash.

open as a page

What are Data Contracts in Confluent Schema Registry, and how do CSFLE and migration/data rules fit in?

level: principalimportance: nice to knowfreq 20%

basics

~20 s

A Data Contract enriches a schema with metadata, tags, and rules. Rules include data quality/transformation rules, migration rules (to transform data across incompatible major versions), and encryption rules — CSFLE (Client-Side Field Level Encryption) encrypts tagged sensitive fields at the producer before they ever hit the broker.

open as a page

showing 31–48 of 48