Schemas and Serialization
How bytes on a topic get meaning: serdes and the wire format, the Schema Registry, Avro and Protobuf, and compatible schema evolution. Interviewers care because topics outlive the services that write to them.
part ofApache Kafkaoverview, primer and where to startread it →on this pageshowhide
explore
- Confluent Schema Registry6 questions
- Avro With Kafka5 questions
- Protobuf and JSON Schema Support5 questions
- Schema Evolution and Compatibility Modes5 questions
- Subject Naming Strategies5 questions
- Schema References, Contexts and Governance6 questions
questions
page 2 of 2Mechanically, when and how does Schema Registry enforce a compatibility check, and how do you configure or test it via the API/CLI?
basics
~20 sWhen a new schema version is registered for a subject, the registry checks it against prior version(s) per the subject's mode and returns HTTP 409 if incompatible. You set the mode with PUT /config/<subject> and can dry-run a check with the /compatibility endpoint.
How do you secure access to a Schema Registry, and why does registry authentication matter for deserialization safety?
basics
~20 sProtect the registry with TLS plus authentication — HTTP Basic auth or mutual TLS — and restrict who can register schemas. It matters because the schema the consumer fetches is itself input; if an attacker can spoof or alter it, they can manipulate how payloads are decoded.
After quarantining poison pills to a DLQ, how do you keep offset progress and replay/reprocessing safe and idempotent?
basics
~20 sOnly commit the offset after the DLQ publish succeeds, so a crash mid-quarantine re-tries instead of skipping silently. Make handling deterministic per record and DLQ writes idempotent/keyed, so replaying the same bad record from an earlier offset produces the same quarantine, not duplicates or new gaps.
How is Protobuf oneof represented, and what are the compatibility rules when evolving a oneof or adding/removing fields in Protobuf and JSON Schema?
basics
~20 sA Protobuf oneof groups fields where at most one is set; setting one clears the others. Adding a new field to a oneof is generally backward/forward compatible, but moving an existing field into or out of a oneof, or removing a field, can break compatibility. JSON Schema models oneof via oneOf/anyOf with $ref.
How do .proto imports and JSON Schema $ref translate to Schema Registry references? How do you register a schema that depends on another?
basics
~20 sEach imported .proto or $ref'd JSON Schema is registered as its own subject/version, then the dependent schema declares schema references: a list of {name, subject, version} so the registry can resolve the import. Register the dependencies first, then the parent with its references array.
How do you govern access to Schema Registry — what controls exist for subjects, contexts, and import/export operations?
basics
~20 sSchema Registry access is governed by ACLs/RBAC scoped to subjects (and contexts). You grant operations like READ, WRITE, and global SUBJECT_COMPATIBILITY/MODE on specific subject (or context) resource patterns, so producers can register schemas, consumers can only read, and only privileged roles can change compatibility, mode, or run import/export.
What is a schema context in Confluent Schema Registry, and how do qualified subjects work?
basics
~20 sA schema context is an independent namespace inside one Schema Registry, with its own set of subjects and schema IDs. You address a schema in a non-default context using a qualified subject like :.contextname:subjectname, letting separate environments or imported schemas coexist without ID/subject collisions.
How does schema linking work between two Schema Registries, and what role do exporters and IMPORT mode play?
basics
~20 sSchema linking continuously replicates schemas from a source registry to a destination by configuring an exporter (a schema link). The exported schemas land in a context on the destination, which runs in IMPORT mode so it can register them with the source's exact IDs and versions, keeping IDs stable for replicated data.
How does Schema Registry store its data durably, and what is the role of the _schemas topic?
basics
~20 sSchema Registry stores all schemas in a special Kafka topic called _schemas. It is a single-partition, log-compacted topic that acts as a commit log; each registry node reads it into an in-memory cache, so the registry itself holds no separate database.
How would you implement a custom Serializer/Deserializer, and what production concerns must it handle?
basics
~20 sImplement Serializer<T>/Deserializer<T>, putting your encoding in serialize()/deserialize(). Handle null in and out, make it thread-safe and stateless, honor the isKey flag in configure(), version your format for evolution, and fail with SerializationException — never crash the whole consumer.
In Kafka Streams, what is a Serde and how does the Serdes factory class fit in?
basics
~20 sA Serde<T> bundles a Serializer<T> and a Deserializer<T> into one object, because Streams both reads and writes data. The Serdes factory class provides ready-made ones, e.g. Serdes.String(), Serdes.Long(), and you set defaults with default.key.serde / default.value.serde.
You need to publish several related event types (OrderCreated, OrderShipped, OrderCancelled) to a single topic to preserve their ordering. How do you configure subject naming and serialization to make this work?
basics
~20 sSwitch off the default per-topic subject by setting value.subject.name.strategy to RecordNameStrategy or TopicRecordNameStrategy, so each event type gets its own subject. Also disable auto-union validation issues by using a schema that allows the multiple types (e.g. an Avro union).
Design question: across many teams with long-retention and compacted topics, how would you choose between non-transitive and transitive compatibility modes, and what failure does the wrong choice cause?
basics
~20 sUse a _TRANSITIVE mode when old-schema data persists in topics (long retention or compaction), so every schema is checked against all prior versions. Non-transitive only guarantees adjacent versions, so version 3 might fail to read version 1's data even though each step passed.
Design a hardened consumer-side deserialization strategy for a service that processes events from a partially-trusted set of producers. What controls would you put in place?
basics
~20 sUse a schema-validated, data-only format; allowlist only the types/subjects you expect; authenticate producers and the registry; bound payload size and depth; and wrap deserialization so a bad record fails safely (poison-pill handling) instead of crashing the consumer.
How would you design a poison-pill strategy that distinguishes truly corrupt records from transient/recoverable deserialization failures, and routes each correctly?
basics
~20 sClassify the failure: genuinely corrupt or wrong-format bytes are permanent — skip to DLQ immediately. Failures from a transiently unavailable Schema Registry or an unregistered-but-fixable schema are recoverable — retry with backoff and do NOT discard, because retrying corrupt bytes wastes effort and discarding recoverable ones causes false data loss.
As a platform architect standardizing subject naming across many teams, what are the trade-offs of RecordNameStrategy (sharing) versus TopicRecordNameStrategy (isolation), and how would you govern them?
basics
~20 sSharing (RecordNameStrategy) gives one canonical type definition reused everywhere but a large blast radius for changes and unclear ownership. Isolation (TopicRecordNameStrategy) limits blast radius per topic but allows drift and many subjects. Govern with ownership, compatibility policy, and CI registration.
What is an Avro schema fingerprint, and how does it differ from the Schema Registry's schema ID?
basics
~20 sA schema fingerprint is a deterministic hash of a schema's canonical form, the same everywhere for identical schemas. A registry schema ID is a small integer the Schema Registry assigns per schema; it's registry-local, not a hash.
What are Data Contracts in Confluent Schema Registry, and how do CSFLE and migration/data rules fit in?
basics
~20 sA Data Contract enriches a schema with metadata, tags, and rules. Rules include data quality/transformation rules, migration rules (to transform data across incompatible major versions), and encryption rules — CSFLE (Client-Side Field Level Encryption) encrypts tagged sensitive fields at the producer before they ever hit the broker.
showing 31–48 of 48