skip to content

Explain Avro's writer schema vs reader schema and how schema resolution works during deserialization.

level: seniorimportance: must knowfreq 60%

answer

  1. writer = encoded bytes, reader = consumer wants
  2. match fields by NAME not position
  3. missing-in-writer → reader default
  4. extra-in-writer → discarded
  5. int→long→float→double promotions

basics

~20 s

The writer schema is the one used to encode the bytes; the reader schema is the one the consumer wants to decode into. Avro resolves the two by matching fields by name, applying defaults for missing fields and dropping unknown ones.

solid answer

~50 s

Every Avro datum is encoded with the **writer schema** — the exact schema the producer used. The consumer decodes with a **reader schema** — the schema its code expects. Avro performs **schema resolution**: it pairs the two schemas, matching record fields **by name** (not position). If the reader has a field the writer omitted, Avro fills it from the reader's **default**; if the writer has a field the reader lacks, Avro skips it; numeric promotions (int→long, float→double) are allowed. On Kafka, the `KafkaAvroDeserializer` always knows the writer schema by resolving the message's schema ID against the registry. The reader schema comes from the generated SpecificRecord class (or, for GenericRecord, defaults to the writer schema). This writer/reader separation is the mechanism behind backward/forward compatibility: a consumer can keep reading old data, and read new data, as long as additive changes carry defaults.

go deeper

for a junior

Know there are two schemas — one to write, one to read — and Avro bridges them.

for a middle

State the core resolution rules: match by name, defaults fill missing fields, extras are dropped.

for a senior

Connect writer/reader resolution to backward/forward compatibility and the registry's role in supplying the writer schema by ID.

for a principal

Design evolution policies (defaults, aliases, enum defaults) so all current and future writer/reader pairs resolve cleanly across many services.

Avro's superpower is that the schema used to **write** data and the schema used to **read** it do not have to be identical — they only have to be *compatible*. This is **schema resolution**. **Definitions:** - **Writer schema**: the schema the producer used to serialize the bytes. It is fixed at write time and travels with the data (on Kafka, via the schema ID embedded in the message). - **Reader schema**: the schema the consuming application expects. It comes from the consumer's own code/version. **Why both are needed:** to decode Avro binary you *must* have the writer schema (the bytes are just values in writer-schema order, with no field names). But the consumer may have evolved to a newer/older schema version. Avro reconciles the two. **Resolution rules (the important ones):** 1. **Records match by name**, recursively, **not by position**. So field order can differ between writer and reader. 2. **Field in reader but not writer** → Avro uses the reader field's `default`. If there is no default, resolution **fails**. 3. **Field in writer but not reader** → Avro reads and **discards** it. 4. **Type promotion** is allowed in a fixed set: int→long→float→double, string↔bytes. 5. **Enums/unions/aliases**: unknown enum symbols can fall back to an enum `default`; `aliases` let a renamed field/record still match. **On Kafka specifically:** the `KafkaAvroDeserializer` extracts the 4-byte **schema ID** from the message, calls the Schema Registry to get the **writer schema**, and caches it. The **reader schema** is determined by config: - With `specific.avro.reader=true`, the reader schema is the schema baked into the generated **SpecificRecord** class on the consumer's classpath. Avro resolves writer→reader, so the consumer can be on a different schema version than the producer. - With GenericRecord (default), there is effectively no separate reader schema — the writer schema is used directly, so you always see exactly what was written. **Edge cases:** - If a producer adds a **required** field (no default) and an old consumer's reader schema lacks it — that's fine (reader ignores it). But if a *consumer* upgrades to a reader schema with a new required field and reads *old* data lacking it, resolution fails unless that field has a default. This is why **adding fields with defaults** is the safe evolution pattern. - Resolution is per-pair; the registry's compatibility checks (BACKWARD etc.) exist precisely to pre-validate that future writer/reader pairs will resolve. **Mental model:** the writer schema says "here is what the bytes mean"; the reader schema says "here is what I want." Avro is the translator that maps one onto the other field-by-field, using defaults to fill gaps.

  • If the reader schema adds a field with no default and reads old data missing it, what happens?
    Schema resolution fails with an error — Avro has nothing to fill the field with. That's why additive fields must carry a default to be backward-compatible.
  • Does Avro match record fields by position or by name?
    By name (recursively). This is why field order can differ between writer and reader and why renames require aliases.
  • With GenericRecord (no specific.avro.reader), what is the reader schema?
    There's effectively none separate — the writer schema is used directly, so you decode exactly what was written.

saying these in an interview costs you the question

  • Saying Avro matches fields by position/index like Protobuf tag numbers
  • Claiming the consumer must use the exact same schema version as the producer
  • Thinking adding a field without a default is always safe

context