skip to content

When choosing between Avro, Protobuf, and JSON Schema for a message registry, what actually differs about how each handles schema evolution, and what should drive the choice?

level: principalimportance: nice to knowfreq 35%

answer

  1. Avro: name-based resolution, writer+reader schema, compact but opaque without schema
  2. Protobuf: numbered tags on the wire, rename-safe, tag reuse is the real danger
  3. JSON Schema: self-describing, debuggable, weaker native compat-check tooling
  4. proto3 fields implicitly optional with defaults
  5. choice driven by ecosystem fit, not a universal 'best'

basics

~30 s

All three can express similar rules for adding/removing fields safely, but they differ in how evolution is tracked: Avro relies on a full schema document and reader/writer resolution, Protobuf uses stable numbered field tags baked into the format itself, and JSON Schema is more flexible but has weaker native tooling for enforcing evolution rules. The right pick depends on your ecosystem, tooling, and how much schema strictness you actually want.

solid answer

~1 min

Avro requires both the writer's and reader's full schema to resolve a message (fields matched by name, with defaults filling gaps), which makes it compact and precise about evolution rules but means you can't parse an Avro binary payload without the exact writer schema, obtained from the registry. Protobuf identifies fields by explicit numbered tags embedded in the wire format itself, not by name, so a field can be renamed freely without breaking wire compatibility, evolution is largely about never reusing or changing a tag number and treating new fields as inherently optional; this makes Protobuf naturally forgiving of renames in a way Avro isn't, at the cost of the registry doing less of the compatibility enforcement (much of it is a Protobuf-language-level discipline). JSON Schema describes plain JSON, so payloads are human-readable and debuggable without any registry lookup, but JSON Schema's validation-oriented design (keywords like additionalProperties, required) doesn't map as cleanly onto registry-style BACKWARD/FORWARD/FULL semantics, and tooling support for automatic compatibility checking is generally less mature than Avro's in the Confluent ecosystem. The choice should be driven by ecosystem fit (Avro is the Kafka-native default with the most mature registry tooling; Protobuf wins when you also need gRPC or cross-language strongly-typed clients and want rename-friendly evolution; JSON Schema wins when human-readability/debuggability and loose coupling to a specific serialization stack matter more than compactness or strict typed evolution).

go deeper

for a junior

Should know these are three different ways to describe message structure and that they're not interchangeable at the wire-format level.

for a middle

Should be able to state that Avro needs the schema to decode while Protobuf and JSON are more self-contained, at a basic level.

for a senior

Should explain the name-based-resolution versus numbered-tag mechanics precisely enough to say why renames behave differently in Avro versus Protobuf.

for a principal

Should make and justify an organization-level format choice given ecosystem constraints (existing gRPC investment, polyglot consumers, partner-facing feeds, tooling maturity) and anticipate the evolution failure modes specific to the chosen format.

## Why the mechanics matter more than a feature list The three formats encode 'what changed is safe' in fundamentally different ways, and understanding that mechanical difference is what lets you predict how each behaves under evolution pressure rather than just memorizing a feature comparison table. ## Avro — resolution by name Avro's model is **schema resolution by name**, with two schemas always in play: - the writer's schema (whatever the producer used when it serialized a specific message), and - the reader's schema (whatever the consumer's code currently expects). Avro binary encoding carries no field names or tags at all in the bytes themselves, just values in the order the writer's schema defines, which is what makes it extremely compact but also means the bytes are completely uninterpretable without the exact writer schema. At read time, Avro's resolution algorithm: 1. walks both schemas, 2. matches fields by name, 3. copies over values present in both, 4. fills reader-only fields from their declared default, 5. and ignores writer-only fields the reader doesn't ask for. This name-based matching is precisely why Avro treats a field rename as a breaking-ish change: renaming `total` to `totalAmount` means the reader looking for `totalAmount` doesn't find a matching name in the writer's schema at all, effectively treating it as if the old field was removed and an unrelated new one added, which is exactly the two-separate-edits problem discussed under compatibility modes generally. ## Protobuf — numbered tags on the wire Protobuf's model is different at the wire level in a way that changes what 'safe evolution' even means. Every field in a `.proto` message definition is assigned an explicit **numeric tag** (e.g., `string order_id = 1;`), and that tag number, not the field name, is what's actually encoded on the wire alongside each value. This means a field can be renamed in the `.proto` source at will, `order_id` to `orderId`, with zero effect on wire compatibility, because renaming doesn't touch the tag number old and new binaries agree on. Protobuf's evolution discipline is therefore centered on **tag stewardship**: - never reuse a tag number that was ever assigned to a removed field (Protobuf even has a `reserved` keyword specifically to prevent accidental reuse), and - understand that in proto3, all fields are implicitly optional with a well-defined default for their type, so adding a new field is inherently safe for both old and new readers in a way that doesn't require a registry to adjudicate the way Avro's default-value bookkeeping does. This makes Protobuf naturally friendlier to renames and arguably requires less registry-side sophistication to get safe evolution, since a lot of the safety is a property of the wire format and language convention rather than something a compatibility mode has to actively check. ## JSON Schema — plain text, validation-shaped JSON Schema takes a third approach: the payload itself is plain, self-describing JSON, readable and debuggable with zero tooling, `curl` and your eyes are enough to understand a message. But JSON Schema as a specification was designed primarily for validation (does this document satisfy these constraints), not for the reader/writer resolution model Avro and Protobuf are built around, so concepts like 'a reader schema can safely read data from an older writer schema' don't map onto JSON Schema's native vocabulary as cleanly; `required`, `additionalProperties`, and type keywords express constraints, but automated BACKWARD/FORWARD/FULL-style compatibility checking has to be layered on top by the registry implementation rather than falling naturally out of the format's own resolution semantics. Confluent Schema Registry does support JSON Schema with compatibility checking, but the ecosystem's maturity and community tooling around it (schema evolution linters, code generators, examples) generally lags Avro's, which has been the de facto Kafka-native default for the longest and has the deepest tooling investment. ## The trade-off surface The trade-off surface, then, isn't 'which is best' but 'which failure mode and which coupling do you want.' | Format | Strength | Price | |---|---|---| | **Avro** | gives compact wire format and precise, registry-enforceable compatibility semantics | at the cost of payloads being opaque without registry access and renames being structurally awkward | | **Protobuf** | gives rename-friendliness, strong multi-language code generation (especially valuable if you're also using gRPC and want one schema definition serving both request/response and event payloads), and wire-format-level safety guarantees | at the cost of evolution discipline living partly in convention (tag stewardship) rather than being fully registry-enforced, and of Protobuf being a slightly heavier tooling investment if your stack wasn't already using it for RPC | | **JSON Schema** | gives maximum debuggability and the lowest barrier to entry (any language can read JSON without a codegen step or SDK) | at the cost of larger message size (field names repeated in every message, no binary compaction), and weaker off-the-shelf compatibility enforcement compared to Avro's mature registry integration | ## Where it shows up A concrete real-world scenario: - A company already running a gRPC-based synchronous service mesh, where every internal API is defined in `.proto` files and code-generated clients exist in Go, Java, and Python, extends the same `.proto` definitions to asynchronous Kafka events, reusing the exact same message types for both RPC and event payloads and getting consistent, rename-tolerant evolution across both without maintaining two parallel schema ecosystems. - A different company with a polyglot, loosely-coupled set of services (including some written by less Kafka-experienced teams, or with external partners consuming a public-facing event feed) chooses JSON Schema specifically so that any partner can read a sample message and understand the contract without adopting Avro or Protobuf tooling, accepting the larger message size and lighter-weight compatibility enforcement as the cost of that broad accessibility.

  • Why does Protobuf's `reserved` keyword matter for safe schema evolution?
    It prevents a future engineer from accidentally reusing a numeric tag that used to belong to a now-removed field; because Protobuf identifies fields by tag number on the wire, reusing a retired tag for a new, differently-typed field would cause old messages (or old binaries) to misinterpret bytes as the wrong field, a silent and hard-to-diagnose form of data corruption that reserving the tag number explicitly prevents at compile time.
  • Why is an Avro binary payload essentially unreadable without the registry, while a Protobuf or JSON payload is comparatively more recoverable?
    Avro binary encoding stores no field names or tags at all, just raw values in schema-defined order, so the exact writer schema is mandatory to make any sense of the bytes. Protobuf embeds numeric tags in the wire format itself, so with just the .proto definitions (which are often checked into a shared repo rather than requiring a live registry call) you can decode a message, and JSON Schema payloads are plain JSON text, readable without any schema at all, just less strictly validated.
  • In what situation would you deliberately choose JSON Schema over Avro despite its larger message size?
    When broad, low-friction accessibility matters more than wire efficiency, for example a public or partner-facing event feed where external consumers may not want to adopt Avro tooling or a Kafka-specific SDK, or an early-stage internal system where engineers value being able to read raw messages directly during debugging without registry access, and the extra bytes per message are not a meaningful cost at the system's actual throughput.

Avro is like a form where blanks are filled in strictly by matching the label printed on both the old and new form; Protobuf is like a form where each blank has a fixed serial number stapled to it so relabeling the blank doesn't confuse anyone filling it in; JSON Schema is like a form written in plain prose that anyone can read without instructions, but that's harder to mechanically double-check for consistency.

saying these in an interview costs you the question

  • Claims one format is unconditionally 'better' with no situational reasoning
  • Doesn't know Protobuf identifies fields by numbered tag rather than by name on the wire
  • Thinks Avro binary payloads are self-describing without needing the writer's schema
  • Cannot explain why Avro treats field renames as effectively breaking while Protobuf tolerates them
  • Assumes JSON Schema gets the exact same automatic BACKWARD/FORWARD/FULL enforcement maturity as Avro in every registry implementation

context