skip to content

questions

6

In a shared schema, why is adding an optional field with a default safe for old readers, while promoting a field to required is not?

level: middleimportance: must knowfreq 66%

answer

  1. two directions, not one
  2. neither side deploys first
  3. an unknown entry costs a skip
  4. absence must resolve to something
  5. required is a claim about old messages

basics

~20 s

Adding an optional field only asks an old reader to skip bytes it never needed, and its default gives new readers a value when old writers omit it. Making a field required invalidates every message already written without it.

solid answer

~60 s

An edit is safe when neither side of a running deployment can be handed bytes it cannot interpret, and that is two checks, not one. A new optional field passes both: a reader on the old schema meets an entry it cannot name and steps over it, and a reader on the new schema meets old bytes with the field absent and substitutes the declared default. Requiredness fails the second check. `required` is a claim about *every* message that exists — including messages sitting in a queue, an archive or a retry buffer, and messages that old writers are still producing right now — so the new reader rejects data that was perfectly valid when it was written. The safe way to get the same guarantee is to add the field optional and enforce the invariant in application code, which can tell an old writer apart from a genuinely missing value. Note that a field is safe on the wire well before it is safe to trust: until every writer populates it, a consumer reading it is reading the default.

code

pseudocode · 12 lines
pseudocode
// reader built against schema v1, decoding bytes written by schema v2
for each (tag, value) in message:
    if v1_schema knows tag:
        fields[tag] = value
    else:
        skip(value)                  // the new optional field costs one skip

for each field f declared in v1_schema:
    if f not in fields:
        fields[f] = default_of(f)    // absence resolves; it does not fail

accept(fields)

go deeper

for a junior

Recall the asymmetry: software can ignore a field that did not exist when it was written, but it cannot conjure a field you have started demanding in data that was written before you demanded it.

for a middle

Explain both decode directions concretely — what an old reader does with an entry it cannot name, and what a new reader substitutes when the entry is absent — and say why the declared default is the part that makes the second direction work.

for a senior

Show that you reason about data at rest. Queues, archives, retry buffers and caches hold bytes written under the old schema long after every writer is upgraded, so a requiredness edit keeps failing in production after the rollout is finished.

for a principal

Argue where an invariant belongs. A schema can only say present or absent, while application code can separate an old writer from a genuinely missing value — and that distinction is what lets a contract be tightened without a coordinated flag day.

## What "safe" actually has to mean A schema edit is **safe** when, for the whole period in which two versions of the contract are live, neither side is handed bytes it cannot interpret — and neither side is required to deploy before the other. That is two separate checks, and a reviewer has to run both: 1. **An old reader over new bytes.** Code built against the previous schema meets a message a new writer produced. 2. **A new reader over old bytes.** Code built against the edited schema meets a message an old writer produced — or a message written months ago that is still sitting in a queue, an archive, a retry buffer or a cache. The second check is the one that gets skipped, because it feels like it is about deployment order and it is not. Bytes at rest never upgrade. A message written under the old schema keeps its old shape forever, so "we will have all producers upgraded by Friday" does not retire the old shape — it only stops new instances of it. ## Why adding an optional field passes both checks - **Check 1.** The old reader meets an entry with an identifier it has no declaration for. Any encoding that supports evolution at all lets that reader find where the entry ends — from a length prefix, a type hint, or the delimiters of a text form — so it steps over the entry and carries on. The cost is one skip. - **Check 2.** The new reader meets a message with the field absent. Its own schema declares the field optional with a default, so absence resolves to a defined value instead of an error. - **Neither side must go first.** Producers and consumers can be upgraded in any order, in any mix, which is the practical property the whole exercise is buying. One honest caveat: the field is safe *on the wire* long before it is safe *to depend on*. Until every writer populates it, a consumer that branches on the value is branching on the default, and a default that looks like plausible business data is indistinguishable from real data. ## Why tightening fails, in both directions Requiredness is not a property of the next message; it is a claim about the whole population of messages that exists. Here is the same table a reviewer runs mentally, edit by edit: | Edit to the shared schema | Old reader over new bytes | New reader over old bytes | Verdict | |---|---|---|---| | Add an optional field with an explicit default | skips an entry it cannot name | substitutes the default | safe | | Add an optional field with no default | skips it | falls back to the format's own absence rule | safe on the wire, ambiguous in the contract | | Add a required field | skips it | rejects every older message | breaking | | Promote an existing optional field to required | unaffected | rejects every message that omitted it | breaking | | Relax a required field to optional | rejects new messages that omit it | accepts everything | breaking the other way | That last row is worth staring at. Relaxing a constraint feels generous, but the reader that still holds the *old*, stricter schema is the one that breaks: a new writer omits the field, the old reader demands it, and the message is rejected. Tightening breaks new readers on old data; loosening breaks old readers on new data. Neither direction is free, which is why the additive edit — a field nobody previously demanded — is the only one that is clean on both. ## The default is half the edit **Optional** says the field may be absent. The **default** says what absence *means*. Leave the second out and the meaning is decided by the encoding family rather than by you: some substitute an empty or null-like value, some substitute the type's zero, and some treat a missing declared field as a decode failure. Families genuinely differ here, so an unqualified "it will just be null" is a claim about one ecosystem, not about the edit. Two consequences follow: - **Pick a default that cannot pass for data.** An unset sentinel or a zero is honest; a plausible business value (a currency, a status, a count of one) is silently wrong and unauditable. - **A default is applied by the reader, from the reader's own schema — it is not carried in the bytes.** Changing a default later therefore changes how already-written messages are interpreted, retroactively and everywhere, which makes an innocuous-looking default edit one of the sharper items in this list. ## Getting the guarantee requiredness was for If the business rule really is "this must always be present", the rule still has a home — just not in the schema, not yet: 1. Add the field optional, with a default that means *unset*. 2. Enforce the invariant where the code can tell the two cases apart: an old writer that never knew the field, versus a current writer that omitted a value it should have supplied. 3. Measure. Count arrivals without the field and attribute them to writer generations; the count going to zero is the evidence, and it is evidence about live traffic only. 4. Tighten the schema only when nothing that can still emit the old shape exists — which, once archived data is in scope, often means the schema stays permissive permanently and the application layer carries the rule.

  • If a newly added field has no declared default, what does a reader on the new schema do with a message that predates the field?
    It has nothing to substitute, so the outcome is whatever the encoding family's absence rule is: an empty or null-like value, the type's zero, or a decode failure. Families differ, which is the point — without an explicit default, the meaning of absence is decided by the format instead of by the contract, and every consumer inherits that decision without being told.
  • Why is "optional in the schema but mandatory by policy" not just requiredness under another name?
    Because enforcement moves somewhere that can distinguish the two causes of absence. A schema can only reject the message. Application code can accept it, recognise that it came from a writer predating the field, apply the documented fallback, and count it — and that count is the only honest signal of when the field has actually become universal.
  • Is a message that a new reader rejects for a missing required field lost, or recoverable?
    That depends entirely on where it was rejected, not on the schema: a rejected read from a durable log can be replayed after the schema is loosened again, while a rejected read from a transient stream is gone. Treating the rejection as recoverable is an assumption about the transport, so a reviewer should never let the schema edit rest on it.

saying these in an interview costs you the question

  • Calls an edit safe because the schema file still compiles
  • Adds a required field and plans to backfill producers afterwards
  • Assumes messages already written are re-encoded when the schema changes
  • Leaves a new field with no default and calls absence an error
  • Thinks relaxing a required field to optional is free for old readers
  • Trusts a newly added field before every writer populates it
open as a page

Why can renaming a field be a non-event in one schema family and a breaking change in another, given the same wire bytes?

level: middleimportance: must knowfreq 70%

basics

~20 s

Identity decides. Where a reader matches fields by number, a rename changes only the label. Where it matches by name, a rename is a delete plus an add, so readers on the old schema quietly stop finding the value.

open as a page

Which changes to a scalar field's declared type survive a mixed-version rollout, and which corrupt values without raising an error?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Only a widening whose on-wire byte form is unchanged survives, and only while writers stay inside the old range. Narrowing, signed-to-unsigned reinterpretation and any edit that changes the byte form corrupt values silently rather than failing.

open as a page

When a field is deleted from a tag-numbered schema, why must its number and its name be reserved rather than freed for reuse?

level: seniorimportance: should knowfreq 60%

basics

~20 s

A retired number and name still live in deployed readers and in bytes already written. Reserving them stops a later edit from handing an old reader a different field under an identifier it already believes it understands.

open as a page

When a shared schema's field must change type, how do you decide between editing it in place and adding a replacement field beside the retired original?

level: principalimportance: should knowfreq 38%

basics

~20 s

Price a one-time risk against a permanent cost. An in-place edit is available only when the layout is unchanged, the direction is safe and the readers are enumerable; a replacement field burns identifiers and adds precedence logic forever, but makes the ambiguity visible per message.

open as a page

What happens when a writer starts sending an enumerated value that a reader on the earlier schema has no name for?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

The value usually arrives intact, so the decode is not where it breaks. It breaks in application logic that branches over the members it knows and has no arm for this one — an additive schema edit that is not additive for consumers.

open as a page