Why can renaming a field be a non-event in one schema family and a breaking change in another, given the same wire bytes?
answer
- ask what the reader matches on
- label versus identity
- bytes unchanged, lookup broken
- a rename is delete plus add
- the loss is silent, not an error
basics
~20 sIdentity decides. Where a reader matches fields by number, a rename changes only the label. Where it matches by name, a rename is a delete plus an add, so readers on the old schema quietly stop finding the value.
solid answer
~50 sThe question a reviewer has to ask first is *what does the reader match on*. In a tag-numbered family the number is the identity and the name is a label for humans and code generators, so renaming leaves the bytes and the match untouched — the wire is a non-event. In a name-keyed family the name **is** the identity, so a rename is two edits at once: the old field is deleted and a new one is added. A reader still on the old schema finds nothing under the old name and applies its absence rule, which means a default, not an error. That is silent data loss — the worst failure shape available. Even in the numbered case a rename is not free: generated accessors, text projections of the same schema, and any name-keyed sink downstream all bind on the name. So the honest answer is "it depends on what identifies a field", followed by the list of places where the name is load-bearing anyway.
go deeper
Recall that a field's human-readable name and its identity are not always the same thing, and that whether a rename is safe depends on which of the two the decoder actually uses.
Explain both families side by side and be specific about the failure shape: under name-keyed identity, the old reader does not error, it finds nothing and falls back — silent loss rather than a visible break.
Show the consequences above the wire. Even a number-keyed rename is a coordinated change across generated code, text projections and name-keyed sinks, and you should be able to list those before approving the diff.
Take the position that identity is a decision, not a detail: choosing a family whose identity is a stable number buys cheap relabelling forever, and the rule that identifiers are append-only is what keeps that property from eroding.
## The only question that matters: what is the key? A field has two things that look like a name: a **human label** and an **identity**. Whether renaming is safe turns entirely on whether those are the same thing in the family you are using. - **Tag-numbered families** put a small integer on the wire for each field. The decoder matches on that integer. The label lives in the schema source and in generated code, and never crosses the wire. - **Name-keyed families** put the name on the wire, or match the writer's declared names against the reader's. The decoder matches on the name; there is no other identity to fall back on. So the same edit — `qty` becomes `quantity` — is a comment change in one family and a delete-plus-add in the other, with identical bytes on the wire in the first case and a different lookup in the second. | | Tag-numbered schema | Name-keyed schema | |---|---|---| | What the decoder matches on | the field's number | the field's name | | Old reader over new bytes after a rename | finds the field, decodes normally | finds nothing under the old name | | What the old reader reports | nothing; the rename is invisible | no error — it applies its absence rule | | Practical verdict | safe on the wire, not free in code | breaking, and silent | ## Why the name-keyed case is worse than it looks A rename in a name-keyed family is *two* edits fused into one line of a pull request: the old field is deleted, and a new field is added. Of the two halves, the delete is the one that hurts, and it hurts immediately: - The **delete half** takes effect the moment a renamed writer ships, against every reader already deployed. Those readers do not fail; they see an absent field, substitute a default or a null-like value, and carry on producing output that looks fine. - The **add half** is inert. No deployed reader knows the new name, so nothing starts working until consumers are rebuilt. That asymmetry is why a rename under name-keyed identity behaves as an outage with no alert. Nothing in the pipeline distinguishes "this field was absent" from "this field was renamed out from under me" — absence is absence. ## Where a "safe" rename still bites Even where the number carries identity and the wire genuinely does not notice, the label is load-bearing in at least four places: - **Generated accessors.** Consumer code built from the schema exposes the field under its name. A rename breaks that build — loudly, which is the good case, but it still couples a schema edit to every consumer's release. - **Text projections of the same schema.** A human-readable rendering, a debug dump or an export usually keys on names, so anything consuming that form is in the name-keyed world even though the binary form is not. - **Name-keyed sinks.** Columns, index fields, stored documents and dashboard series populated by field name go on receiving values under the old name until somebody rebuilds them, or silently stop. - **Anything that stored the name.** Configuration that refers to a field path, a mapping table, a rule engine, a query, a saved filter — none of these are rebuilt by a schema compiler. The net advice for a review: a rename is cheap on the wire only in one family, and it is never zero-cost above the wire. ## The reviewer's routine 1. **Ask what identifies a field here.** If nobody in the room can answer, that is the finding; stop there. 2. **If the name is the identity, do not rename.** Treat it as what it is — add the new field, keep populating the old one while readers still ask for it, then retire the old name and reserve it. The cost is visible, which is the point. 3. **If the number is the identity, still enumerate the name's consumers.** Generated code, text forms, name-keyed sinks, stored references. The rename is safe for the bytes and a coordinated change everywhere else. 4. **Never pair a rename with a renumber.** Two identity edits in one change make the resulting breakage impossible to attribute. A final subtlety worth carrying: renaming and re-meaning are different edits that look alike in a diff. Changing `amount` to `amount_minor_units` is a label edit only if the values were always minor units. If the rename documents a *change* in what the field carries, then no identity rule saves you — old readers match the field perfectly and misinterpret every value, which is the silent-corruption case in a different costume.
- In a tag-numbered schema, name a place where a rename still breaks something.Anywhere the name is the key above the wire: generated accessors in consumer code, a text projection of the same schema, name-keyed sinks such as columns or index fields, and stored references like a mapping rule or a saved query. The binary decode is untouched; everything built from the label is a coordinated change.
- If a rename in a name-keyed family is really a delete plus an add, which half hurts first?The delete. It takes effect against every already-deployed reader the moment a renamed writer ships, and those readers report nothing — they just see an absent field and substitute their fallback. The add half is inert until consumers are rebuilt, so the damage window is entirely one-sided.
- How is renaming a field different from changing what it means under the same name?Renaming is an identity question; re-meaning is not. If a field keeps its identifier but starts carrying different units, a different scale or a different scope, every reader matches it perfectly and misinterprets every value. No identity rule protects against that, which is why it is the more dangerous of the two edits despite the smaller diff.
saying these in an interview costs you the question
- Says a rename is always safe because the number is unchanged
- Says a rename always breaks, without asking what identifies a field
- Expects a missing name to raise an error rather than an absence
- Forgets generated accessors and text projections carry the name
- Frees the old name for reuse right after renaming
- Treats a re-meaning as a rename because the diff looks the same