skip to content

When a field is deleted from a tag-numbered schema, why must its number and its name be reserved rather than freed for reuse?

level: seniorimportance: should knowfreq 60%

answer

  1. deleting edits the file, not the fleet
  2. the number is the contract's key
  3. old bytes outlive the field
  4. a recycled tag decodes silently
  5. retire the name as well

basics

~20 s

A retired number and name still live in deployed readers and in bytes already written. Reserving them stops a later edit from handing an old reader a different field under an identifier it already believes it understands.

solid answer

~50 s

Deleting a field edits the current schema; it does not edit the readers already running, and it certainly does not edit the messages already written. In a tag-numbered family the number *is* the field's identity, so if a later edit hands that number to a new field, a reader still on the old schema does not see a new field — it sees the old one. Where the new value's wire shape still fits the old declaration, it decodes cleanly into the wrong field and no error is raised anywhere; where the shapes differ, you get a loud failure instead, which is better but still an outage. Reserving the number makes that mistake impossible to re-make by a reviewer who was not present for the deletion, because the schema tool rejects the reuse. The **name** is reserved for the same reason at a different layer: generated accessors, text projections of the same schema and name-keyed sinks all bind on the name, so a recycled name rebinds them silently.

code

pseudocode · 11 lines
pseudocode
// tag 7 held "quantity" (integer) in v1, was deleted in v2,
// and a later edit reused 7 for "note" (text) in v3

writer_v3:   emit(tag = 7, bytes_of_text("backordered"))

reader_v1 sees tag 7 and consults its OWN schema:
    declared_type = integer             // v1 was never told the field retired
    if wire_shape(bytes) fits integer:
        quantity = as_integer(bytes)    // wrong number, no error raised
    else:
        fail_or_skip(bytes)             // loud only when the shapes differ

go deeper

for a junior

Recall the rule and its reason: a field identifier that has been used is retired forever, because software built before the deletion still believes that identifier means the old field.

for a middle

Explain what the identifier does — it is how a decoder decides which declaration a run of bytes belongs to — and therefore why recycling it makes an old reader misread new data rather than reject it.

for a senior

Demonstrate the diagnosis. A recycled identifier produces plausible values, no errors and no alerts, so you should be able to say what evidence would surface it: a value distribution that shifts at a deploy boundary, or a consumer whose totals diverge from the producer's.

for a principal

Set the standing rule and make it machine-checkable. Identifiers append-only, reservations required in the same change as the deletion, no renumbering — enforced by the tool rather than by whoever remembers the outage.

## The delete that is only half a delete Removing a field from a schema file removes it from *that file*. Three things survive the edit and none of them are in your working tree: - **Readers already deployed**, built against the schema as it was, still holding a declaration for the field. - **Messages already written** — in queues, logs, archives, caches and retry buffers — still carrying the field's bytes. - **Generated code and downstream projections** built from the old schema, still exposing the field by name. The reserve-the-identifier rule exists because of the first two. It is not bookkeeping hygiene; it is the only thing standing between a future edit and a category of outage that produces no error at all. ## The identifier is the contract's primary key Every evolvable encoding needs a way to say *which* field a run of bytes belongs to, and there are two families: - **Tag-numbered.** Each field carries a small integer on the wire. The number is the identity; the human-readable name exists only in the schema and in generated code. - **Name-keyed.** Each field carries its name, or the reader matches the writer's field names against its own. The name is the identity. In both cases the identifier behaves like a primary key that has been handed out to every reader ever built. Reusing a primary key for a different row is a familiar disaster; reusing a field identifier for a different field is the same disaster, with the twist that the affected readers are not in your repository and cannot be updated in the same change. ## The failure mode is silence Walk an old reader over bytes from a schema that recycled a number. The reader reads the identifier, consults **its own** schema — which was never told the field retired — and decodes according to the declaration it holds. What happens next depends only on whether the new value's on-wire shape happens to fit the old declaration: | Later edit after deleting field 7 | What a reader on the old schema does | Verdict | |---|---|---| | Number reserved, name reserved | sees no entry 7; its absence rule applies | safe | | Number recycled, new value's wire shape fits the old type | decodes the new value as the old field | silent corruption | | Number recycled, wire shapes differ | fails the decode, or skips, depending on the family | loud outage | | Number reserved, name recycled | binary readers are fine; name-keyed readers and generated code rebind | silent in text projections | The third row is the lucky one. A count that suddenly reads as a nonsense integer because a text value was laid down under its number will be carried into a total, a threshold or a bill long before anybody notices, and the data it corrupts is not recoverable from the message, since the message is exactly what the writer intended to send. The subtler trap is recycling a number for a field of *the same type*. Nothing in the bytes is anomalous, no shape check can fire, and the old reader gets a well-formed value of the right type under the wrong meaning. That is the worst case in the table, not the mildest. ## Why the name is retired too In a tag-numbered family it is tempting to reserve the number and let the name go. Three surfaces disagree: - **Generated accessors** in every consumer carry the name. A recycled name compiles against the new field while the surrounding code still means the old one. - **Text projections** of the same schema — a human-readable rendering of the same messages, a debug dump, an export — key on the name rather than the number, so they rebind silently. - **Name-keyed sinks** downstream: columns, index fields, dashboard series and stored documents that were populated from the field's name and now receive a different field's values under it. So a deletion retires both halves of the identity, and it does so *declaratively*, inside the schema, where the tooling can enforce it. A comment saying "do not reuse 7" is not a reservation; it is a hope about who reads comments. ## Reviewing a deletion 1. **Establish that nothing still reads the field.** The schema cannot tell you this. It is an empirical question about consumers and stored data, and being unable to answer it is itself a reason to stop. 2. **Delete the declaration and reserve both the number and the name** in the same change, so the reservation cannot be forgotten in a follow-up that never happens. 3. **Check what the value's disappearance means to readers that still declare it.** They now apply their absence rule to a field that used to be populated, which is a semantic change even though the decode is clean. 4. **Never renumber the remaining fields.** Compacting a schema's numbers to close the gap re-identifies every field after the hole, which is the same outage repeated once per field. The standing rule that falls out of all this is short: **identifiers are append-only**. A schema's numbers and names may be retired, never recycled — the space is cheap and the collision is not.

  • Why can recycling a number for a field of the same type be worse than recycling it for a different type?
    Because the only accident that saves you is a shape mismatch. A different type may not fit the old declaration, so the decode fails loudly and somebody investigates. The same type always fits: the old reader gets a well-formed value of exactly the type it expected, under a meaning that changed, and nothing anywhere has a reason to complain.
  • In a family where fields are matched by name rather than by number, what is the equivalent of reserving a tag?
    Reserving the name. The rule is identical once you ask what the reader matches on: the retired identifier must never be rebound to a new field, a new type or a new meaning. Name-keyed families make this easier to get wrong, because a name looks like documentation rather than like a key, and reusing a descriptive word feels harmless.
  • Does reserving the identifier make a deletion safe on its own?
    No. Reservation protects future edits from re-identifying the field; it says nothing about consumers that still read the value today. Those readers now apply their absence rule to something that used to be populated, which is a semantic change with a clean decode — the least visible kind. Establishing that nobody depends on the value is a separate, empirical step.

Reissuing a retired room number in a building whose old floor plans are still in circulation. Deliveries do not fail; they arrive, confidently, at the wrong door.

saying these in an interview costs you the question

  • Believes a deleted field's number is free for the next field
  • Expects a recycled identifier to fail loudly rather than decode
  • Reserves the number but reuses the name in a text projection
  • Assumes no consumer reads the field once it is deleted
  • Treats a comment in the schema as a reservation
  • Renumbers the remaining fields to close the gap