How do inline field names, numeric tags, and purely positional fields differ in what the wire carries and what a rename costs?
answer
- how does the reader know which field
- name, number, or place in line
- the key costs bytes every record
- identity moves, breakage moves with it
- a spent tag number is never reused
basics
~20 sInline names put identity in every record, so a rename changes the bytes and breaks readers. Numeric tags carry identity in one or two bytes and leave names free to change. Positional fields carry no identity, so order becomes the contract.
solid answer
~50 sThree strategies exist for answering "which field is this?" on the wire. **Inline names** repeat the key with every value: readable, self-sufficient, and the most expensive per record — and because the name *is* the identity, renaming a field is a wire-breaking change for every reader. **Numeric tags** replace the name with a small number, typically one to two bytes for low field numbers, plus a marker that lets an unrecognised field be stepped over; identity is now the number, so a rename is free on the wire, though generated accessors and downstream column names still move. **Positional fields** carry nothing: values are laid out in the writer's declared order, which makes the layout minimal and makes *order itself* part of the contract — inserting a field in the middle silently reinterprets everything after it. The design question is not which is best, it is which mistake you would rather be able to make.
code
pseudocode · 11 lines// tag-keyed decode: identity is a number, never a name
while bytes remain:
tag = readVarint() // which field, as a number
length = readLength() // how far this value runs
raw = readBytes(length)
if tag in readerSchema.byTag:
field = readerSchema.byTag[tag]
record[field.name] = parse(raw, field.type) // name comes from the schema
else:
discard(raw) // unknown tag: length already told us how fargo deeper
Recall that a field can be identified on the wire by its name, by a number, or by nothing but its place, and that repeating names costs bytes in every single record.
Explain the overhead and the breakage of each strategy, and in particular why moving identity from a name to a number changes what a rename costs.
Demonstrate the operational habits that follow: reserving spent tag numbers, refusing mid-record insertions in positional layouts, and treating a rename as a contract change wherever names are on the wire.
Frame it as buying the mistakes you can survive: name-carrying wire makes renames loud and bytes expensive, positional wire makes bytes free and ordering mistakes silent and unrecoverable.
## Three ways to say "which field is this?" Every encoding has to answer one question for each value it writes: **how will the reader know which field this is?** There are only three answers in practice, and the choice between them is what the self-describing / schema-driven split is made of. - **By name.** The key travels beside the value in every record. Identity is a string. - **By tag.** A small integer travels beside the value. Identity is a number, and the schema maps it to a name. - **By position.** Nothing travels. Identity is the value's ordinal place in the record, fixed by the writer's schema. ## What each costs on the wire | Strategy | Identity carried by | Overhead per field | Unknown field can be skipped | Order is part of the contract | |---|---|---|---|---| | Inline names | the key string, repeated | the key plus punctuation, every record | yes, structure is inline | no | | Numeric tags | a small integer | one to two bytes for low numbers, plus a marker | yes, the marker gives its length | no | | Positional | nothing | none | no, there is nothing to skip past | yes, absolutely | The overhead column is why the argument is loud at volume. A field named `transactionTimestamp` costs more bytes in its key than most values cost in total, and that cost is paid on **every record**, forever, not once per record type. ## What a rename does in each scheme 1. **Inline names.** The bytes change. A reader that matches on the old key stops finding the field and falls back to whatever it does with an absent value — usually silently. This is a wire-breaking change dressed up as a refactor, and it is the single most common way a "cosmetic" pull request takes down a consumer. 2. **Numeric tags.** The bytes do not change. Old and new readers both match on the tag and keep working, which is why tag-keyed encodings are so often chosen where producers and consumers upgrade at different speeds. It is not *costless*: generated accessors, log field names, and any column derived from the name still move, so the change has a blast radius above the wire. 3. **Positional.** The name was never on the wire, so a rename is invisible — but the same property means a **reorder or an insertion** is catastrophic and undetectable. A reader given a record with one extra field in the middle does not fail; it reads the next value as the wrong field, with whatever garbage that implies. ## The thing you must never reuse Moving identity onto a number creates a new invariant: **a tag number, once used, is spent**. If field `7` is retired and later reassigned to a different field, old bytes still contain a value under tag `7`, and a new reader will decode that value into the new field. Nothing errors. The type marker may even match. This is why mature tag-keyed contracts carry a graveyard of reserved numbers that nobody may claim, and why "we renumbered the fields to tidy them up" is a sentence that should stop a review. ## Where a positional layout is still the right answer Positional encodings look indefensible in isolation, and they are still everywhere, for one reason: **they are the floor**. When both ends are generated from the same schema, deployed together, and the record is small and enormously repeated, the bytes you do not spend on identity are pure gain — no key, no tag, no marker, just the values. The trade is explicit rather than naive: you are declaring that the writer's schema will always be available to the reader and that field order will be managed as a contract. Where those two statements are true, position is the cheapest identity there is. Where either is doubtful — an archive, an external consumer, a long retention horizon — it is the most dangerous.
- If names are off the wire, what is the thing a tag-keyed contract must never do?Reuse a retired tag number. Old bytes still carry values under that number, so a reader built from the new schema decodes an old value straight into the new field, with no error and possibly a matching type marker. Retired numbers are reserved permanently rather than recycled, which is the price of having moved identity onto a small integer.
- Under positional encoding, why is inserting a field worse than appending one?Appending puts the new value after everything an old reader knows about, so the old reader reads its fields correctly and simply stops early. Inserting shifts every subsequent value one place, so each is decoded as its neighbour. Because nothing identifies a field, the reader has no way to notice; it returns confidently wrong values rather than failing.
saying these in an interview costs you the question
- Says numeric tags are only a size trick with no effect on renames.
- Believes renaming a field is always a cosmetic, wire-neutral change.
- Assumes a positional reader can detect that the writer inserted a field.
- Claims retired tag numbers can be reassigned once no writer uses them.
- Treats field order in a tag-keyed stream as part of the contract.