skip to content

In schema-driven binary encodings, what must reach the decoder under tag-numbered identity versus writer-schema resolution?

level: middleimportance: must knowfreq 55%

answer

  1. one schema, or two
  2. a number agreed, or a name matched
  3. nothing travels, or an identifier does
  4. build-time binding versus read-time pairing
  5. the writer's exact layout must be reachable

basics

~20 s

Tag-numbered identity needs only the reader's own compiled schema, because the numbers in the bytes are the field identities. Writer-schema resolution additionally needs the exact schema the writer used, reachable by an identifier that travels with the data.

solid answer

~50 s

This family splits two ways. In the **tag-numbered** style, each field is declared with a number, the bytes carry that number, and a reader that compiled its own copy of the definition can decode alone — nothing about the writer's version has to reach it. In the **writer-schema** style, no numbers are used: fields are matched by name, and decoding requires *two* schemas — the one the writer actually used and the one the reader expects — paired up before any value is read. The writer's schema therefore has to be obtainable, either written once into the file that holds the records or identified per message so the reader can fetch and cache it. The first style optimises for point-to-point links where both ends are deployed from the same definition; the second optimises for data whose readers were not built alongside its writer.

code

pseudocode · 10 lines
pseudocode
# binding A: tag-numbered identity — the reader decodes alone
reader_schema = compiled_into_this_binary()
record = decode(bytes, using = reader_schema)   # numbers in the bytes are identity

# binding B: writer-schema resolution — two schemas are paired first
writer_schema = fetch_schema(bytes.schema_id)   # id rode with the bytes
if writer_schema is missing:
    stop("cannot decode: the writer's exact schema is unreachable")
plan   = pair(writer_schema, reader_schema)     # fields matched by name
record = decode(bytes, using = plan)

go deeper

for a junior

Hold on to the headline: in one style the message carries numbers and the reader needs only its own definition; in the other the reader must also get hold of the exact definition the writer used.

for a middle

Explain what has to be in the decoder's hands in each case, and name the three ways a writer's schema is made reachable: a file header, a per-message identifier resolved from a store, or an out-of-band agreement.

for a senior

Show that you have operated both. A per-message identifier puts a resolvable store on the read path, and you should be able to say what your consumers do when that lookup fails and how caching keeps it off the hot path.

for a principal

Decide per hop rather than per company: synchronous links that deploy together want build-time binding, while data that outlives its producer wants bytes that carry their own explanation.

## Two ways to bind bytes to meaning Every encoding in this family shares the premise that the data's shape is declared outside the data. What they do **not** share is how a particular stream of bytes is connected to the declaration that explains it. There are two answers, and the operational difference between them is one of the most commonly probed points in this material. **Tag-numbered identity.** Each field is given a number when it is declared. The writer emits the number with the value; the reader looks the number up in the definition it compiled. Field identity is therefore a small integer agreed in advance, and a reader needs exactly one schema — its own — to decode. **Writer-schema resolution.** No numbers are declared. Fields are identified by name, and the decoder is handed two schemas: the **writer's schema**, describing precisely how these bytes were laid out, and the **reader's schema**, describing what this program wants. The two are paired before decoding starts, and the values are read in the writer's layout and delivered in the reader's shape. A reader here cannot decode from its own definition alone, because the byte layout is the writer's, not its own. | | tag-numbered identity | writer-schema resolution | |---|---|---| | what identifies a field | a number allocated in the IDL | the field's name in the writer's schema | | schemas needed to decode | the reader's | the writer's *and* the reader's | | what must accompany the bytes | nothing per message | the schema itself, or an identifier for it | | what is load-bearing in the contract | the number | the name | | natural carrier | request and response payloads on a link | records in a bulk file, or a stream with a schema identifier | ## What "must reach the decoder" means in practice The phrase is operational, not theoretical. Under tag-numbered identity, the only thing that must reach the decoder is the bytes: the schema travelled at build time, inside the binary. That is why this style suits two services exchanging millions of small messages — there is no per-message schema overhead and no lookup on the read path. Under writer-schema resolution, the writer's schema must be reachable at the moment of decoding. There are three usual arrangements: 1. **Written into the container.** A bulk file writes the writer's schema once in its header and then millions of records after it. The cost is amortised to nothing per record, and the file is interpretable years later by a tool that never saw the producing program. 2. **Identified per message.** Each message carries a short identifier for its schema, and the reader resolves that identifier against a store the first time it sees it, caching the result. Per-message cost is a few bytes; the read path gains a dependency that must be cached and must have a defined behaviour when the store is unreachable. 3. **Agreed out of band.** Both ends are pinned to one agreed schema by configuration. This works, and it quietly gives up the property the style exists for, since a writer that changes its schema now breaks readers that were never told. ## Why the split exists at all The two styles are optimising different things, and each pays for it. - Tag-numbered identity buys **zero per-message schema overhead and a decode path with no external dependency**. It pays by making the number permanent: a retired number must never be reused, and a reader compiled from an older definition simply has no entry for numbers added later. - Writer-schema resolution buys **exactness about the bytes in hand**: whatever the writer did, the reader is told precisely, so data written long ago by a program that no longer exists is still fully interpretable. It pays with a schema that must be stored, distributed and kept as long as the data, plus the per-message identifier or per-file header. ## The consequence people miss Because identity is a number in one style and a name in the other, **the same edit has opposite consequences**. Renaming a field is invisible on the wire under tag-numbered identity and is a change of identity under name-matched resolution. Renumbering a field is a catastrophe under tag-numbered identity and is meaningless under name-matched resolution, where there are no numbers. When an interviewer asks you to place an encoding, this is the axis they are usually after: not which is faster, but **what a decoder in that style needs in its hands before it can turn bytes into fields** — and therefore what your operations team has to keep alive, for as long as the data lives.

  • Which of the two styles makes a long-lived archive easier to read, and why?
    Writer-schema resolution, when the schema is written into the file's header. The file then explains itself: a tool built years later can open it and recover every field's name and type without access to the producing program's source or build. A tag-numbered archive is only as readable as the definition someone kept beside it.
  • What new failure mode appears when a per-message schema identifier is resolved from a store?
    The decode path acquires a runtime dependency. An unreachable store stalls consumers that meet an identifier they have not cached, so readers normally cache resolved schemas aggressively and the team must decide up front whether an unresolvable identifier blocks, skips or dead-letters the message.
  • Can one system use both bindings at once?
    Yes, and large systems commonly do: a tag-numbered encoding for synchronous service-to-service calls, where both ends deploy together and per-message overhead matters most, and a writer-schema container for the bulk files those services emit for later analysis. They are different hops with different readers, not competing choices for the same hop.

saying these in an interview costs you the question

  • Thinks every schema-driven encoding sends field numbers on the wire
  • Believes the reader's own schema alone can decode name-matched bytes
  • Assumes the writer's schema must be embedded in every single message
  • Says a per-message schema identifier is as large as the schema itself
  • Treats the two bindings as fast versus slow rather than differently bound