skip to content

questions

5

In an IDL-first binary encoding, what rides on the wire in place of field names, and what does dropping them buy?

level: middleimportance: must knowfreq 68%

answer

  1. identity without spelling it out
  2. both sides already hold the definition
  3. one small integer per field
  4. tag plus value, no key strings
  5. renames are free, renumbering is not

basics

~20 s

A small numbered tag plus the value's bytes ride instead of the name. Both sides compiled the same schema, so the name is already known. That buys much smaller messages, cheaper decoding, and a field name that is free to change.

solid answer

~50 s

Both endpoints compile the same interface definition, so the decoder already knows that tag `7` means the customer identifier and that it holds an integer. The wire therefore carries an identifier and the value, with no key strings, quotes or separators. Three consequences follow. Messages shrink: a five-field record whose keys average ten characters spends roughly 65 bytes per message on quoted keys and separators in a self-describing text form, and about five bytes on tags here — at ten million messages a day that is around 650 MB of wire traffic. Decoding gets cheaper, because the reader dispatches on a small integer instead of hashing and comparing key strings. And the *name* leaves the contract while the *number* joins it: renaming a field regenerates code on both sides without changing a byte, whereas renumbering one breaks every reader.

code

pseudocode · 13 lines
pseudocode
# the wire carries tag numbers and values, never field names
schema = compiled_schema_for(RequestMessage)   # tag -> (name, type)
record = empty_map()

while more_bytes():
    tag = read_tag()
    if not schema.has(tag):
        stop("cannot name or type this field: no schema entry for tag")
    kind  = schema.type_of(tag)        # only the schema knows the type
    value = read_value(kind)
    record[schema.name_of(tag)] = value

return record

go deeper

for a junior

Remember the one-line shape: the definition is compiled into both programs, so the message on the wire needs only a number per field and the value itself.

for a middle

Be able to explain both halves — where the name went, and the three things that buys: fewer bytes, an integer dispatch instead of string comparison, and no per-field string allocation while decoding.

for a senior

Show that you know which identifier is now load-bearing. Renames are free; reusing a retired number silently corrupts data for older readers, which is a review discipline on the schema file, not something the bytes can catch.

for a principal

Frame it as a trade you are choosing for the whole fleet: density and decode cost against data that no tool, log or archive can interpret without the definition in hand.

## What the bytes actually contain A **schema-driven binary encoding** starts from an **interface definition** — an IDL file that declares a message, its fields, each field's type, and, in the tag-numbered branch of this family, a **number** for each field. Both the sending and the receiving side compile that definition before anything runs, so both hold the same mapping from number to name and type. At encode time the writer emits, for each field it is sending, that field's number followed by the value laid out the way the schema says. At decode time the reader reads the number, looks it up in the schema it compiled, and thereby learns which field this is and how to interpret the bytes that follow. **The field's name exists only in the two compiled schemas — it is never transmitted.** The wire carries no field names and only the minimum of type information the layout itself needs to find where a value ends. Contrast this with a self-describing text object, where every message repeats every key it contains, spelled out and quoted, for every occurrence of that field. | what a field costs | self-describing text | tag-numbered binary | |---|---|---| | identity | the key spelled out and quoted, on every message | one small integer | | type | recovered by the parser from punctuation | fixed in advance by the schema | | value | decimal digits, escaped text | the binary layout the schema names | ## Why the saving is larger than it looks Take a record with five fields whose names average ten characters. In a text object each field pays roughly thirteen bytes of pure identity — two quote characters, the ten-character key, a colon — plus a separator, so about **65 bytes per message** goes to keys alone before a single value is written. The tag-numbered form pays about **one byte per field** for the low field numbers, so roughly **five bytes**. The difference is about 65 bytes on every message; across **ten million messages a day** that is **650 MB** of traffic that carries no information at all. On top of the identity saving: - **Values shrink too.** A 32-bit quantity written as decimal digits can run to ten characters plus a sign; in a binary layout it is at most four bytes, and variable-length integer layouts make small values smaller still. - **Decoding gets cheaper.** The reader compares small integers instead of hashing key strings and comparing them character by character, and it does not have to scan for quotes, escapes or separators. - **Allocation drops.** A text parser typically materialises a key string per field before it can decide what to do with the value; a tag lookup materialises nothing. ## What leaves the contract, and what joins it This is the part interviewers actually probe, because it inverts the intuition people bring from text formats. 1. **The name stops being load-bearing on the wire.** Rename the field in the IDL and both sides regenerate; the bytes are byte-for-byte identical, because the name was never in them. Only the source code that reads the generated accessor changes. 2. **The number becomes load-bearing.** The tag *is* the field's identity. Reuse a retired number for a different field and old readers will happily decode the new value into the old meaning — a silent, typed-looking corruption rather than a parse error. 3. **Tag allocation becomes an editing discipline on the schema file**, because the bytes themselves cannot detect a collision. Two teams that independently assign the same number to different fields produce messages that decode without complaint and mean different things. ## Where the win shrinks Dropping names is not a universal size argument, and a good answer says where it stops paying: - **Messages dominated by one large value.** If a record is a single multi-kilobyte blob, the keys were never the cost; the saving is proportional to the *number of fields*, not the size of the payload. - **Batches under a general-purpose compressor.** Repeated keys across many records are exactly what a dictionary-based compressor such as DEFLATE removes, so the size gap narrows considerably in a compressed bulk file. What compression does **not** remove is the per-message CPU of parsing text, and small lone messages compress poorly because there is little history to exploit. - **Deeply nested sparse structures**, where the framing overhead of nesting starts to rival what the keys cost. ## The price of the trade Everything above is bought with one standing liability: the bytes are meaningless without the matching schema. Nobody can read a captured payload by eye, a stored message outlives the code that wrote it only if the schema outlives it too, and every tool that touches the data needs the definition. That is the deal this family of encodings offers — maximum density and speed on a link where both ends are known, in exchange for data that is opaque on its own.

  • What stops two teams from independently assigning the same field number to different fields?
    Nothing on the wire — the bytes cannot detect a collision, and a reader will decode the other team's value into its own field's meaning. The only defence is that the IDL file is a single reviewed artifact where numbers are allocated once and retired numbers are marked as never to be reused.
  • Does dropping field names help a message that is one large binary attachment?
    Barely. The identity saving scales with the number of fields, not the size of the payload, so a record dominated by one multi-kilobyte value is essentially the same size either way. The remaining wins there are decode cost and the absence of an escaping or text-transfer step for binary data.
  • If a general-purpose compressor already removes repeated keys, why not compress a text format instead?
    Compression narrows the size gap on large batches, where repeated keys are highly redundant, but it does not help a small standalone message with little history to exploit, and it adds CPU on both ends. It also leaves the text parser's per-message cost — scanning, unescaping and materialising key strings — completely untouched.

A row of numbered lockers: the doors carry numbers, and the sheet mapping number to owner hangs in the office. Reading a locker's contents tells you nothing about whose it is unless you hold that sheet.

saying these in an interview costs you the question

  • Thinks the encoding compresses the field names rather than omitting them
  • Believes a field's name can be recovered from the captured bytes alone
  • Assumes renaming a field in the definition breaks readers on the wire
  • Treats field numbers as cosmetic and renumbers them freely
  • Claims the size win scales with payload size rather than field count
open as a page

In schema-driven binary encodings, what must reach the decoder under tag-numbered identity versus writer-schema resolution?

level: middleimportance: must knowfreq 55%

basics

~20 s

Tag-numbered identity needs only the reader's own compiled schema, because the numbers in the bytes are the field identities. Writer-schema resolution additionally needs the exact schema the writer used, reachable by an identifier that travels with the data.

open as a page

A schema-driven binary encoding carries no field names, so an on-call engineer who captures a failing request's bytes can read nothing from them. What standing costs does that opacity impose?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Every place bytes are read needs the matching schema: incident tooling, logs, dead-letter queues and archives. That means decoded projections in logs, a decoder in the on-call toolkit, and schemas retained at least as long as the data they explain.

open as a page

A lead argues that generating stubs for every language from one interface definition makes the internal contract identical everywhere. What does that actually guarantee, and where must teams still be aligned by hand?

level: principalimportance: should knowfreq 38%

basics

~20 s

It guarantees identical bytes and identical field identity: any generated reader decodes what any generated writer wrote. It does not guarantee identical meaning — units, invariants, required-ness, error behaviour and ownership of the definition are agreements no schema expresses.

open as a page

The same schema-driven binary encoding appears as a remote call's payload and as the records inside a bulk analytics file. How does the schema reach the reader in each setting?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

On a remote call the schema reaches the reader at build time, compiled into both peers from one interface definition, with nothing sent per message. In a bulk file it reaches the reader inside the file, written once in the header and amortised over every record.

open as a page