skip to content

questions

6

What makes an encoding self-describing, and what must a reader of schema-driven bytes obtain before it can decode them?

level: middleimportance: must knowfreq 70%

answer

  1. can a stranger read these bytes
  2. structure inline or out of band
  3. names and types travel, or do not
  4. values only, identity from a schema
  5. walkable is not the same as nameable

basics

~20 s

Self-describing bytes carry their own structure: field identity and type travel beside the values, so any reader can parse them unaided. Schema-driven bytes carry values only, so the reader must first obtain the schema that says which value is which.

solid answer

~50 s

An encoding is **self-describing** when the bytes carry their own structure — each field's identity plus enough type or length information to find where the value ends — so a reader that has never seen this record type can still parse it into named fields. A **schema-driven** encoding strips that out: the writer emits values keyed by a small numeric tag, or by nothing at all but position, and the reader recovers identity from a schema it obtained separately. The honest test is whether a stranger holding only the bytes can interpret them. It is a spectrum rather than a switch: a tag-keyed stream can still be walked and skipped without a schema, it just cannot be *named*. Self-description costs payload size and parse work in every record; dropping it buys those back and takes on a distribution problem instead — the schema has to reach every reader, including readers that do not exist yet.

go deeper

for a junior

Know that some byte streams explain themselves and some do not, and be able to say which side a document of keys and values sits on and why a stranger can read it.

for a middle

Explain what a self-describing stream spends bytes on — inline identity, type markers, boundaries — and what a tag-keyed or positional writer puts in a schema instead.

for a senior

Show that you treat the schema as an artefact with a distribution and retention problem, and can say which side a given hop belongs on and what that costs its consumers.

for a principal

Own the consequence across an estate: which hops may emit bytes that are unreadable on their own, what must travel with them, and what permanent tax self-description everywhere is worth paying.

## The question the split answers Serialization turns a live value into a flat stream of bytes. The moment those bytes leave the process that produced them, a second question appears: **can whoever picks them up work out what they mean?** An encoding is **self-describing** when the answer is yes from the bytes alone — the stream carries its own structure, so a reader that has never seen this record type can still recover named, typed fields. It is **schema-driven** when the answer is no — the stream carries values, and identity comes from a separate description, the **schema**, that the reader must already hold. The reserved setting for this leaf makes the stakes concrete: an archived event file written today, opened five years from now by a consumer that did not exist when it was written. Everything below is about what that consumer needs in its hands on the day it opens the file. ## What a self-describing stream carries A self-describing writer spends bytes on structure so that the reader needs nothing external: - **Field identity inline.** The key or name of each field sits next to its value, repeated in every record. - **Enough type information to parse.** Markers that separate a string from a number from a nested object, so the decoder knows where each value ends and what kind of thing it just read. - **Explicit boundaries.** Delimiters or length prefixes that let a reader walk the stream without knowing the shape in advance. Text encodings are the obvious case — a `JSON` object repeats every key in every record — but **self-description is not the same as readability**. Binary encodings exist that keep inline keys and type markers and are simply not printable. The property that matters is self-sufficiency of the bytes, not whether a human can read them in a terminal. ## What a schema-driven stream leaves out A schema-driven writer removes exactly what the self-describing writer paid for. Two strategies dominate, and they remove different amounts: 1. **Tag-keyed.** Each field is written as a small numeric **tag** plus its value, usually with a wire-type or length marker. Names never appear; the schema maps tag to name and to the declared type. 2. **Purely positional.** Nothing identifies a field at all. Values are written back to back in the order the writer's schema declares, and identity is recovered from *position*. Without that schema the stream is an undifferentiated run of bytes. ## It is a spectrum, and most systems live in the middle | What the bytes carry | Self-describing text | Self-describing binary | Tag-keyed | Purely positional | |---|---|---|---|---| | Field names inline | yes | yes | no | no | | Type or length markers | yes | yes | yes | no | | Walkable with no schema | yes | yes | yes | no | | Nameable with no schema | yes | yes | no | no | | Bytes spent on structure | highest | high | low | none | The middle column pair is where the interview usually lands. A tag-keyed stream is **structurally self-describing and semantically schema-driven**: any reader can find field boundaries and skip what it does not recognise, but only the schema turns tag `7` into `customerId`. ## What the choice actually decides - **Who can read the bytes.** An external or unknown consumer argues for self-description; a pipeline whose readers you deploy argues the other way. - **What a byte costs.** Repeating names in every record is a permanent tax that scales with volume, not with the number of distinct record types. - **Where the contract lives.** With a schema, the contract is an artefact you can review, version and diff. Without one, the contract is whatever the producer happened to emit last week. - **What happens on rename and reorder.** Names on the wire make a rename a wire event; tags make it a source-level one; position makes ordering itself the contract. ## The failure each side owns Self-description fails by **drift**: because nothing has to agree in advance, producers quietly add, drop and re-shape fields, and consumers accumulate defensive coercion. Nothing ever breaks loudly, so nobody fixes it. Schema-driven encoding fails by **separation**: the bytes outlive the thing that explains them. If the schema for a five-year-old positional archive is gone — deleted, never retained, or held only by a system that was decommissioned — the archive is not hard to read, it is unreadable. That asymmetry, not payload size, is what makes this the first question in almost every format-selection discussion.

  • Can a binary encoding be self-describing, or does going binary always require a schema?
    It can. Self-description is about whether identity and type travel with the values, not about whether the result is printable. Binary encodings exist that keep inline keys and type markers; they are more compact than text and still parse with no external schema. Text versus binary is a separate axis from self-describing versus schema-driven, and the four combinations all exist.
  • Where does a tag-keyed binary encoding sit on this split?
    In the middle, and saying so is the sign of a good answer. The numeric tag and a wire-type or length marker are inline, so any reader can walk the stream and step over fields it does not recognise with no schema at all. What it cannot do without the schema is name those fields or know their declared types. Structurally self-describing, semantically schema-driven.
  • Why is this distinction the first thing raised in a format-selection discussion?
    Because it decides what a consumer must have in hand, which is the constraint the other choices bend around. An unknown or external consumer pushes towards self-describing bytes; a high-volume internal hop with a reliable way to distribute schemas pushes the other way. Payload size, decode cost and how renames behave all follow from that first answer rather than driving it.

Self-describing bytes are a crate where every part is individually labelled; schema-driven bytes are a crate of unlabelled parts with the packing list mailed separately. The second ships lighter and is worthless if the list goes astray.

saying these in an interview costs you the question

  • Treats self-describing as a synonym for human-readable text.
  • Assumes every schema-driven message embeds its schema in each record.
  • Says a reader's own schema is always enough to decode the writer's bytes.
  • Claims self-describing bytes need no shared agreement with the producer.
  • Thinks dropping field names only matters for payload size.
open as a page

What does schema-on-write guarantee about an archived event file that schema-on-read leaves to whoever opens it five years later?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Schema-on-write checks the record against a declared schema before the bytes are stored, so everything in the archive conforms to some known version. Schema-on-read stores whatever arrived and leaves structure and validation to each consumer, years later.

open as a page

How do inline field names, numeric tags, and purely positional fields differ in what the wire carries and what a rename costs?

level: middleimportance: should knowfreq 55%

basics

~20 s

Inline names put identity in every record, so a rename changes the bytes and breaks readers. Numeric tags carry identity in one or two bytes and leave names free to change. Positional fields carry no identity, so order becomes the contract.

open as a page

Why can a reader of purely positional schema-driven bytes not decode them with only its own current schema?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Because the bytes were laid out by the writer's schema, not the reader's. The reader's schema says what it wants; only the writer's says what is actually there, in what order and at what width, so decoding needs both.

open as a page

You own the encoding policy for an event archive whose consumers are unknown five years out: which side of this split do you mandate, and on what grounds?

level: principalimportance: should knowfreq 35%

basics

~20 s

Mandate that anything entering long-term storage is interpretable from storage alone: either self-describing bytes, or schema-driven bytes with the writer schema in the same durable object. Decide per hop, not per estate, and defend it on asymmetry of failure.

open as a page

A team says archived documents need no schema because field names are in the bytes; what does self-description still not supply?

level: middleimportance: nice to knowfreq 27%

basics

~20 s

Names give a reader structure, not meaning. Self-describing bytes never say which fields are required, what an absent one means, what units or which closed set of values apply, or which shapes are legal but simply rare.

open as a page