skip to content

A team says archived documents need no schema because field names are in the bytes; what does self-description still not supply?

level: middleimportance: nice to knowfreq 27%

answer

  1. shape is visible, contract is not
  2. names without units or obligation
  3. what does a missing field mean
  4. inferred from a sample, not declared
  5. tells you what is, not what should be

basics

~20 s

Names give a reader structure, not meaning. Self-describing bytes never say which fields are required, what an absent one means, what units or which closed set of values apply, or which shapes are legal but simply rare.

solid answer

~50 s

Self-describing is not self-explaining. Inline keys let a reader recover the shape of the records it is holding, which is genuinely useful and is the whole reason the encoding costs what it does. What they cannot carry is the **contract**: which fields are mandatory, what the absence of one signifies, the closed set a categorical field may draw from, the units and precision behind a bare number, and how any of that has changed over the archive's life. A late consumer that has only the bytes ends up inferring all of it from the records it happens to sample — and inference from a sample is not a schema. A field optional and rare looks nonexistent; a field whose meaning shifted three years in looks uniform. That gap is why teams publish a schema for self-describing data even though nothing forces them to.

go deeper

for a junior

Know that seeing a field's name tells you it exists, not what it means, whether it was required, or what its number is measured in.

for a middle

List what a schema pins that the bytes cannot — obligation, closed sets, units, the meaning of absence — and explain why a shape inferred from a sample misses rare optionals.

for a senior

Argue for publishing and enforcing a contract over readable bytes, and recognise the drift signature of an archive where every consumer inferred its own reading.

for a principal

Decide what must accompany data into long-term storage beyond the bytes, and who owns the semantics that no schema language can express.

## Structure is not meaning A self-describing document hands a reader real information: the field names present, the nesting, and enough marking to know a string from a number. That is why the encoding is chosen, and it is not nothing. But there is a widespread slide from *the reader can see the shape* to *the reader knows the contract*, and the second does not follow from the first. Everything below is something a schema states and the bytes cannot. ## What a schema pins that inline names do not - **Obligation.** Which fields must be present. A document with a field missing is indistinguishable from a document where the field is legitimately optional. - **The meaning of absence.** Whether a missing field means unknown, not applicable, or unchanged — a distinction another leaf owns in depth, and one the bytes never carry. - **Closed sets.** That a categorical field draws from exactly five permitted values, so a sixth is a defect rather than new data. - **Units and precision.** That a bare number is minutes and not seconds, or that a timestamp's precision and reference point are what the reader assumes. - **Cardinality and nesting rules.** That a field may repeat, or must not, or is meaningful only alongside another. - **History.** Which of these statements were true in which period of the archive. ## Inference from a sample is not a contract When the schema does not exist, every late consumer derives one, and derivation has failure modes that are invisible from inside it: 1. **Rare optionals disappear.** A field present in one record in ten thousand is absent from any sample the consumer draws, so the inferred shape simply lacks it — until production hands it one. 2. **A widened type reads as consistent.** If a field was numeric for three years and became a string afterwards, a sample from either side looks perfectly uniform, and a sample spanning both looks like corruption. 3. **Coercion becomes private knowledge.** Each consumer writes its own handling for the shapes it has met, none of that reaches any other consumer, and no artefact records which reading was correct. 4. **Nobody can be wrong.** Without a declared contract, a producer that changes a shape has not violated anything, which means there is no conversation to have and no review to fail. ## What it costs the five-year-later consumer The reserved setting makes this sharp. A consumer opening an archived event file today can parse every record — that is exactly what self-description bought — and still not know whether the `amount` field it is summing was ever in a different unit, whether the records missing `channel` are a defect or a legitimate older shape, or whether the four values it sees in `status` are the whole set. It will produce a number. The number will look fine. ## How teams close the gap The fix is not to abandon the encoding; it is to stop treating the encoding as a substitute for the contract: - **Publish a schema for self-describing data anyway,** and validate against it at write time. The bytes stay readable by anything; the shape stays enforced by something. - **Store the contract with the data, not only in a code repository,** so the archive and its explanation share a retention policy and a blast radius. - **Write down the semantics the schema cannot express** — units, the meaning of absence, the reason a field exists — where the consumer will look, because a field named `value` with no note is a five-year mystery. - **Treat an inferred schema as a hypothesis,** and promote it to a contract by agreement with the producer rather than by having shipped a job that works. The short form, and the answer an interviewer is listening for: **self-describing bytes tell you what is there; they never tell you what was supposed to be there.**

  • Why is a schema inferred from an archive treated as a hypothesis rather than a contract?
    Because it describes the records that happened to be sampled, not the records that are permitted. Rare optional fields are missing from it, a field that changed type mid-history is flattened or looks corrupt, and nothing in it is binding on the producer. It documents the past imprecisely and constrains the future not at all, so it can be a starting point for an agreement but never the agreement itself.
  • If nothing enforces it, what is the point of publishing a schema for self-describing data?
    It creates something that can be reviewed, versioned and pointed at when a change would break a consumer, and it gives every consumer one shared reading instead of several private ones. Pair it with validation at write time and it becomes enforced as well — self-describing bytes and an enforced contract are not in tension, they are the combination most mature pipelines settle on.

saying these in an interview costs you the question

  • Says a document containing field names is effectively its own schema.
  • Treats a shape inferred from sampled records as the producer's contract.
  • Assumes an absent field always means the same thing across an archive.
  • Believes readable bytes rule out silent changes to a field's meaning.
  • Thinks publishing a schema is pointless when the encoding is self-describing.