Why can a reader of purely positional schema-driven bytes not decode them with only its own current schema?
answer
- whose schema laid out these bytes
- the reader wants, the writer wrote
- two schemas participate, not one
- ship it, header it, or point at it
- a dangling pointer loses the archive
basics
~20 sBecause the bytes were laid out by the writer's schema, not the reader's. The reader's schema says what it wants; only the writer's says what is actually there, in what order and at what width, so decoding needs both.
solid answer
~50 sA positional stream has no identity in it, so the layout comes entirely from the schema the **writer** held at the moment it encoded. The reader's own schema is a statement of what it wants today, which may have gained, lost or reordered fields since. Decoding therefore is not parsing one schema but **resolving two**: walk the stream in the writer's declared order, keep the fields the reader wants, step past the ones it does not, and fill in the reader's fields the writer never wrote. That makes getting the writer's schema to the reader a first-class problem, solved either by shipping it with the data — typically once in a container header rather than per record — or by referencing it with an identifier the reader resolves. If it is lost, positional bytes are not merely hard to read, they are unrecoverable.
code
pseudocode · 13 lines// positional bytes: only the WRITER's schema says what is where
cursor = start_of_record
record = {}
for each wfield in writerSchema.fields: // must follow writer order
value = decode(bytes, cursor, wfield.type) // advances cursor either way
if readerSchema.has(wfield.name):
record[wfield.name] = value // wanted: keep it
// else: decoded only to move past it, then dropped
for each rfield in readerSchema.fields:
if not writerSchema.has(rfield.name):
record[rfield.name] = rfield.declaredFallback // writer wrote no bytesgo deeper
Remember that positional bytes are meaningless on their own, and that the description needed to read them belongs to whoever wrote them, not whoever reads them.
Explain the resolution walk: traverse in writer order, keep what the reader wants, step past the rest, and supply the reader's own fallback where no bytes exist.
Show you manage schema availability as an operational concern — retention matched to the data, self-contained headers for durable files, and a cold-read drill that proves it.
Own the invariant that nothing durable may depend on a live service to be interpretable, and decide where pointers are acceptable and where self-containment is mandatory.
## The bytes were laid out by someone else's schema A purely positional stream is a run of values with nothing marking where one ends and the next begins. What cuts it into fields is the **writer's schema**: the declared order of fields and the width or encoding of each. That schema was fixed at the moment the bytes were produced, possibly years ago. The **reader's schema** is a different artefact answering a different question. It describes the record the reader wants *now*: the fields it will use, under the names its code expects. Between the two writes and reads, fields may have been added, retired or reordered. Using the reader's schema to walk the writer's bytes means cutting the stream at the wrong offsets and returning confidently wrong values. ## Resolution is matching two schemas, not parsing one Decoding therefore has a shape that surprises people the first time they see it: 1. **Walk in writer order.** The stream can only be traversed in the order the writer declared, because that order is the only thing that says where each value ends. 2. **Keep what the reader asked for.** A field present in both is decoded and placed under the reader's name for it. 3. **Advance past what it did not.** A field the writer wrote and the reader has no use for still has to be decoded far enough to move the cursor. 4. **Fill what the writer never wrote.** A field the reader declares and the writer's schema does not contain has no bytes at all, so its value comes from the reader's declaration. The important consequence is that **the writer's schema is not optional and not replaceable**. How a specific format resolves mismatched types between the two is that format's own business; the structural point here is that two schemas participate, always. ## Getting the writer's schema to the reader There are three ways, and they trade size against coupling: 1. **Per record.** Prepend the schema to each message. Self-sufficient and absurdly wasteful, since it re-pays the cost that dropping names was meant to save. Essentially nobody does this for volume data. 2. **Per container.** Write the schema once in the header of a file or block holding many thousands of records. The amortised cost is negligible, and the file becomes readable with nothing else in hand — which is exactly the property an archive wants. 3. **By identifier.** Put a small id in the message and resolve it to a schema held elsewhere. Smallest on the wire, and it creates a hard dependency: the bytes are interpretable only as long as whatever holds that mapping is reachable and retains the entry. The design of that authority is a separate subject; what matters here is that the id is a pointer, and a pointer can dangle. ## What survives if the schema is lost | Encoding | Recoverable without the writer's schema | |---|---| | Self-describing | Everything structural: names, nesting, boundaries, coarse types | | Tag-keyed | Field boundaries, coarse wire types, tag numbers — no names, no meaning | | Purely positional | Nothing dependable: a run of bytes with no boundaries to find | The last row is the one to say out loud in an interview. Losing the schema for a tag-keyed archive is a painful reverse-engineering exercise with partial success; losing it for a positional archive means the data is gone while the storage bill continues. ## What this implies operationally - **Retention of the schema must match retention of the data.** A five-year archive referencing schemas that are pruned after one year is a five-year archive with a one-year lifespan. - **Prefer the container header for anything durable.** Self-containment costs one copy per file and removes the dangling-pointer failure entirely. - **Identifiers are fine for live hops.** A message on a queue read within seconds by a service you deploy can afford a pointer; a file nobody will open for five years cannot. - **Test the cold path.** The only way to know an archive is readable is to read one, from a process holding nothing but the file and the resolution mechanism — a drill teams almost never run until the day it matters.
- What is actually recoverable if the writer's schema is lost?It depends entirely on how much identity was inline. For a tag-keyed stream the markers still give field boundaries, tag numbers and coarse wire types, so a determined engineer can reverse-engineer a partial picture and never the names or meaning. For a purely positional stream there is nothing: no boundaries, no types, no way to tell where one value stopped. The bytes remain, the data does not.
- Why put the writer's schema in a container header rather than in every record?Amortisation. One copy in front of hundreds of thousands of records costs effectively nothing per record, and it makes the file interpretable with nothing else in hand — the property that matters for an archive. Repeating it per record would re-pay exactly the overhead that dropping inline names was meant to avoid, which defeats the reason for choosing the encoding.
saying these in an interview costs you the question
- Says the reader's current schema is sufficient to decode any old record.
- Assumes skipping an unwanted field costs nothing in a positional layout.
- Treats a schema identifier as if it were the schema itself.
- Keeps schemas on a shorter retention than the data they describe.
- Claims positional bytes can be reverse-engineered as easily as tagged ones.