You own the encoding policy for an event archive whose consumers are unknown five years out: which side of this split do you mandate, and on what grounds?
answer
- interpretable from storage alone
- self-describing, or schema in the container
- pointers for live hops only
- a percentage against a total loss
- drill the cold read yearly
basics
~20 sMandate that anything entering long-term storage is interpretable from storage alone: either self-describing bytes, or schema-driven bytes with the writer schema in the same durable object. Decide per hop, not per estate, and defend it on asymmetry of failure.
solid answer
~50 sThe defensible answer is not a format, it is an **invariant**: no bytes enter long-term storage whose interpretation depends on a service that may be decommissioned before the retention window ends. That admits two compliant shapes — self-describing bytes, or schema-driven bytes with the writer's schema written into the same container — and rules out the tempting third, a bare schema identifier resolved elsewhere. I would decide per hop rather than per estate, because the forces genuinely reverse: a live hop between two services deployed together, at high volume and read within seconds, should be tag-keyed and pointer-resolved, while the archive behind it should be self-contained. The grounds are the asymmetry of being wrong. Over-spending on self-description costs a predictable, budgetable tax on bytes and CPU. Under-spending costs an unreadable archive, discovered years later, with no remediation available at any price.
go deeper
Understand that the choice is about what a future reader must have in hand, and that bytes stored without their explanation may be unreadable later.
Compare the concrete costs on both sides — repeated structure and decode work against a dependency on the schema still being available years afterwards.
Argue the per-hop decision with a specific retention window and consumer set, and insist the schema is retained as long as the data it describes.
Write the invariant, not the format; justify it on the asymmetry between a payable rate and an unrecoverable cliff; name where it flips and how exemptions expire.
## What the decision actually is The question is posed as a format choice and is not one. Formats change; the thing you are actually deciding is **what a future reader is required to have in its hands**, and that is a policy that outlives every format in it. Stated as a policy it has three candidate positions: 1. **Self-describing everywhere.** Every durable byte carries its own structure. Maximum reach, permanent overhead. 2. **Schema-driven, self-contained.** Compact encoding, with the writer's schema written into the header of the same file or object as the data. 3. **Schema-driven, pointer-resolved.** Compact encoding with a small identifier, resolved against whatever holds schemas at read time. Smallest, and dependent on that thing existing in five years. ## The forces, and which way each pushes | Force | Pushes towards self-describing | Pushes towards schema-driven | |---|---|---| | Consumers unknown or external | strongly | — | | Retention measured in years | strongly | only if self-contained | | Very high volume per record | — | strongly | | Both ends deployed together | — | strongly | | Need for an enforcement point | — | moderately | | Tolerance for producer friction is low | moderately | — | | Cost of a wrong reading is high | — | moderately, via the contract | Read the table honestly and the two columns are not symmetric in weight. Volume arguments are about a rate: they scale with traffic and they are payable. Retention arguments are about a cliff. ## The asymmetry of being wrong This is the heart of a principal-level answer, and it is what separates it from "it depends": - **Over-spending on self-description** produces a larger archive and more decode CPU. It is measurable today, it appears in a budget line, and it can be corrected later by re-encoding — the data is readable, so it can be migrated. - **Under-spending** produces an archive that cannot be read. It is not measurable today, it appears in no budget, and it cannot be corrected later, because the correction would require the information that was lost. The failure is also **silent until exercised**: nothing tells you the pointer will dangle until someone follows it. When one branch's worst case is a percentage and the other's is a total loss, the two do not get weighed on the same scale, and a policy that treats them as if they do is the actual defect. ## A policy I would defend 1. **The invariant.** Bytes in long-term storage must be interpretable from the storage system alone. Compliance has exactly two shapes: self-describing, or schema-in-the-container. 2. **Pointers are for live hops only.** A message read within seconds by a service you deploy may reference its schema by identifier; the same message on its way to durable storage is rewritten into a compliant shape. 3. **Schema retention matches data retention, by construction rather than by policy.** Writing the schema into the container makes the requirement structural, so nobody has to remember it. 4. **One cold-read drill per year.** Open an old object with a process that holds nothing but the storage credentials and the bytes. An untested claim of readability is an assumption, and this one has a five-year feedback loop. 5. **Exemptions are explicit and expiring.** A team that must break the invariant for cost reasons records what it costs to be wrong and revisits it on a date. ## What I would not put in the policy - **A named format.** The invariant survives a format migration; a format mandate has to be rewritten every time the estate moves, and it invites arguments about syntax rather than about readability. - **A blanket ban on compact encodings.** Most bytes in a large estate never reach long-term storage, and taxing them to protect the minority that does is the mirror-image mistake. - **"Prefer self-describing" as guidance without a rule.** Preferences lose to a spreadsheet showing the storage saving. The invariant survives that conversation because it is stated as a property of the archive rather than a taste. ## Where the answer flips On a high-volume internal hop with short retention, reachable readers and a deployment you control, the pointer-resolved form is not a compromise but the right answer: the coupling it creates is real, bounded, and already present in every other way those two services depend on each other. A principal answer names that flip explicitly, because a policy that never flips is not calibrated, it is just a preference with authority.
- What single invariant would you write down, in one sentence?No bytes may enter long-term storage whose interpretation depends on a system that could be decommissioned before their retention expires. It is deliberately format-free, so it survives migrations; it admits both compliant shapes, so it does not tax teams needlessly; and it is checkable by inspection of a stored object, which means compliance can be audited rather than asserted.
- How does your answer change for a hop between two services you deploy together?It reverses. Retention is seconds, the readers are known and reachable, the volume is high, and the coupling a schema identifier creates already exists through every other dependency between those services. Compact schema-driven bytes with a resolved identifier is correct there, and the archive written behind that hop is where the self-containment requirement attaches instead.
- How would you know the policy is working before the five years are up?By exercising the property rather than trusting it: an annual cold read of the oldest stored object from a process holding only credentials and bytes, plus an audit that samples stored objects for compliant shape. Both give a feedback loop measured in months against a failure mode whose natural feedback loop is measured in years.
saying these in an interview costs you the question
- Answers with a single format mandate for the whole estate.
- Weighs an unreadable archive on the same scale as a storage cost.
- Treats a schema identifier as durable because it has never dangled yet.
- Assumes readable bytes today will still be readable after a migration.
- Sets one policy for live hops and archives without distinguishing them.
- Claims self-description everywhere removes the need for any contract.