skip to content

Two producers write the same logical record, one omitting a field and one sending its default; why do their content ids differ?

level: seniorimportance: should knowfreq 38%

answer

  1. two legal spellings, one logical record
  2. omitted field versus default-valued field
  3. the profile must decree a presence rule
  4. omit-at-default erases absence
  5. always-emit ties ids to schema version

basics

~20 s

Both spellings are valid — one omits the field, the other writes it at its default — so the bytes differ and so do the ids. A canonical profile must pick one presence rule and apply it before hashing.

solid answer

~50 s

A reader that materialises defaults resolves both documents to the same value, but the id is a digest of **bytes**, and the two byte strings are not the same. Nothing is misbehaving: presence at default is one of the freedoms a canonical profile has to close. The two options are **omit every field sitting at its default** or **always emit every field**, and they are not symmetric. Omit-at-default keeps ids stable as defaulted fields are added to the schema, but it collapses "absent" and "present at the default" onto one id — fatal if your contract gives absence its own meaning. Always-emit preserves nothing about the original presence either, and it makes the id a function of the schema version: add a defaulted field and every recomputed id moves. Pick deliberately, write it in the profile, and apply it before hashing rather than per writer.

go deeper

for a junior

Know that an omitted field and a field written at its default are different bytes even when a reader resolves both to the same value, and that anything hashing the bytes will therefore see two things.

for a middle

Explain presence as one of the freedoms a canonical profile closes, state the two available rules, and describe what each does to a record whose field is sitting at the default value.

for a senior

Show the consequence in an operating system: duplicated artefacts, cache misses across producers, retries that fail to dedup. Walk the always-emit case when a defaulted field is added and say exactly which ids move and when.

for a principal

Own the coupling between the content-id space and schema evolution. Decide whether presence may carry meaning at all, version the profile, and plan the migration for the day the rule has to change.

## Why both writers are legal An encoding that supports optional fields generally accepts both documents: the field omitted, and the field present carrying the value a reader would have supplied anyway. A reader that fills defaults treats them as the same value — and that is exactly why the divergence is so confusing when it reaches a content-addressed store. The id is not computed from what the reader concluded; it is computed from the octets the writer produced. Two legal spellings, two byte strings, two ids, one logical record. The practical symptoms are unglamorous and expensive: a dedup table that stores the same artefact twice, a cache that misses for one producer and hits for another, an idempotency key that fails to suppress a retry issued by a different service, and a diff tool that reports two artefacts as distinct while showing no differing field. ## Picking a presence rule A canonical profile must decree one rule. There are two, and the choice has consequences well beyond bytes: | Rule | What the canonicaliser does | What you gain | What you pay | |---|---|---|---| | Omit at default | Drops every field whose value equals its declared default | Ids do not move when a defaulted field is added, as long as writers leave it unset | "Absent" and "present at the default" collapse onto one id | | Always emit | Writes every field in the schema, filling unset ones from their defaults | The encoded form is self-contained and uniform | The id is a function of the schema version; recomputed ids move when the schema grows | Walk the always-emit case explicitly, because candidates get its direction backwards. A new optional field with a default is added. Nothing about the stored octets of an old record changes. But the moment anyone **recomputes** that record's id — a re-upload, a migration, a rebuild from storage — the canonicaliser now emits the new field too, the bytes are longer, and the id is different. Under omit-at-default, the same recomputation leaves the bytes untouched while the field is unset, so the id survives. That asymmetry is the whole reason to choose deliberately rather than by accident. ## The distinction you may be erasing Omit-at-default is the safer default *only* when your contract does not distinguish absence from the default value. Where it does — a partial update whose omitted fields mean "leave alone", a measurement whose absence means "not taken" rather than zero — an omit-at-default canonicaliser destroys information before the digest is computed, and two records with genuinely different meanings receive one id. The fix is not to change the presence rule; it is to stop leaning on defaults for a distinction that matters. Model the difference **in the value**: a field whose type can hold "not supplied" as a value of its own, so that presence stops being a spelling decision and becomes part of what is being encoded. ## Designing it out 1. **Normalise the value before the canonical encoder sees it.** Resolve which fields are semantically present, then hand the encoder a value with no remaining ambiguity. 2. **Prefer shapes with no defaults** for anything that is content-addressed. A schema whose fields are all explicitly carried has no presence freedom left to pin. 3. **Record the profile version next to the digest.** Any change to the presence rule changes every recomputed id, so a verifier needs to know which rules produced the id it is holding. 4. **Test the pair explicitly.** Keep a fixture of the same record in both spellings, and assert that canonicalising both yields one byte sequence. ## What to check in review - Does any producer compute the id from its own writer's output rather than from the canonical encoder? That is the usual root cause. - Does a hop fill defaults on the way through, so a record that arrived sparse leaves dense? - Does the contract anywhere give absence a meaning of its own, while the profile omits defaults? - If the schema gained a defaulted field last quarter, did any ids move — and did anything depend on them not moving?

  • Your contract gives an absent field a meaning distinct from its default. Which presence rule do you choose?
    Neither is adequate on its own — the right move is to stop encoding that distinction as presence. Give the field a type that can carry "not supplied" as a value, so the difference lives in the value and survives any presence rule. Then omit-at-default is safe, because the canonicaliser is no longer deciding anything meaningful.
  • You must change the presence rule on a live content-addressed store. How do you do it?
    Version the profile and treat the change as a new id space. Record the profile version alongside every digest, keep verifying old artefacts under the old rules, and compute new ids under the new ones. If old and new ids must coexist for the same artefact, store the mapping explicitly; silently recomputing ids under new rules orphans every reference that already exists.

saying these in an interview costs you the question

  • Says an omitted field and its default are interchangeable
  • Assumes the id depends only on the business record
  • Fills in defaults before computing the content id
  • Treats presence as each writer's private choice
  • Assumes a new defaulted field cannot move any existing id
  • Leaves the presence rule unwritten and per service