When is mandating one canonical encoding across a platform the wrong call for artefacts that must be deduplicated and verified?
answer
- canonicalisation is code you must own
- the cheapest canonical form is none
- address the received bytes as the artefact
- cross-producer dedup forces the profile
- keep verification off the canonicaliser
basics
~20 sCanonicalisation is a second encoder every implementation must reproduce exactly, forever. Where the received bytes can simply be kept and addressed as they are, keeping them is cheaper and fails safer; mandate a profile only when producers must agree on an id without exchanging bytes.
solid answer
~50 sA canonical profile is not a configuration flag — it is **specified behaviour every writer in every language must reproduce bit for bit**, and a change to it moves every recomputed id. It has to be written down, versioned, tested across implementations and treated as correctness-critical, because a canonicaliser bug is silent: it produces plausible bytes that simply are not the agreed ones. So the first question is whether you need one at all. If each artefact arrives as bytes and you can **keep those bytes** and id them as received, verification needs no profile and the whole class of bugs disappears; the price is that one value uploaded in two spellings becomes two artefacts. You pay for a profile when independent producers must agree on an id without exchanging bytes, when records are rebuilt from storage rather than retained, or when cross-language dedup is a product requirement — and then you canonicalise once, at the edge, and record the profile version with the digest.
go deeper
Understand that a canonical encoding is a set of rules somebody has to implement and maintain, not a switch. Knowing that it has a cost at all is the useful takeaway at this stage.
Contrast the two options concretely: keeping the received bytes and addressing them as they are, against specifying rules every writer must reproduce. Say what each one buys and what it gives up.
Show the operational consequence of each. Describe how a canonicaliser defect presents — plausible bytes, drifting ids, failures over a subset of records nobody can characterise — and why retained bytes remove that failure mode entirely.
Make and defend the call. Name the conditions that genuinely force a profile, propose the hybrid that keeps verification independent of it, and state how the profile is versioned and how cross-implementation agreement is proven.
## What canonicalisation actually costs A canonical profile looks like a paragraph of rules and behaves like a distributed specification. Its real costs are: - **A second encoder to own.** Every implementation in every language on the platform must produce byte-identical output, including for the awkward inputs: extreme numeric values, characters with more than one code-point sequence, deeply nested structures, and the field-presence rule. - **Silent failure.** A canonicaliser that is subtly wrong does not throw. It emits well-formed bytes that are simply not the agreed ones, so ids drift, dedup quietly degrades and checks fail for a subset of records nobody can characterise. - **A versioned, immovable contract.** Changing any rule changes every id recomputed afterwards. That makes the profile harder to evolve than most schemas, and it means every stored digest needs to carry which profile produced it. - **Correctness weight.** Where canonical bytes underpin verification, "two different values canonicalise to the same bytes" and "one value has two canonical forms" are both serious defects, so the code needs the review and fuzzing budget of security-relevant code rather than that of a utility. ## The cheaper discipline: keep the bytes The cheapest canonical form is none. If an artefact arrives as a byte string, treat **those octets as the artefact**: store them unmodified, name them by their digest, verify against them, and never regenerate them. The check then depends on nothing but the storage layer, and no encoder anywhere can break it. | Strategy | What it needs | What it gives up | |---|---|---| | Address the bytes as received | Every hop forwards the payload unchanged | Cross-producer dedup: two spellings of one value are two artefacts | | Canonicalise at the edge, keep both | One profile, applied once on ingest, plus the original bytes | Storage for both, and a rule about which one is authoritative | | Canonicalise everywhere | A versioned profile reproduced bit for bit in every implementation | Evolvability, and a large silent-failure surface | ## When you must pay for a profile Four conditions genuinely force it, and they are worth naming out loud in a design discussion: 1. **Independent producers must agree on an id** for a value without exchanging the bytes — two services computing the same dedup or idempotency key from their own copies of a record. 2. **The bytes cannot be retained.** A record assembled from several stores, or republished long after the original writer is gone, has no received octets to fall back on. 3. **Dedup across producers is a product requirement**, not a nice-to-have. If duplicated artefacts merely cost storage, that is often cheaper than owning the profile. 4. **A verifier must re-derive the covered bytes** rather than being handed them — for example, checking a value it reconstructed rather than one it received. If none of these holds, the profile is a cost with no buyer. ## A hybrid that fails safe Where dedup matters but verification matters more, split the two jobs so a canonicaliser bug cannot take down the important one: - **Verification runs over the bytes as received.** It never touches the canonicaliser. - **Dedup runs over a canonical projection**, stored as a secondary index key next to the artefact. Then a defect in the profile degrades dedup quality — duplicates survive, or two records collide on an index key you then confirm by comparing bytes — while nothing that says "this artefact is intact" ever depends on re-encoding. That asymmetry is the design judgement this question is really probing: **push determinism to the least critical consumer that needs it.** ## How to decide in a review Ask, in order: can we keep the bytes? If yes, can we forbid every hop from transforming them, and enforce that? If we can, canonical form is optional and should be justified by a specific dedup requirement rather than adopted on principle. If we cannot keep the bytes, the profile is not optional — and the plan must then include who owns it, how it is versioned, which digest records its version, and how cross-implementation agreement is tested on a shared corpus rather than asserted.
- You decide to keep bytes as received. What still has to be true of the pipeline?That nothing between producer and store rewrites the payload. Carry it as an opaque octet string, keep routing metadata in an envelope so no hop needs to open it, make transformation a contract violation rather than a convention, and log the payload digest at each hop so a mismatch identifies the hop that changed it.
- A team proposes canonicalising everything now, in case a future product needs cross-producer dedup. How do you respond?Ask what it costs to add later against what it costs to own now. Retaining the received bytes does not block adding a canonical projection afterwards — it is a secondary index you can build over stored artefacts. Adopting the profile early buys nothing yet and commits every implementation to a specification that is harder to change than a schema.
saying these in an interview costs you the question
- Mandates a canonical encoding platform-wide without costing it
- Assumes every encoding publishes a deterministic profile
- Treats the canonicaliser as ordinary non-critical library code
- Changes canonical rules without versioning the profile
- Discards received bytes once the value is stored
- Makes verification depend on re-encoding for convenience