An intermediary decodes each message into objects and re-encodes it before forwarding; why does that break a downstream integrity check over the bytes?
answer
- the check covers bytes, not meaning
- decoding discards the spelling
- re-encoding re-chooses every free option
- keep the octets as received
- carry the payload as an opaque blob
basics
~20 sRe-encoding rebuilds bytes from a decoded model, which has already dropped what the check covered: member order, spacing, number spelling, presence. The check is over bytes, so identical meaning still yields different bytes and fails.
solid answer
~50 sThe check covers the **octets that were on the wire**, and decoding throws most of that away. A decoded model remembers the values; it does not remember which of several valid spellings carried them — the member order, the whitespace, whether `1` was written as `1.0`, which characters were escaped, whether a default-valued field was present. When the intermediary re-encodes, its writer makes those choices again, from its own defaults, so the forwarded bytes differ from the originals even though nothing was tampered with. The rule is: **verify against the bytes as received**, and decode only afterwards. Anything that must survive a hop is carried as an opaque payload — a length-delimited byte field, or a text-safe blob inside an envelope — so no intermediary can helpfully normalise it. Re-encoding is only safe when both ends run one agreed canonical profile.
code
pseudocode · 10 linesprocedure verify_wrong(raw_bytes, expected_digest):
value = decode(raw_bytes) # spelling is gone from here on
rebuilt = encode(value) # writer re-chooses order, spacing, number form
return digest(rebuilt) == expected_digest
# holds only while this writer happens to reproduce the original spelling
procedure verify_right(raw_bytes, expected_digest):
if digest(raw_bytes) != expected_digest:
return REJECT
return decode(raw_bytes) # decode only what already checked outgo deeper
Remember that an integrity check is computed over a byte sequence. If the bytes change for any reason, the check fails — even when the information they carry is unchanged.
Explain what decoding discards: the member order, the whitespace, the number spelling, the escapes and whether a default-valued field was present. Then explain why a re-encode has to invent all of it again.
Diagnose it. Recognise the byte-different-but-semantically-identical signature, explain why the failures are intermittent, and name the discipline — check the octets as received, decode second, keep the payload opaque across hops.
Set the contract for the whole pipeline: which octets a check covers, where bytes are retained, that transformation in transit is forbidden and how a violating hop is detected, and which producers are allowed to rely on a canonical profile instead.
## What the check is actually over When a detached digest or an integrity tag accompanies an artefact, it covers **a byte sequence**, not a value. Whoever produced it ran the computation over the exact octets it emitted; whoever checks it must run the same computation over the same octets. Any step that changes the octets — even in a way no reader would notice — changes the outcome. This is the whole reason canonical form exists, and the reason a pipeline's shape matters as much as its cryptography. ## What a decode-then-re-encode hop destroys Decoding maps bytes to a model. The model is deliberately lossy about presentation, because presentation is what the encoding was hiding. After the hop, the writer must reinvent every choice the original writer made: - **Member order** — reconstructed from whatever order the model iterates in, not from the wire. - **Whitespace and layout** — the new writer's defaults, not the old one's. - **Number spelling** — a value decoded as a number is re-emitted in the writer's preferred form; an integer that arrived in a wide binary form may go out in the shortest one, or the reverse. - **String escaping and normalization** — the decoder resolved escapes into characters, so the re-encoder picks fresh ones, and any normalization applied on the way in is baked in. - **Presence** — a decoded model that materialises defaults cannot tell you whether a field was on the wire, so the re-encoder guesses. - **Unknown fields** — a model that dropped what it did not recognise re-emits a strictly smaller document. The result is the diagnostic signature people remember: **the payload is byte-different and semantically identical**, so every debugging instinct that compares decoded values reports "no difference" while the check keeps failing. Worse, it is often *intermittent* — if the intermediary happens to re-emit the same spelling, the check passes, and the failures cluster around whichever records exercise a freedom the two writers treat differently. ## The rule: check the octets you received Two disciplines, in order of preference: 1. **Retain and verify the received bytes.** Treat the payload as an opaque octet string from the moment it arrives. Check it *before* decoding, and keep the original bytes for as long as anyone downstream may need to re-check them. This makes the check independent of every encoder in the system. 2. **Canonicalise once, at the producer, and agree the profile.** Where the bytes cannot be retained — the record is rebuilt from storage, or assembled by a service that never saw the original — both ends must encode under one written-down canonical profile so that re-encoding reproduces the bytes exactly. | Approach | What must be true | Main failure mode | |---|---|---| | Verify bytes as received | Every hop forwards the payload unchanged | A well-meaning proxy normalises or pretty-prints it | | Canonicalise on both sides | One agreed, versioned profile in every implementation | A writer drifts from the profile and ids move silently | ## Designing the pipeline so nobody can re-encode 1. **Carry the payload as an opaque blob**, inside an envelope that holds the metadata a router actually needs — so routing never requires opening the payload. 2. **Give the check a defined scope.** Say precisely which octets are covered: the payload alone, or the payload plus named envelope fields in a stated order. 3. **Forbid transformation in transit** as an explicit contract term, and make it detectable: log the digest of the payload as received at each hop, so a mismatch names the hop that changed it rather than the endpoint that noticed. 4. **Check before you decode.** Reversing that order not only breaks the check, it also hands untrusted bytes to a decoder before anything has vouched for them. ## When you genuinely have no original bytes Some producers legitimately cannot keep them: a record is assembled from several tables, or a value is republished years later. That is exactly the case canonical form was invented for, and it comes with an obligation — the profile must be specified precisely enough that an independent implementation reproduces it, versioned so it can change, and recorded alongside the digest so a verifier knows which rules to apply. A digest whose profile is implicit is a digest nobody can re-derive once the original writer is gone.
- The failures are intermittent — most messages pass and a few fail. What does that pattern tell you?That the two writers agree on most spellings and diverge on a specific freedom. Group the failures and look for what those records have in common: a member name that sorts differently, a number that one writer emits with a fractional part, a character one writer escapes, a field left at its default. Intermittency points at a freedom exercised by only some values, not at a corrupt channel.
- A hop must read one routing field out of the payload. How do you let it, without letting it re-encode?Lift that field into the envelope so the hop never opens the payload, and keep the payload an opaque octet string. If it truly must peek inside, it may decode a copy for its own decision, but it forwards the original octets untouched — reading is harmless, re-emitting is not.
- Should the digest cover the envelope as well as the payload?Decide it explicitly and write it down. Covering the payload alone lets routing metadata change in transit, which is usually what you want, but then nothing binds the payload to its envelope. Covering named envelope fields too binds them, at the cost that any hop legitimately rewriting one breaks the check. What is fatal is leaving the scope unstated, because the two ends will then disagree about which octets were covered.
saying these in an interview costs you the question
- Says the payload must have been tampered with
- Believes decode then encode reproduces the exact bytes
- Verifies against a freshly encoded copy of the model
- Lets a proxy normalise the payload it forwards
- Decodes untrusted bytes before checking them
- Assumes any two writers emit identical bytes for one value