skip to content

Several components in one request path decode the same payload: what would you change so a decoder disagreement cannot become a bypass?

level: principalimportance: should knowfreq 34%

answer

  1. count the decodes, not the quirks
  2. two decodes are the precondition
  3. what leaves is downstream of the check
  4. canonical form costs byte fidelity
  5. matching settings decays silently

basics

~20 s

Reduce the number of decodes the payload meets. Decode once at the boundary and pass the decoded or canonically re-encoded form onward, so no component re-derives meaning from attacker bytes; where that is impossible, make the boundary decoder strict and fail closed.

solid answer

~50 s

Treat the count of decodes per payload as the design variable. The strongest change is **decode once**: the boundary parses the bytes, applies policy, and forwards either a structured internal form or a canonical re-encoding, so downstream components have nothing left to reinterpret. Where a second decode is unavoidable — an inspecting component that must forward the client's exact bytes for signature reasons — make the boundary **strict and fail-closed**: reject repeated member names, reject content after the document, reject lenient extensions, reject numbers outside the consumer's exact range. Then remove the incentive: forward decisions rather than the fields decisions were made from, so no downstream parse is security-relevant. Verify with differential testing across the decoders actually deployed. Each option has a cost — processor time, broken byte signatures, coupling to one boundary — and naming those costs is the answer.

go deeper

for a junior

Take away the core idea: the fewer times a payload is parsed on its way through a system, the fewer chances there are for two components to disagree about it.

for a middle

Be able to describe what decoding once at the boundary means in practice, and why forwarding a canonical re-encoding removes the second interpretation.

for a senior

Choose under constraints and defend it: which path can afford re-encoding, which needs a strict fail-closed boundary, and how you prove either holds over time.

for a principal

Own the trade-off between one decoder of record and the bottleneck it creates, decide what the platform mandates versus what each team chooses, and make the property verifiable rather than aspirational.

## Frame it as a count, not a list of bugs Every parser differential needs the same two ingredients: an input the format leaves under-specified, and **two decodes of the same bytes** by code that resolves it differently. Teams usually attack the first ingredient, closing one ambiguity at a time — duplicate names this quarter, trailing bytes the next. That list never ends, because it is the intersection of two implementations' liberties rather than a property of the format. The variable a lead can actually move is the second ingredient. Ask of each path: **how many times are these bytes parsed, and by how many distinct pieces of code?** A path with one decode has no differential, whatever the format permits. ## The options, strongest first | Option | What it removes | What it costs | |---|---|---| | Decode once, forward the decoded form | the second interpretation | an internal representation and a trusted channel to carry it | | Decode once, forward a canonical re-encoding | the second interpretation | processor time; byte signatures over the client's body break | | Forward a decision, not the data | the security relevance of any later decode | the boundary must own the policy, and it becomes a bottleneck | | Strict, fail-closed boundary decoder | most of the ambiguous input set | still two decodes; strictness rejects some lenient clients | | Pin both sides to one decoder build | today's divergence | expires on the next upgrade or the next new consumer | | Detect and alert on ambiguous documents | nothing | tells you it happened, after it happened | The first three share one property worth stating plainly: what leaves the boundary is **downstream of the check**. The last three are mitigations that leave the shape intact. ## The trade-offs a lead actually argues about 1. **Byte fidelity against uniformity.** Canonical re-encoding guarantees one interpretation, but the forwarded bytes are the boundary's, not the client's. Any component that authenticates the client's exact body — a signature, a content hash, a receipt — has to move to the canonical form, computed identically on both sides. That is a contract change with real migration cost, and it is the reason teams settle for a strict boundary instead. 2. **Centralisation against blast radius.** One decoder of record is easier to reason about and easier to harden, and it is also a single point of failure and a deployment bottleneck. A regression in it is a regression everywhere, and every team now waits on it to accept a new field. 3. **Strictness against clients you do not control.** Rejecting lenient input is free on a machine-to-machine path where both ends are yours. On a public interface it breaks integrations whose clients emit a byte-order mark or a trailing comma, and the rollout needs a measurement period before enforcement. 4. **Uniformity by policy against uniformity by construction.** "Everyone uses the same library and settings" is the option that looks cheapest and holds worst: settings drift silently, nothing fails when they do, and the property is unverifiable at review time. Prefer a control that holds no matter what the other side is configured to do. ## Making the property verifiable A design that cannot be checked will decay, so pair the decision with evidence: - **Differential testing over the decoders actually deployed.** Feed both a corpus of ambiguous documents — repeated names, appended documents, high-precision numbers, values at the exact-representation boundary, exotic escapes — and fail the build on any difference in the resulting value trees. - **An architectural test on the count.** Assert that a payload crossing the boundary is parsed by exactly one component, in the same way module-boundary rules are asserted elsewhere. - **A review question with a yes-or-no answer**: does this change add a component that parses a client-supplied body? If yes, it needs the boundary form, not the raw bytes. ## What a strong answer sounds like It does not open with a list of format quirks. It opens with the invariant — *the thing that is checked must be the thing that is used* — derives the design rule from it — *minimise the number of decodes, and make what leaves the boundary downstream of the check* — and then prices each option honestly, including the one the candidate would not choose. Mentioning that an inspection proxy is structurally in the two-decoder shape, and that this is an argument for putting policy inside the consumer instead, is the observation that marks a lead rather than a practitioner.

  • When is a strict boundary decoder the right choice over canonical re-encoding?
    When something downstream must see the client's exact bytes — a signature, a content hash, a stored receipt — re-encoding is not available without a contract change. A strict boundary then buys most of the protection at none of that cost: it refuses the ambiguous documents rather than resolving them, so the second decode still exists but has a far smaller input set to disagree about.
  • Why is pinning both components to the same decoder build a weak control?
    It holds only for as long as nobody upgrades, reconfigures or adds a consumer, and nothing fails visibly when it stops holding. The property is also unverifiable at review time, because it is a claim about two deployments rather than about one piece of code. Controls that depend on the other side's configuration decay; controls that reject ambiguous input at the boundary do not.
  • What does this argue about inspection proxies as a security control in general?
    An inspecting component that parses bytes and forwards them is, by construction, the two-decoder shape. It can still be useful for coarse controls, but a decision that depends on the fine structure of a payload belongs in the component that consumes it, where the checked value and the used value are the same object. Treat body-inspecting policy as defence in depth, not as the boundary of record.

saying these in an interview costs you the question

  • Enumerates format quirks instead of reducing the number of decodes
  • Presents canonical re-encoding without mentioning broken byte signatures
  • Relies on both sides sharing a library version as the primary control
  • Treats alerting on ambiguous documents as prevention
  • Ignores that a single decoder of record is also a bottleneck
  • Assumes strict rejection is free on a public interface