Several components in one request path decode the same payload: what would you change so a decoder disagreement cannot become a bypass?
answer
- count the decodes, not the quirks
- two decodes are the precondition
- what leaves is downstream of the check
- canonical form costs byte fidelity
- matching settings decays silently
basics
~20 sReduce the number of decodes the payload meets. Decode once at the boundary and pass the decoded or canonically re-encoded form onward, so no component re-derives meaning from attacker bytes; where that is impossible, make the boundary decoder strict and fail closed.
solid answer
~50 sTreat the count of decodes per payload as the design variable. The strongest change is **decode once**: the boundary parses the bytes, applies policy, and forwards either a structured internal form or a canonical re-encoding, so downstream components have nothing left to reinterpret. Where a second decode is unavoidable — an inspecting component that must forward the client's exact bytes for signature reasons — make the boundary **strict and fail-closed**: reject repeated member names, reject content after the document, reject lenient extensions, reject numbers outside the consumer's exact range. Then remove the incentive: forward decisions rather than the fields decisions were made from, so no downstream parse is security-relevant. Verify with differential testing across the decoders actually deployed. Each option has a cost — processor time, broken byte signatures, coupling to one boundary — and naming those costs is the answer.
go deeper
Take away the core idea: the fewer times a payload is parsed on its way through a system, the fewer chances there are for two components to disagree about it.
Be able to describe what decoding once at the boundary means in practice, and why forwarding a canonical re-encoding removes the second interpretation.
Choose under constraints and defend it: which path can afford re-encoding, which needs a strict fail-closed boundary, and how you prove either holds over time.
Own the trade-off between one decoder of record and the bottleneck it creates, decide what the platform mandates versus what each team chooses, and make the property verifiable rather than aspirational.
## Frame it as a count, not a list of bugs Every parser differential needs the same two ingredients: an input the format leaves under-specified, and **two decodes of the same bytes** by code that resolves it differently. Teams usually attack the first ingredient, closing one ambiguity at a time — duplicate names this quarter, trailing bytes the next. That list never ends, because it is the intersection of two implementations' liberties rather than a property of the format. The variable a lead can actually move is the second ingredient. Ask of each path: **how many times are these bytes parsed, and by how many distinct pieces of code?** A path with one decode has no differential, whatever the format permits. ## The options, strongest first | Option | What it removes | What it costs | |---|---|---| | Decode once, forward the decoded form | the second interpretation | an internal representation and a trusted channel to carry it | | Decode once, forward a canonical re-encoding | the second interpretation | processor time; byte signatures over the client's body break | | Forward a decision, not the data | the security relevance of any later decode | the boundary must own the policy, and it becomes a bottleneck | | Strict, fail-closed boundary decoder | most of the ambiguous input set | still two decodes; strictness rejects some lenient clients | | Pin both sides to one decoder build | today's divergence | expires on the next upgrade or the next new consumer | | Detect and alert on ambiguous documents | nothing | tells you it happened, after it happened | The first three share one property worth stating plainly: what leaves the boundary is **downstream of the check**. The last three are mitigations that leave the shape intact. ## The trade-offs a lead actually argues about 1. **Byte fidelity against uniformity.** Canonical re-encoding guarantees one interpretation, but the forwarded bytes are the boundary's, not the client's. Any component that authenticates the client's exact body — a signature, a content hash, a receipt — has to move to the canonical form, computed identically on both sides. That is a contract change with real migration cost, and it is the reason teams settle for a strict boundary instead. 2. **Centralisation against blast radius.** One decoder of record is easier to reason about and easier to harden, and it is also a single point of failure and a deployment bottleneck. A regression in it is a regression everywhere, and every team now waits on it to accept a new field. 3. **Strictness against clients you do not control.** Rejecting lenient input is free on a machine-to-machine path where both ends are yours. On a public interface it breaks integrations whose clients emit a byte-order mark or a trailing comma, and the rollout needs a measurement period before enforcement. 4. **Uniformity by policy against uniformity by construction.** "Everyone uses the same library and settings" is the option that looks cheapest and holds worst: settings drift silently, nothing fails when they do, and the property is unverifiable at review time. Prefer a control that holds no matter what the other side is configured to do. ## Making the property verifiable A design that cannot be checked will decay, so pair the decision with evidence: - **Differential testing over the decoders actually deployed.** Feed both a corpus of ambiguous documents — repeated names, appended documents, high-precision numbers, values at the exact-representation boundary, exotic escapes — and fail the build on any difference in the resulting value trees. - **An architectural test on the count.** Assert that a payload crossing the boundary is parsed by exactly one component, in the same way module-boundary rules are asserted elsewhere. - **A review question with a yes-or-no answer**: does this change add a component that parses a client-supplied body? If yes, it needs the boundary form, not the raw bytes. ## What a strong answer sounds like It does not open with a list of format quirks. It opens with the invariant — *the thing that is checked must be the thing that is used* — derives the design rule from it — *minimise the number of decodes, and make what leaves the boundary downstream of the check* — and then prices each option honestly, including the one the candidate would not choose. Mentioning that an inspection proxy is structurally in the two-decoder shape, and that this is an argument for putting policy inside the consumer instead, is the observation that marks a lead rather than a practitioner.
- When is a strict boundary decoder the right choice over canonical re-encoding?When something downstream must see the client's exact bytes — a signature, a content hash, a stored receipt — re-encoding is not available without a contract change. A strict boundary then buys most of the protection at none of that cost: it refuses the ambiguous documents rather than resolving them, so the second decode still exists but has a far smaller input set to disagree about.
- Why is pinning both components to the same decoder build a weak control?It holds only for as long as nobody upgrades, reconfigures or adds a consumer, and nothing fails visibly when it stops holding. The property is also unverifiable at review time, because it is a claim about two deployments rather than about one piece of code. Controls that depend on the other side's configuration decay; controls that reject ambiguous input at the boundary do not.
- What does this argue about inspection proxies as a security control in general?An inspecting component that parses bytes and forwards them is, by construction, the two-decoder shape. It can still be useful for coarse controls, but a decision that depends on the fine structure of a payload belongs in the component that consumes it, where the checked value and the used value are the same object. Treat body-inspecting policy as defence in depth, not as the boundary of record.
saying these in an interview costs you the question
- Enumerates format quirks instead of reducing the number of decodes
- Presents canonical re-encoding without mentioning broken byte signatures
- Relies on both sides sharing a library version as the primary control
- Treats alerting on ambiguous documents as prevention
- Ignores that a single decoder of record is also a bottleneck
- Assumes strict rejection is free on a public interface