skip to content

Why is validating a decoded copy of a request body unsafe when the original bytes are forwarded unchanged?

level: seniorimportance: must knowfreq 50%

answer

  1. the check holds over a value tree
  2. only bytes cross the boundary
  3. the consumer re-derives its own meaning
  4. check-to-use gap without time
  5. forward what you decoded, not the input

basics

~20 s

Validation applies to one decoder's interpretation of the bytes, and that interpretation is thrown away when the raw bytes are forwarded. The next decoder derives its own meaning, so the payload that is used was never the payload that was checked.

solid answer

~50 s

A policy check is a statement about a **decoded value**, not about bytes. When a proxy decodes a body, approves the result and forwards the bytes it received, nothing carries the approved interpretation forward: the service decodes the same bytes independently and may resolve every ambiguity — a repeated member name, a number that rounds, content after the end of the document — differently. The result is a time-of-check to time-of-use gap expressed in bytes rather than in time. The fix is to make the checked thing and the used thing the same object: decode once at the boundary and forward a canonical re-encoding, or pass a structured internal representation, or a decision rather than the data. The trade-off is byte fidelity — a re-encoding breaks any signature or hash taken over the client's original bytes, so those must move to the canonical form.

code

pseudocode · 11 lines
pseudocode
raw = read_body()

// vulnerable: the approved interpretation is discarded
doc = decode_A(raw)
if policy_rejects(doc) then deny() end
forward(raw)                  // consumer runs decode_B(raw)

// safe: the approved interpretation is what travels
doc = decode_A(raw)
if policy_rejects(doc) then deny() end
forward(encode_canonical(doc))

go deeper

for a junior

Hold on to the shape: a check is about the value someone decoded, and forwarding the raw bytes lets the next component decode them again for itself.

for a middle

Be able to name two concrete ambiguities that make two decoders disagree about one body, and explain why the approved interpretation does not travel with the bytes.

for a senior

Demonstrate the remedy under constraints: decode once and forward a canonical encoding, or fail closed on ambiguous syntax, and say what each costs in signatures, fidelity and processor time.

for a principal

Argue about where meaning is allowed to be derived at all. A path that re-parses attacker bytes at several layers is an architectural choice, and centralising the decode has its own coupling and availability costs.

## What a check actually binds to When a component evaluates a policy over a request body, the object it inspects is not the bytes. It is a value tree that some decoder produced from those bytes: a mapping with a `role` member, a number for `amount`, a list of items. The policy holds over **that tree**. Forwarding the original bytes forwards the input to a computation, not its output — and the computation is re-run downstream by different code. That is the whole bug. It is a time-of-check to time-of-use gap where nothing changed in time: the bytes on the wire are identical at both ends. What differs is the function applied to them. ## The ambiguities that make the two functions disagree A decoder pair only diverges where the format leaves room, and there is more room than people expect: - **Repeated member names**, resolved first-wins, last-wins, by rejection, or by keeping both. - **Numbers**, mapped onto a binary floating-point value, a fixed-width integer, or an arbitrary-precision decimal, with different rounding, ranges and overflow behaviour. - **Content after a complete document** — one decoder stops at the end of the first value, another insists on end-of-input, a third reads a sequence. - **Lenient extensions** accepted by one side and rejected by the other: comments, trailing commas, unquoted names, a leading byte-order mark. - **Text handling** — escape sequences, surrogate handling, and whether unpaired or overlong sequences are replaced or refused. Any one of these lets an attacker write a body whose meaning to the checker differs from its meaning to the consumer. ## Three chain shapes, three amounts of exposure | Chain shape | What travels | Exposure | |---|---|---| | Decode, check, forward original bytes | bytes | full: the consumer re-derives meaning with its own decoder | | Decode, check, forward a canonical re-encoding | the checker's interpretation | none from this class; byte signatures break | | Decode, check, forward the decision only | a verdict or a structured value | none from this class; needs a trusted internal channel | The second and third rows share one property: the artefact that leaves the boundary is downstream of the check, not upstream of it. That is the invariant to aim for, and it is easier to state than any list of ambiguities to close. ## The re-encoding trade-off, stated honestly Re-encoding is not free, and a candidate who presents it as a pure win has not deployed it. 1. **Byte-level authentication breaks.** If the client signs or hashes the exact body, the proxy's re-encoded bytes will not verify — whitespace, member order and the spelling of numbers may all change. The signature has to move to a canonical form that both sides compute identically, which is a contract change, not a configuration change. 2. **Fidelity is lost on purpose.** A canonical encoding normalises number spelling and may not round-trip a value the client cared about, such as a decimal with trailing zeros or an integer beyond the precision of the target type. 3. **It costs processor time and allocation** on every request, at the busiest point in the path. When those costs are unacceptable, the cheaper mitigation is to make the boundary decoder **strict and fail-closed**: reject repeated member names, reject anything after the document, reject lenient extensions, reject numbers outside the range the consumer can represent. This does not remove the second decode, so it is a narrowing rather than an elimination; it works because the attacker's raw material is the set of documents the two decoders read differently, and a strict boundary refuses most of that set outright. ## Testing for it The defect is invisible to ordinary functional tests, because every well-formed request behaves identically on both sides. What finds it is **differential testing**: feed the same corpus to both decoders in the chain, compare the resulting value trees, and treat any difference as a finding. A generator that mutates valid documents into the ambiguous corners above — repeated names, huge and high-precision numbers, trailing bytes, exotic escapes — produces those inputs cheaply. ## What a strong answer sounds like State the invariant first: *the thing you check must be the thing you use.* Then name the mechanism by which it is violated here — only bytes cross the boundary, so the interpretation does not survive — and give one concrete divergence. Then offer both remedies with their costs: re-encode and lose byte fidelity, or decode strictly and accept that a second decode still exists. Naming the trade-off is what separates this from a textbook answer.

  • If the client signs the body, where should the signature be taken so re-encoding is possible?
    Over a canonical form both sides can compute — a defined ordering, defined number spelling, no insignificant whitespace — so the proxy's re-encoding hashes to the same value the client signed. Failing that, verify the signature at the boundary, drop it, and let the internal hop carry its own authentication, so no downstream component needs the client's exact bytes.
  • Is a stricter schema check at the proxy an adequate substitute for re-encoding?
    It narrows the gap without closing it. A schema constrains the value tree the proxy decoded, which is the interpretation that is about to be thrown away; it says nothing about what the second decoder will build from the same bytes. Strictness helps most when it rejects the ambiguous *syntax* — repeats, trailing content, lenient extensions — rather than when it constrains the decoded values.
  • How would you find this class in an existing system rather than reason about it?
    Count the decodes per payload. Any path where the same bytes are parsed by two different pieces of code is a candidate, whether they are separate services or a validator and a binder in one process. Then run differential testing: feed a mutated corpus to both decoders and compare the value trees, treating any difference as a finding rather than a curiosity.

saying these in an interview costs you the question

  • Believes forwarding unmodified bytes is the safest option because nothing is altered
  • Thinks a hash of the body proves both sides interpret it the same way
  • Says a stricter schema on the decoded copy closes the gap
  • Treats it as a transport problem rather than an interpretation problem
  • Presents canonical re-encoding as free, ignoring broken byte signatures
  • Assumes functional tests would have caught a divergence