skip to content

questions

5

A JSON object carries the same member name twice: why might a filtering proxy and the backend behind it read different values?

level: middleimportance: must knowfreq 55%

answer

  1. one name on the wire, two values
  2. the mapping holds only one
  3. the specification leaves it undefined
  4. first-wins against last-wins
  5. checked one value, applied the other

basics

~20 s

The grammar allows a repeated member name but the data model holds one value per name, so each decoder invents its own resolution: last wins, first wins, reject, or keep both. Two decoders can pick differently.

solid answer

~50 s

A JSON document may repeat a member name; the specification only says names `should` be unique and leaves the behaviour undefined when they are not. Every decoder still has to return something, so the resolution falls out of its implementation: a decoder that builds a map and overwrites keeps the **last** occurrence, one that inserts only when the name is absent keeps the **first**, a strict mode rejects the document, and an order-preserving or event-streaming decoder exposes both. None of those is a bug in isolation. The defect is the chain: a filtering proxy decodes the body, sees `role: user`, approves it and forwards the bytes it received; the backend decodes the same bytes, resolves the repeat the other way and acts as `role: admin`. The check bound to one interpretation, but only bytes travelled.

code

json · 6 lines
json
{
  "role": "user",
  "amount": 10,
  "amount": 1000000,
  "role": "admin"
}

go deeper

for a junior

Remember that a document can spell the same member name twice even though the resulting mapping holds one value, and that the specification does not say which value survives.

for a middle

Explain the four resolutions — last wins, first wins, reject, keep both — and say which implementation shape produces each. That is the mechanics tier of this question.

for a senior

Show the chain: name the two components that decode the same bytes, state which value each applies, and propose decoding once at the boundary or rejecting repeats there.

for a principal

Frame it as a class, not an instance: any input the grammar permits and the data model leaves undefined is a divergence point, so the design question is how many decoders a payload meets at all.

## The grammar allows what the data model does not define A self-describing text encoding defines two separate things, and conflating them is what makes this question interesting. The **grammar** says which byte sequences are a document: an object is an opening brace, zero or more `name: value` members separated by commas, a closing brace. The **data model** says what a document *means*: an object denotes a mapping, and a mapping holds one value per name. Nothing in the grammar prevents a document from spelling the same name twice. JSON's own text says member names **should** be unique and warns that behaviour is unpredictable when they are not — that is guidance to whoever writes the document, not an obligation the decoder is required to enforce. So `{"role": "user", "role": "admin"}` is syntactically valid and semantically undefined. Undefined does not mean nothing happens: a decoder still has to hand something back to its caller, so it resolves the repeat in whatever way its implementation makes natural. ## The resolutions you actually meet | Resolution | What the caller receives | How it arises | |---|---|---| | Last occurrence wins | the final value | build a map, assign on each member, overwrite silently | | First occurrence wins | the earliest value | insert only when the name is not already present | | Reject the document | a decode error | a strict mode that tracks the names it has seen | | Keep every occurrence | a list, or two callbacks | an order-preserving structure, or an event-streaming decoder | A fifth shape appears in decoders that bind a document straight into a typed structure: the repeat hits one field twice, and whether the second assignment overwrites the first, is ignored, or raises an error is a property of the binder, not of the format. Configuration matters as much as the implementation — the same decoder often has a strict switch that changes which row of the table applies. Every row is a defensible reading of an under-specified input. The vulnerability is not in any one of them. ## Two decoders in one path The shape that turns this into a finding is a request body parsed twice by different code: 1. A filtering proxy reads the body, decodes it with its own decoder, and evaluates a policy against the result — say, the request may proceed only when `role` is `user`. 2. The proxy approves and forwards **the bytes it received**, unchanged. 3. The service behind it decodes those same bytes with a different decoder and acts on *its* resolution. If the proxy keeps the first occurrence and the service keeps the last, `{"role": "user", "role": "admin"}` is approved as a user request and executed as an administrative one. Swap the two resolutions and the attacker simply swaps the order of the members. Neither component is misbehaving by its own specification; the security property lived in the gap between them. The sentence worth carrying out of this: **a check binds to an interpretation, but only bytes travel.** Anything that re-derives meaning downstream is free to derive a different one. ## Variations of the same shape - **Audit divergence.** A logging or audit component decodes the body one way and records `role: user`, while the actor decodes it the other way. The trail then faithfully records a value that was never applied, which is worse than no trail. - **Two decoders in one process.** A schema validator and an object binder are often different pieces of code over the same bytes. The validator may see one resolution and the binder another, so the validated document is not the constructed one. - **Nesting.** The repeat can sit several levels down, inside a member the outer component never inspects because it only reads the fields its policy mentions. ## What to propose - **Decode once and forward what you decoded.** Re-encode the document the proxy accepted, or pass a structured form onward, so no second interpretation exists. The cost is byte fidelity: anything downstream that authenticates the client's exact bytes has to be reworked. - **Fail closed on the ambiguity itself.** Configure strict decoding that rejects a repeated member name outright, at the boundary, so the undefined case never enters the system. This is cheap and does not break byte-level signatures. - **Do not authorise on a field someone re-parses.** Where a decision has been made, forward the decision — not the bytes it was made from — so the downstream component has nothing to reinterpret. ## The general rule Wherever a grammar permits an input that the data model leaves undefined, two independent implementations may diverge on it, and any pair of them in one request path is a candidate bypass. Repeated member names are the most reliable example because the divergence is total: the two decoders return different values, with no error anywhere to notice.

  • Does configuring one decoder to reject repeated member names fix the chain?
    It fixes the chain only if that decoder sits in front of every other one, so the document never reaches a lenient reader. A strict setting on the *backend* alone still leaves an inspecting proxy approving documents that are then rejected — safe, but noisy. A strict setting at the boundary is the useful placement: the ambiguous document is refused before any component has a chance to disagree about it.
  • How does an event-streaming decoder differ from a map-building one here?
    A streaming decoder emits a callback per member, so the caller genuinely sees both occurrences and decides what to do — it can reject, or it can silently keep whichever it assigns last, which reproduces the same divergence one layer up. A map-building decoder makes the decision itself and hands back a single value, so the caller cannot tell a repeat ever happened.
  • Do schema-driven binary encodings have the same problem with repeated fields?
    They have the same input but a defined answer: several schema-driven binary encodings specify that a repeated occurrence of a scalar field means the last one wins, so every conforming decoder agrees. That is the point — the danger is not repetition, it is repetition with no defined resolution, which leaves each implementation to choose.

A form arrives with two lines both labelled 'amount'. The clerk who approves it stamps the first line; the cashier who pays it reads the last. Both follow their own rule, and the money that leaves is the number nobody approved.

saying these in an interview costs you the question

  • Says repeated member names are a syntax error, so no decoder accepts them
  • Assumes every decoder keeps the last occurrence
  • Assumes a specification conformance claim makes two decoders agree on undefined input
  • Calls it a bug in one decoder rather than a property of the chain
  • Thinks validating the decoded copy protects the bytes that are forwarded
  • Believes logging the anomaly is a remedy rather than detection
open as a page

Why is validating a decoded copy of a request body unsafe when the original bytes are forwarded unchanged?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Validation applies to one decoder's interpretation of the bytes, and that interpretation is thrown away when the raw bytes are forwarded. The next decoder derives its own meaning, so the payload that is used was never the payload that was checked.

open as a page

A proxy compares a numeric field against a limit and forwards the body: how can the backend see a different number?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A JSON number is decimal text of unbounded precision, but each decoder maps it onto a concrete type: a binary floating-point value, a fixed-width integer, or an arbitrary-precision decimal. Different targets round, overflow and coerce differently, so one payload yields two numbers.

open as a page

Several components in one request path decode the same payload: what would you change so a decoder disagreement cannot become a bypass?

level: principalimportance: should knowfreq 34%

basics

~20 s

Reduce the number of decodes the payload meets. Decode once at the boundary and pass the decoded or canonically re-encoded form onward, so no component re-derives meaning from attacker bytes; where that is impossible, make the boundary decoder strict and fail closed.

open as a page

Why can bytes placed after a complete JSON document change what a second decoder in the chain sees?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Decoders disagree about what must follow a complete value: some stop and ignore the remainder, some require end-of-input and reject, some read on and return a sequence. The same bytes therefore yield one document, an error, or several.

open as a page