skip to content

questions

10

A JSON object carries the same member name twice: why might a filtering proxy and the backend behind it read different values?

level: middleimportance: must knowfreq 55%

answer

  1. one name on the wire, two values
  2. the mapping holds only one
  3. the specification leaves it undefined
  4. first-wins against last-wins
  5. checked one value, applied the other

basics

~20 s

The grammar allows a repeated member name but the data model holds one value per name, so each decoder invents its own resolution: last wins, first wins, reject, or keep both. Two decoders can pick differently.

solid answer

~50 s

A JSON document may repeat a member name; the specification only says names `should` be unique and leaves the behaviour undefined when they are not. Every decoder still has to return something, so the resolution falls out of its implementation: a decoder that builds a map and overwrites keeps the **last** occurrence, one that inserts only when the name is absent keeps the **first**, a strict mode rejects the document, and an order-preserving or event-streaming decoder exposes both. None of those is a bug in isolation. The defect is the chain: a filtering proxy decodes the body, sees `role: user`, approves it and forwards the bytes it received; the backend decodes the same bytes, resolves the repeat the other way and acts as `role: admin`. The check bound to one interpretation, but only bytes travelled.

code

json · 6 lines
json
{
  "role": "user",
  "amount": 10,
  "amount": 1000000,
  "role": "admin"
}

go deeper

for a junior

Remember that a document can spell the same member name twice even though the resulting mapping holds one value, and that the specification does not say which value survives.

for a middle

Explain the four resolutions — last wins, first wins, reject, keep both — and say which implementation shape produces each. That is the mechanics tier of this question.

for a senior

Show the chain: name the two components that decode the same bytes, state which value each applies, and propose decoding once at the boundary or rejecting repeats there.

for a principal

Frame it as a class, not an instance: any input the grammar permits and the data model leaves undefined is a divergence point, so the design question is how many decoders a payload meets at all.

## The grammar allows what the data model does not define A self-describing text encoding defines two separate things, and conflating them is what makes this question interesting. The **grammar** says which byte sequences are a document: an object is an opening brace, zero or more `name: value` members separated by commas, a closing brace. The **data model** says what a document *means*: an object denotes a mapping, and a mapping holds one value per name. Nothing in the grammar prevents a document from spelling the same name twice. JSON's own text says member names **should** be unique and warns that behaviour is unpredictable when they are not — that is guidance to whoever writes the document, not an obligation the decoder is required to enforce. So `{"role": "user", "role": "admin"}` is syntactically valid and semantically undefined. Undefined does not mean nothing happens: a decoder still has to hand something back to its caller, so it resolves the repeat in whatever way its implementation makes natural. ## The resolutions you actually meet | Resolution | What the caller receives | How it arises | |---|---|---| | Last occurrence wins | the final value | build a map, assign on each member, overwrite silently | | First occurrence wins | the earliest value | insert only when the name is not already present | | Reject the document | a decode error | a strict mode that tracks the names it has seen | | Keep every occurrence | a list, or two callbacks | an order-preserving structure, or an event-streaming decoder | A fifth shape appears in decoders that bind a document straight into a typed structure: the repeat hits one field twice, and whether the second assignment overwrites the first, is ignored, or raises an error is a property of the binder, not of the format. Configuration matters as much as the implementation — the same decoder often has a strict switch that changes which row of the table applies. Every row is a defensible reading of an under-specified input. The vulnerability is not in any one of them. ## Two decoders in one path The shape that turns this into a finding is a request body parsed twice by different code: 1. A filtering proxy reads the body, decodes it with its own decoder, and evaluates a policy against the result — say, the request may proceed only when `role` is `user`. 2. The proxy approves and forwards **the bytes it received**, unchanged. 3. The service behind it decodes those same bytes with a different decoder and acts on *its* resolution. If the proxy keeps the first occurrence and the service keeps the last, `{"role": "user", "role": "admin"}` is approved as a user request and executed as an administrative one. Swap the two resolutions and the attacker simply swaps the order of the members. Neither component is misbehaving by its own specification; the security property lived in the gap between them. The sentence worth carrying out of this: **a check binds to an interpretation, but only bytes travel.** Anything that re-derives meaning downstream is free to derive a different one. ## Variations of the same shape - **Audit divergence.** A logging or audit component decodes the body one way and records `role: user`, while the actor decodes it the other way. The trail then faithfully records a value that was never applied, which is worse than no trail. - **Two decoders in one process.** A schema validator and an object binder are often different pieces of code over the same bytes. The validator may see one resolution and the binder another, so the validated document is not the constructed one. - **Nesting.** The repeat can sit several levels down, inside a member the outer component never inspects because it only reads the fields its policy mentions. ## What to propose - **Decode once and forward what you decoded.** Re-encode the document the proxy accepted, or pass a structured form onward, so no second interpretation exists. The cost is byte fidelity: anything downstream that authenticates the client's exact bytes has to be reworked. - **Fail closed on the ambiguity itself.** Configure strict decoding that rejects a repeated member name outright, at the boundary, so the undefined case never enters the system. This is cheap and does not break byte-level signatures. - **Do not authorise on a field someone re-parses.** Where a decision has been made, forward the decision — not the bytes it was made from — so the downstream component has nothing to reinterpret. ## The general rule Wherever a grammar permits an input that the data model leaves undefined, two independent implementations may diverge on it, and any pair of them in one request path is a candidate bypass. Repeated member names are the most reliable example because the divergence is total: the two decoders return different values, with no error anywhere to notice.

  • Does configuring one decoder to reject repeated member names fix the chain?
    It fixes the chain only if that decoder sits in front of every other one, so the document never reaches a lenient reader. A strict setting on the *backend* alone still leaves an inspecting proxy approving documents that are then rejected — safe, but noisy. A strict setting at the boundary is the useful placement: the ambiguous document is refused before any component has a chance to disagree about it.
  • How does an event-streaming decoder differ from a map-building one here?
    A streaming decoder emits a callback per member, so the caller genuinely sees both occurrences and decides what to do — it can reject, or it can silently keep whichever it assigns last, which reproduces the same divergence one layer up. A map-building decoder makes the decision itself and hands back a single value, so the caller cannot tell a repeat ever happened.
  • Do schema-driven binary encodings have the same problem with repeated fields?
    They have the same input but a defined answer: several schema-driven binary encodings specify that a repeated occurrence of a scalar field means the last one wins, so every conforming decoder agrees. That is the point — the danger is not repetition, it is repetition with no defined resolution, which leaves each implementation to choose.

A form arrives with two lines both labelled 'amount'. The clerk who approves it stamps the first line; the cashier who pays it reads the last. Both follow their own rule, and the money that leaves is the number nobody approved.

saying these in an interview costs you the question

  • Says repeated member names are a syntax error, so no decoder accepts them
  • Assumes every decoder keeps the last occurrence
  • Assumes a specification conformance claim makes two decoders agree on undefined input
  • Calls it a bug in one decoder rather than a property of the chain
  • Thinks validating the decoded copy protects the bytes that are forwarded
  • Believes logging the anomaly is a remedy rather than detection
open as a page

Why does a byte-size cap on a request body fail to protect a recursive-descent decoder from deeply nested input?

level: middleimportance: must knowfreq 56%

basics

~20 s

A byte is cheap and a stack frame is not: a megabyte of opening brackets is about a million nesting levels, and a decoder that recurses per level exhausts its stack long before it exhausts that input. Depth needs its own ceiling.

open as a page

On a public upload endpoint, why must a decoder's size and nesting ceilings be enforced during the parse rather than after it?

level: middleimportance: must knowfreq 62%

basics

~20 s

By the time a parse finishes, the memory, CPU and stack the sender asked for have already been spent, so a later check only reports the damage. A ceiling has to be a counter the parser itself tests as it reads, aborting mid-stream.

open as a page

Why is validating a decoded copy of a request body unsafe when the original bytes are forwarded unchanged?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Validation applies to one decoder's interpretation of the bytes, and that interpretation is thrown away when the raw bytes are forwarded. The next decoder derives its own meaning, so the payload that is used was never the payload that was checked.

open as a page

A binary decoder allocates a buffer from the four-byte length prefix it just read. What does that enable?

level: middleimportance: should knowfreq 48%

basics

~20 s

A few bytes declaring four gigabytes make the server reserve four gigabytes, and the sender never has to produce them. A declared length is a claim about the stream, not authorisation to reserve memory on the sender's behalf.

open as a page

A proxy compares a numeric field against a limit and forwards the body: how can the backend see a different number?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A JSON number is decimal text of unbounded precision, but each decoder maps it onto a concrete type: a binary floating-point value, a fixed-width integer, or an arbitrary-precision decimal. Different targets round, overflow and coerce differently, so one payload yields two numbers.

open as a page

Several components in one request path decode the same payload: what would you change so a decoder disagreement cannot become a bypass?

level: principalimportance: should knowfreq 34%

basics

~20 s

Reduce the number of decodes the payload meets. Decode once at the boundary and pass the decoded or canonically re-encoded form onward, so no component re-derives meaning from attacker bytes; where that is impossible, make the boundary decoder strict and fail closed.

open as a page

You own a public upload API used by many client teams: how do you choose its decode ceilings and handle a caller that legitimately exceeds one?

level: principalimportance: should knowfreq 32%

basics

~20 s

Derive each ceiling from measured traffic rather than a round number, set it per endpoint, roll it out in report-only mode before enforcing, and answer a legitimate outlier with a different request shape instead of a raised cap or a per-caller exemption.

open as a page

Why can bytes placed after a complete JSON document change what a second decoder in the chain sees?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Decoders disagree about what must follow a complete value: some stop and ignore the remainder, some require end-of-input and reject, some read on and return a sequence. The same bytes therefore yield one document, an error, or several.

open as a page

An upload endpoint accepts compressed request bodies. How do you bound what a small compressed body can expand to?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Count decompressed bytes as they are produced and abort the moment the running total crosses an absolute cap, with a secondary check on the output-to-input ratio so a bomb dies early. Measuring the size after decompression has already paid for it.

open as a page