Why can bytes placed after a complete JSON document change what a second decoder in the chain sees?
answer
- what happens after the last brace
- stop, reject, or read on
- one buffer, several complete values
- the inspector reads one, the consumer another
- demand end-of-input at the boundary
basics
~20 sDecoders disagree about what must follow a complete value: some stop and ignore the remainder, some require end-of-input and reject, some read on and return a sequence. The same bytes therefore yield one document, an error, or several.
solid answer
~50 sA decoder has to decide what to do once it has read one complete value. Three behaviours are all in the wild: **stop and return**, leaving the rest of the buffer unread; **require end-of-input**, rejecting anything that follows; and **read a sequence**, returning several values or the last one. So `{"role":"user"}{"role":"admin"}` is one object to the first, an error to the second, and two records to the third. Put two of those in one path — an inspecting component that stops at the first value and a consumer that keeps the last — and the approved document is not the applied one. Lenient extensions widen the same seam: a leading byte-order mark, comments, trailing commas or control characters may be tolerated on one side and refused on the other. The boundary should demand end-of-input and reject extensions.
go deeper
Note that a decoder must decide what to do with bytes after the first complete value, and that different decoders decide differently.
Describe the three behaviours — stop, reject, read a sequence — and what each returns for a buffer holding two complete objects.
Explain why the pairing matters in a real path, and pick the control: refuse anything after the document at the boundary rather than matching settings across services.
Treat decoder leniency as a contract the platform sets once, not a per-service option, and decide what the organisation accepts on machine-to-machine paths.
## The question every decoder has to answer After a decoder reads one complete value from a buffer, something must happen with whatever bytes remain. There is no single right answer, because decoders serve different callers: - A decoder built for **a single document per payload** naturally requires end-of-input and reports an error when bytes follow the closing brace. - A decoder built as **a streaming reader over a long buffer** naturally stops at the end of the value and hands back a cursor, leaving the caller to decide whether to read again. If the caller does not look, the remainder is silently discarded. - A decoder built for **a sequence of documents** — the line-delimited style used for logs and bulk records — reads value after value and returns them all, or, through a convenience entry point, the last one. Each of these is correct for its purpose, and any two of them in one request path disagree about the same bytes. ## The shape of the divergence Consider a body carrying two complete objects back to back: `{"role":"user"}{"role":"admin"}` | Decoder behaviour | Result | |---|---| | Stop after the first value | one object, `role` is `user`, remainder ignored | | Require end-of-input | a decode error, the document is refused | | Read a sequence, keep the last | one object, `role` is `admin` | An inspecting proxy of the first kind approves the request. A consumer of the third kind applies the second object. Nothing logged an anomaly anywhere, because from each component's own point of view the payload parsed cleanly. ## Related leniency that lives in the same seam The trailing-bytes question is one instance of a larger disagreement about what is tolerated around and inside a document: - **A byte-order mark or stray whitespace before the opening brace**, skipped by one decoder and treated as a syntax error by another. - **Comments and trailing commas**, accepted by decoders that implement a common superset of the format and rejected by strict ones. - **Control characters inside strings** that the grammar requires to be escaped, silently accepted by some decoders. - **Truncated input**: a decoder that returns whatever it has assembled so far, against one that demands a complete value. Every item on that list is a way for one payload to mean two things, which is the only ingredient a parser differential needs. ## Why this one is under-tested Functional tests never produce these inputs. Client libraries emit a single well-formed document and stop, so the trailing region is empty in every legitimate request, and the divergence is invisible until someone constructs the input deliberately. Fuzzing that only mutates *inside* the document misses it too — the mutation has to append a second complete value, which a structure-aware generator will not do unless it is told to. The cheap test is direct: take a valid body, append a second complete document that differs in one security-relevant member, send it, and observe which value the system acted on. If the two differ from what the inspecting component logged, the chain has the defect. ## Closing it - **Demand end-of-input at the boundary.** The component that inspects the payload should use a decoder mode that refuses anything after the first complete value, so ambiguous bodies are rejected rather than resolved. - **Turn lenient extensions off on the inspecting side**, and prefer them off everywhere: comments and trailing commas buy nothing on a machine-to-machine path and cost a divergence class. - **Forward the decoded document, not the received buffer**, which removes the trailing region from the consumer's view entirely. - **Do not rely on matching the two decoders' settings by policy.** Settings drift with upgrades and with whoever writes the next service; a boundary that rejects the ambiguous document does not depend on the other side's configuration at all. The framing to keep: the security-relevant property is not "is this document well formed" but "do both decoders agree on where it ends and what it contains". The second is the stronger question, and only the boundary can answer it for the whole chain.
- Why do ordinary tests and structure-aware fuzzing both miss this?Legitimate clients emit exactly one document, so the region after it is empty in every test fixture. A fuzzer that mutates within the document's structure keeps producing one value too. The input that exposes the divergence is a second *complete* document appended to the first, which has to be generated on purpose.
- Is matching both decoders' leniency settings an adequate control?It is a control that expires. Settings drift when a library is upgraded, when a service is rewritten, or when a new consumer joins the path, and nothing fails visibly when they diverge again. A boundary that refuses anything after the first complete value holds regardless of what the other side is configured to do, which is why it is the stronger control.
saying these in an interview costs you the question
- Assumes any decoder rejects bytes after a complete document
- Thinks a well-formed prefix makes the whole payload unambiguous
- Believes ignoring the remainder is the safe behaviour
- Says matching library versions removes the divergence permanently
- Expects ordinary functional tests to surface it