When does a streaming parser that emits events beat materialising the whole decoded document in memory?
answer
- memory ceiling versus payload size
- one forward pass, no going back
- work starts before the last byte
- hand-written state machine is the price
- buffering it yourself rebuilds the tree
basics
~20 sStreaming wins when the document is large relative to memory, or unbounded, and the reader consumes it in one forward pass — peak memory then tracks the largest single value, not the payload. It loses when the work needs random access or back-references.
solid answer
~50 sA tree-building reader decodes the entire payload into a connected in-memory structure before the application sees anything; peak memory scales with the payload and the whole cost is paid before the first result. An event-emitting reader hands out one token at a time — field start, value, container end — so peak memory is bounded by the largest single value plus the reader's own state, and work can begin before the last byte arrives. Streaming wins for large or unbounded payloads consumed in a single forward pass, especially when the reader retains only an aggregate. It loses when the computation needs random access, back-references or ordering that the byte order does not provide, because then you buffer the data yourself and have rebuilt the tree by hand, minus the library's care. For a two-kilobyte request body the tree is almost always the right answer.
code
pseudocode · 12 linestotal = 0
reader = open_event_reader(input_bytes)
while reader.has_next():
event = reader.next()
if event.kind == FIELD and event.name == "amount":
total = total + reader.read_number()
else:
reader.skip_value() # nothing is retained
return total
# peak memory: one event, one open-container stack, one accumulatorgo deeper
Know the two shapes: one reader hands you the whole decoded document, the other hands you one piece at a time. The second uses far less memory for big payloads.
Explain the memory bound precisely — largest single value plus reader state, not payload size — and name the cost you pay for it: a forward-only pass and a state machine you write yourself.
Show the operating judgment: the ceiling is concurrency times per-request peak, malformed input may surface mid-computation, and a streaming consumer must be written so partial work can be discarded.
Decide where the boundary lives. Often the better answer is to change the payload shape or push aggregation upstream, so no consumer needs a hand-rolled state machine at all.
## Two shapes of reader Every decoder sits somewhere between two shapes. - A **tree-building (document) reader** consumes the whole payload and returns one connected structure — maps, lists and scalars — that the application then navigates. Everything is available at once, in any order. - An **event-emitting (streaming) reader** exposes the payload as a sequence: container start, field name, scalar value, container end. Either the reader calls into your handler (push) or you ask for the next event (pull). Nothing is retained unless your code retains it. The pull form is usually the easier one to build real logic on, because the application keeps its own call stack and can stop; the push form inverts control and forces you to carry state in a handler object. ## What streaming actually buys 1. **A memory ceiling that does not track the payload.** Peak memory is roughly the largest single value the reader must present, plus its own stack of open containers. A payload larger than memory becomes processable rather than impossible. 2. **Work that overlaps arrival.** The first events are available after the first bytes, so decoding can proceed while the rest of the payload is still in flight, and a rejectable document can be rejected early instead of after it has been fully materialised. 3. **Allocation you control.** In a tree reader every value is allocated whether or not you look at it. Streaming makes retention an explicit decision by your code, which is what turns a large document into a small running aggregate. ## What streaming costs - **A hand-written state machine.** "Which container am I inside, and which field was that?" becomes your problem. That code is where the bugs go. - **No random access.** An event stream is one forward pass. Anything that needs to look backwards must have been remembered on the way past. - **Retention you may not escape.** If the computation ultimately needs most of the document, you buffer it yourself — and a hand-rolled buffer is a tree without the reader's validation, sharing and lifecycle care. - **Deferred structural errors.** A tree reader can reject a malformed payload before any application logic runs. With streaming, half the work may already have side effects when the failure surfaces, so a streaming consumer usually needs to be written so partial work is discardable. | Property | Tree reader | Event/streaming reader | |---|---|---| | Peak memory | scales with the payload | largest single value plus reader state | | First result available | after the last byte | after the first relevant event | | Access pattern | random, repeatable | one forward pass | | Application complexity | low — navigate a structure | higher — explicit state machine | | Failure timing | before your logic runs | possibly mid-computation | ## Choosing, in practice The honest decision rule has two inputs: **how big the payload can get**, and **what fraction of it the computation must hold at once**. - Small payload, arbitrary access — build the tree. A few kilobytes of request body costs a handful of short-lived objects, and the state machine buys nothing. - Large or unbounded payload, single forward pass, small retained result — stream. This is the case the mechanism exists for: a running total, a filter, a projection into a narrower shape. - Large payload, but the work needs sorting, joining or resolving a reference to something defined later — streaming does not save you. Either buffer deliberately, or restructure the payload so the ordering the consumer needs is the ordering the producer writes. - Medium payload in a latency-bound request path — measure. Streaming removes the tree's allocation but adds per-event overhead, and for moderate sizes the two can land close enough that clarity should decide. One more caution about scale: the ceiling that matters is not one document but **concurrency times document size**. A tree reader that is comfortable at one request becomes a memory problem at two hundred concurrent ones, and that is the threshold at which teams usually discover streaming. Ecosystems also differ in what their default reader is — some libraries give you an event reader first and a tree on top of it, others only expose the tree — so the choice may cost more in one stack than another, but the trade-off itself does not change.
- Which computations cannot be done in a single streaming pass without buffering?Anything that needs data the pass has not reached yet or has already left behind: sorting, joining two collections, resolving a reference to a field defined later, or any aggregate over a window wider than what you chose to remember. You can still stream, but you must buffer the part you need, and that buffer is the real memory bound.
- Does streaming reduce total processor time, or only memory?Mostly memory and time-to-first-result. Total processor time often improves because fewer values are converted and allocated, but per-event dispatch adds overhead, so a streaming reader that ends up materialising everything anyway can be slower than the tree it replaced. Measure rather than assume.
- How does concurrency change the decision?The memory bound is concurrent requests multiplied by per-request peak. A tree reader that is fine for one payload can exhaust memory at high concurrency, so the threshold for switching to streaming drops as the concurrency target rises. Size the ceiling against the load you intend to serve, not one request.
saying these in an interview costs you the question
- Claims a streaming reader never buffers anything
- Says streaming is always faster than building a tree
- Uses a tree reader for payloads with no size bound at all
- Ignores that concurrency multiplies per-request peak memory
- Forgets that a forward-only pass cannot resolve back-references
- Assumes malformed input is still rejected before any work happens