skip to content

Chunked & Streamed Bodies

Producing a body incrementally: a writer the framework drains, explicit flushes, headers committed at first write, a writer stalled by a slow reader. Asked because a buffer makes a stream a batch.

on this pageshow

questions

5

Why would a web framework handler stream a large response body incrementally instead of building the whole body in memory first?

level: juniorimportance: must knowfreq 60%

answer

  1. size times concurrency, not size
  2. bytes moving, not accumulating
  3. time to first byte
  4. small write buffer per request
  5. length unknown until the end

basics

~20 s

Streaming keeps memory flat and bytes moving. The handler writes pieces the framework sends as they are produced, so a large response costs a small buffer per request instead of its full size, and the client sees data sooner.

solid answer

~40 s

A buffered handler builds the entire body and hands it over, so the process holds the whole thing in memory once per in-flight request — the cost is body size times concurrency, which is what turns a large export endpoint into an out-of-memory incident. A streamed handler writes into a sink the framework drains, so only a small write buffer plus the piece being produced is resident, and the first bytes reach the client almost immediately instead of after the full build. The tradeoffs: the length is not known up front, so no size-based progress is possible, the response commits as soon as the first bytes flush, and any cursor or connection the handler reads from stays open for the whole transfer.

go deeper

for a junior

Recall the core contrast: buffered means the whole body sits in memory before anything is sent; streamed means pieces go out as they are produced. Memory per request and time to first byte are the two things that change.

for a middle

Explain the arithmetic — peak memory is body size times concurrent requests, not size alone — and name what streaming costs: no length up front, an early commit, and a source held open for the whole transfer.

for a senior

Show that you decide from measurements: peak memory per in-flight request, time to first byte, and how long a pooled connection or cursor is pinned by the slowest reader. Say where you would keep buffering deliberately.

for a principal

Frame it as a budget question. Per-request memory and hold time set the concurrency ceiling of an instance, and streaming trades a clean failure path for that ceiling — decide which endpoints are worth that trade and make it a standard, not a per-handler improvisation.

A response body can be produced in one of two shapes, and the choice decides how much memory the service holds, how soon the client can start working, and what the framework is still able to do about the response later. ## The two shapes In the **buffered** shape the handler builds the whole body first — a string, a byte array, a serialized object — and hands that finished value to the framework. The framework knows the exact size before anything leaves the process: it can set a length header, compress the whole payload, and write it in one pass. The handler has already returned by the time the first byte reaches the network. In the **streamed** shape the framework hands the handler a sink — a writer, an output stream, a generator it pulls from, or a callback it invokes with each piece — and the handler produces the body in parts. Bytes leave the process while the handler is still running, and the handler is done only when it closes the sink. ## Where the memory goes The decisive number is rarely the size of one response. It is **size x concurrency**, because every in-flight request holds its own copy: - A 200 MB export built in memory costs about 200 MB of live data per concurrent download; twenty of them is roughly 4 GB, on top of everything else the process is doing. - Serialization usually costs more than the final size, because an intermediate representation and the encoded bytes coexist for a moment. - The streamed version holds only a small write buffer (typically kilobytes) plus whatever single row, record or chunk is being produced right now — a per-request cost that does not grow with the body. - The failure is non-linear. Memory looks fine at low concurrency and the service falls over the moment concurrency crosses a line, which is why this shows up as an incident rather than as a slow regression. Large, short-lived buffers are also the worst possible shape for a garbage-collected runtime: they are allocated in one piece, survive long enough to be promoted, and are then discarded, which is exactly the pattern that produces long pauses. ## What the client experiences Streaming does not make the network faster and does not reduce the bytes transferred. What it moves is **when** the first useful byte arrives. A buffered build of a large body looks to the client like a hang — nothing arrives until the whole thing is ready, which can exceed a client or intermediary read timeout even though the server is working correctly. A streamed body starts arriving almost immediately, so a consumer can parse, render or write to disk progressively, and the connection shows steady activity instead of silence. ## What streaming costs | Concern | Buffered body | Streamed body | |---|---|---| | Peak memory per request | Proportional to the body | A small write buffer | | Time to first byte | After the whole body is built | Almost immediately | | Length known up front | Yes, the framework can set it | No, unless you know it another way | | Resources held during transfer | Released when the handler returns | Held until the last byte is written | | Reacting to a mid-body failure | Full freedom, nothing has been sent | Constrained, the response is already committed | Three of those deserve emphasis. First, without a known length the framework must use a framing that does not need one, and the client cannot show a percentage-of-total progress bar. Second, a cursor, connection or file handle the handler reads from stays open for the whole transfer, so the slowest client now decides how long a scarce resource is pinned. Third, the response commits as soon as the first bytes are flushed, which changes what can still be done if the producer fails halfway. ## When buffering is the better choice Streaming is not a free upgrade. Prefer a buffered body when: 1. The body is small — for typical record-sized responses the buffer is irrelevant and the buffered path is simpler and easier to fail out of. 2. The exact length or a digest of the whole body is required before sending. 3. Failing cleanly matters more than starting early, because a buffered response can still be replaced by an error response. 4. The result is cacheable or reusable, in which case something has to hold the whole value anyway. ## A working rule Stream when the body is unbounded or large relative to the per-request memory budget, or when the first part is useful before the last part exists. Buffer otherwise. And measure the real thing: peak memory per in-flight request and time to first byte are the two numbers that tell you which shape you actually have, regardless of which one the code appears to describe.

  • Does streaming make the response arrive faster overall?
    No. The same bytes cross the same link, so the time to the *last* byte is roughly unchanged. What streaming improves is the time to the *first* byte and the memory held while the transfer runs. Claiming a throughput win is a common overstatement.
  • When is buffering the whole body still the right call?
    When the body is small, when the exact length or a digest of the whole payload is needed before sending, when the value is cached or reused anyway, or when failing cleanly matters more than starting early — a buffered response can still be replaced by an error response, a committed one cannot.
  • What resource cost does streaming add that buffering does not have?
    Hold time. A cursor, connection or file handle the handler reads from stays open until the last byte is written, so the slowest client decides how long that resource is pinned. With a buffered body the source is released as soon as the value is built, well before the transfer finishes.

saying these in an interview costs you the question

  • Thinks streaming makes the total transfer faster or smaller
  • Believes streaming removes the need for any buffer at all
  • Streams every response, including small record-sized ones
  • Reasons about one response's size instead of size times concurrency
  • Assumes a percentage progress bar still works without a known length
  • Forgets the source cursor stays open for the whole transfer
open as a page

In a web framework, what becomes fixed the moment a streaming handler's first bytes are flushed to the client?

level: middleimportance: must knowfreq 58%

basics

~20 s

Headers travel ahead of the body, so the first flush freezes the status code, every header field, cookies and the body's framing choice. Later changes to them are lost; only more body bytes can follow.

open as a page

A streaming handler writes rows as it produces them, yet the client receives them all at once at the end — why?

level: middleimportance: should knowfreq 50%

basics

~10 s

Writing is not sending. The framework buffer, a content encoder, the socket, an intermediary and the client's parser can each hold bytes, so incremental writes arrive as one batch unless every layer is flushed.

open as a page

A client reads a streamed response far slower than the handler produces it — what does the framework do with the unsent data?

level: seniorimportance: should knowfreq 46%

basics

~20 s

A framework does one of three things: the write blocks and pins whatever runs the handler, unsent bytes queue in memory and grow per slow connection, or the write signals not-ready so the handler must pause producing.

open as a page

How would you decide how many minutes-long, held-open streaming responses one instance should carry, and how would you shed them safely?

level: principalimportance: should knowfreq 36%

basics

~10 s

Measure what one open stream costs on the real execution model, find the limit that binds first, cap admission below it, bound every stream's lifetime so the fleet stays drainable, and bound per-stream buffers.

open as a page