skip to content

A streaming handler writes rows as it produces them, yet the client receives them all at once at the end — why?

level: middleimportance: should knowfreq 50%

answer

  1. a write is not a send
  2. buffers between handler and consumer
  3. flush pushes one layer down
  4. intermediaries buffer whole responses
  5. the client parser can hold too

basics

~10 s

Writing is not sending. The framework buffer, a content encoder, the socket, an intermediary and the client's parser can each hold bytes, so incremental writes arrive as one batch unless every layer is flushed.

solid answer

~40 s

A write only hands bytes to the next layer down, and several layers are allowed to accumulate them: the framework's response buffer, a compression encoder that needs a block's worth of input, the socket's own batching, a reverse proxy or edge cache that buffers the whole upstream response, and a client whose parser surfaces nothing until it has a complete record. An explicit `flush` pushes one layer's contents down, so on a path with an encoder you must flush both, and no application flush can override an intermediary that buffers by policy. Diagnose it by bisecting the path — read directly from the server with a byte-level consumer that timestamps arrivals, then repeat through each hop, and the hop where the timing collapses is the culprit.

go deeper

for a junior

Take away one sentence: a write only reaches the next buffer, and something has to push it further. If output arrives all at once, some layer on the path is holding it.

for a middle

Name the layers and what each one waits for — framework buffer, encoder block, socket batching, intermediary policy, client parser — and explain that a flush acts on exactly one layer.

for a senior

Show the bisection: timestamped byte-level reads directly from the process, then through each hop, so you find the buffering layer instead of guessing. Know that edge buffering is configuration, not code.

for a principal

Decide the flush unit as a policy — per record for low-rate streams, per batch or interval for high-volume ones — and make the end-to-end path, including edge buffering rules, part of what the streaming feature owns and tests.

Writing is not sending. A handler's write call hands bytes to the layer below it, and every layer between the handler and the consumer's parser is allowed to hold onto those bytes until it has a reason to pass them on. When all of them hold, the stream still *works* — it is simply delivered as one batch at the end, which defeats the entire point of streaming. ## The layers that can hold your bytes 1. **The framework's response buffer.** Most frameworks accumulate small writes and pass them down when the buffer fills, because one syscall per row would be ruinous. A handler writing short rows may fill that buffer only after thousands of them, or never. 2. **A content encoder.** Compression works on blocks and needs input before it can emit anything. An encoder placed in front of the output has its own buffer, and flushing the writer without flushing the encoder simply moves the data from one buffer to another. 3. **The socket send buffer and the transport's own batching.** The operating system may coalesce small writes to avoid sending many tiny segments, which adds a delay that is invisible from application code. 4. **An intermediary.** A reverse proxy, gateway, load balancer or content-delivery network may buffer the whole upstream response before forwarding any of it — often the default, because it frees the upstream connection sooner. This is one of the most common causes in production, and **no application-level flush can override it**. 5. **The consuming client.** A client that reads complete lines, complete records or a fully parsed document surfaces nothing to the caller until its own parser is satisfied, even when the bytes arrived on time. ## What a flush does, and what it does not A **flush** is an explicit instruction to push whatever is buffered at that layer down to the next one. It only affects the layer you call it on, so on a path with an encoder and a writer, both must be flushed. And it is always a request: a flush can push bytes into the socket, but it cannot make an intermediary forward them, and it cannot make a client's parser emit an incomplete record. Flushing is not free, either. Every flush costs at least one write down the stack and typically one network segment, so flushing per byte or per tiny record trades throughput and CPU for latency. The usual compromise is to flush on a **meaningful unit** — a complete record, a batch of rows, or a time interval — rather than on every write. | Symptom | Likely layer | What to change | |---|---|---| | Nothing until the very end, any size | An intermediary buffering the whole response | Disable response buffering on that hop for this route | | Output arrives in large, even blocks | Framework or encoder buffer | Flush at record boundaries, flush the encoder too | | Bytes arrive on time but the app sees nothing | The consuming client's parser | Read incrementally, parse per record | | Fine locally, batched in production | Only production has the intermediary | Compare a direct call with a call through the edge | ## How to diagnose it in the right order The path has several buffers and guessing wastes a day, so bisect it instead: 1. Read the endpoint **directly from the server process**, with a consumer that prints bytes as they arrive with timestamps. If the timing is correct here, the handler and framework are fine and the problem is in front of them. 2. Repeat **through each hop** — the reverse proxy, then the edge — and the hop where the timing collapses is the buffering one. 3. If even the direct read batches, remove the encoder and re-test to separate the compressor from the writer, then check whether the handler flushes at all. 4. Only then look at the client library, by comparing what a raw byte-level reader sees against what the high-level parser surfaces. ## Design consequences - **Choose a flush unit deliberately.** Per record is right for a low-rate event stream; per batch or per time window is right for a high-volume export where per-record flushing would dominate the cost. - **Treat the whole path as part of the feature.** A streaming endpoint that is never exercised through the same hops as production is not tested. Buffering at the edge is configuration, and configuration is where the behaviour actually lives. - **Be careful with compression on low-rate streams.** An encoder needs input to emit a block, so it can add latency precisely where latency is the feature; either flush it per record, accepting worse compression, or leave the route uncompressed. - **Do not confuse a hold with a hang.** A stream held in a buffer and a stalled producer look identical from the outside. Timestamped byte-level reads are what distinguish them, which is why the diagnosis starts there. The one-line version worth remembering: a write reaches the next buffer, a flush pushes it one layer further, and delivery to the consumer requires **every** layer on the path to cooperate.

  • Why not flush after every single write?
    Because each flush costs a push down the stack and usually a network segment, so per-byte flushing burns CPU and bandwidth for latency nobody asked for. Flush on a meaningful unit instead — a complete record, a batch, or a time interval — matched to how fast the stream actually produces.
  • The stream is incremental locally but batched in production. Where do you look first?
    At the hops only production has. A reverse proxy or edge cache that buffers the full upstream response is the common cause, and it is usually the default because it frees the upstream connection sooner. Compare a direct read against a read through the edge to confirm before touching the handler.
  • How can compression turn a working stream into a batched one?
    An encoder emits output only when it has enough input for a block, so it holds records back regardless of what the writer does. Either flush the encoder at each record and accept a worse compression ratio, or leave that route uncompressed.

saying these in an interview costs you the question

  • Assumes a write reaches the client immediately
  • Flushes the writer but not the content encoder in front of it
  • Believes an application flush overrides a buffering proxy
  • Never tests the stream through the production path
  • Flushes on every byte and calls the throughput loss unavoidable
  • Blames the handler when the client's parser is what is holding records