A client reads a streamed response far slower than the handler produces it — what does the framework do with the unsent data?
answer
- the difference has to go somewhere
- block, buffer or drop
- unbounded queue per connection
- a not-ready signal only helps if honoured
- fast test readers hide all of it
basics
~20 sA framework does one of three things: the write blocks and pins whatever runs the handler, unsent bytes queue in memory and grow per slow connection, or the write signals not-ready so the handler must pause producing.
solid answer
~50 sOnce the transport stops accepting bytes the difference between produce rate and drain rate has to land somewhere. A blocking write API parks the call, so memory stays flat but the thread or worker is held for as long as the slowest consumer takes. A non-blocking write with an unbounded queue returns immediately and accumulates unsent bytes per connection — it tests perfectly and fails as an out-of-memory event in production. A write that returns a not-ready signal or a completion handle is the only shape that bounds memory and the producer without discarding anything, and it works only if the handler actually stops producing when told. For live updates, a bounded queue with an explicit discard policy is usually more honest than holding a backlog nobody wants, as long as consumers are told discards can happen.
go deeper
Learn the three outcomes by name: the write waits, the bytes pile up in memory, or something is discarded. A slow consumer always forces one of them.
Explain which shape follows from which write API, and why the queueing shape is the dangerous one — it grows per connection and is invisible against fast test readers.
Show the operational side: bounded queues with a chosen policy, write and idle-progress timeouts, per-connection buffered-bytes as a distribution, and a throttled reader in the test plan.
Decide the policy once, as a platform rule. Which streams may block, which must bound and discard, what consumers are promised about completeness, and what per-instance stream cap follows from the memory that policy implies.
A streaming handler produces at whatever rate its source allows. The consumer reads at whatever rate its network, its CPU and its own patience allow. When the second is slower than the first, the difference has to go somewhere, and there are essentially three places it can go: into the producer (it waits), into memory (the bytes pile up), or into the bin (something is discarded). In practice a framework's write API lands in one of these three shapes, and knowing which one you are on decides how the service fails under load. ## Why the difference cannot just vanish The socket's send buffer is finite. Once it is full the transport will not accept more bytes until the consumer acknowledges what it already has, which it does at its own pace. From the handler's point of view the write call has nowhere to put the data. What happens next depends entirely on the shape of the write API the framework hands the handler. ## The three outcomes | Shape of the write API | What happens on a slow reader | The resource that runs out | |---|---|---| | Blocking write | The call parks until space frees up | The thread or worker, held for the slowest client's duration | | Non-blocking write with an unbounded queue | Unsent bytes accumulate in process memory | Memory, and the whole instance with it | | Write that signals "not ready" | The producer is expected to pause and resume | Nothing, if the handler honours the signal | 1. **It blocks.** The write parks until the transport accepts more. Nothing is lost and memory stays flat, but whatever executes the handler is pinned for as long as the slow consumer takes. In a model that dedicates a worker per in-flight request this converts one slow reader into one unavailable worker, and enough of them into a service that accepts no new requests at all. It is also why a single slow reader can hold a pooled database cursor open for minutes. 2. **It queues.** The write returns immediately and the framework holds the unsent bytes. Memory now grows at the difference between produce rate and drain rate, per slow connection. This is the most dangerous shape, because it looks perfect in testing — fast readers barely queue anything — and it fails as an out-of-memory event under exactly the conditions that matter: mobile clients, congested links, a consumer doing expensive work per record. 3. **It pushes back.** The write returns a signal, a completion handle or a "wait for drain" callback, and the contract is that the handler stops producing until the transport is ready again. This is the only shape that bounds memory and the producer without discarding anything — but only if the handler actually honours it. A handler that ignores the signal and keeps writing has silently opted into shape 2. ## Discarding, and when it is correct For a stream of live updates, buffering is often the wrong answer anyway: a consumer minutes behind does not want the backlog, it wants the current state. Those streams are built on a **bounded** per-connection queue plus a discard policy: - **Drop oldest** — keep the newest state, right for telemetry, prices, positions and progress updates. - **Drop newest** — keep an intact prefix, right when early records are the meaningful ones. - **Collapse** — replace the queued value for a key with the latest one, so a slow consumer gets a compacted view rather than a backlog. - **Disconnect the consumer** — the honest option when the data must not be silently incomplete: close the stream and let the client come back for a fresh, consistent read. The important half of that decision is the contract. A consumer that does not know discards are possible will treat the stream as complete and be quietly wrong, so the discard policy has to be part of the documented interface, not an implementation detail. ## Operating this in production - **Measure per-connection buffered bytes**, not just total memory. The distribution is what tells you a few consumers are dragging, long before the aggregate looks alarming. - **Bound every queue.** An unbounded buffer is a deferred out-of-memory event with a slow fuse; a bounded one with an explicit policy fails in a way you chose. - **Set a write or idle-progress timeout**, so a consumer that has effectively stopped reading does not hold resources indefinitely. - **Cap concurrent streams per instance**, because each one holds a connection, a buffer and, depending on the execution model, something more expensive. - **Watch for the load-test blind spot.** Fast local consumers never trigger any of this; you have to test with a deliberately throttled reader to see which of the three shapes you are actually running. ## The takeaway Ask one question of any streaming code: *when the consumer is slower than the producer, what gives?* If the answer is "the thread waits", you have a concurrency ceiling. If it is "memory grows", you have an incident waiting for a bad network. If it is "the producer pauses", check that the handler really honours the signal — and if the data is live, ask whether the right answer is to discard rather than to hold anything at all.
- Why does the queueing shape pass load tests and fail in production?Because load generators read as fast as the link allows, so nothing ever queues. Queue growth needs a consumer slower than the producer — a congested mobile link, or a client doing real work per record. Test it with a deliberately throttled reader, not more concurrency.
- Which metric exposes this before the instance dies?Buffered bytes per connection, as a distribution rather than a total. A handful of connections holding large buffers is the early signal; aggregate memory only moves once the problem is already large. Pair it with the count of streams that have made no write progress recently.
- When is discarding records the right design rather than a bug?When the consumer wants current state rather than history — telemetry, prices, positions, progress. Keep a bounded queue, drop oldest or collapse per key, and make the policy part of the documented contract so no consumer treats the stream as complete when it is not.
- Does a blocking write ever lose data?No. It preserves every byte and keeps memory flat; the cost is paid in hold time instead. Whatever executes the handler stays pinned for the slowest consumer's duration, along with any source cursor or pooled connection it is reading from.
A conveyor belt feeding a packer who cannot keep up: either the belt stops and the line behind it waits, boxes pile up on the floor beside it, or someone pushes the surplus off the end. There is no fourth option, and the pile on the floor is the one that looks fine until the room is full.
saying these in an interview costs you the question
- Assumes the framework handles slow readers automatically
- Treats an unbounded per-connection buffer as safe because tests pass
- Ignores a not-ready signal and keeps writing
- Thinks a blocking write loses data rather than hold time
- Load-tests only with readers as fast as the producer
- Discards records without telling consumers it can happen