skip to content

How would you decide which endpoints buffer their response fully and which stream, given that buffering keeps failures mappable?

level: principalimportance: should knowfreq 38%

answer

  1. unflushed bytes are still reversible
  2. buffering trades memory for error fidelity
  3. memory cost is threshold times concurrency
  4. stream only when size or duration forces it
  5. streaming endpoints owe a completeness contract

basics

~10 s

Treat it as buying error fidelity with memory. Buffer anything small and bounded so late failures still map to a real status; stream only where size or duration forces it, under a completeness contract.

solid answer

~50 s

The trade is memory and time-to-first-byte against the ability to change your mind. While nothing is flushed, a late exception still reaches the central mapper and the client gets a real status; the moment you stream, every failure after the first byte becomes a truncation. So the default should be **buffered with a bounded limit**, and streaming should be an explicit, justified exception for responses that are large, unbounded, or long-lived — exports, downloads, and event feeds. For those, the price is paid elsewhere: a completeness contract the consumer can check, an abort rather than a clean ending on failure, and a post-commit failure signal in the platform's telemetry. The useful policy knob is a **buffer size threshold**: below it a response stays replaceable; above it the framework flushes and the endpoint has opted into the streaming rules, whether its author meant to or not.

go deeper

for a junior

Know the two shapes: a response held in memory until it is finished, and one written out as it is produced. The first can still be replaced by an error response; the second cannot once it starts.

for a middle

Explain the trade concretely: memory and first-byte latency against the ability to return a proper error status for a late failure, and note that memory scales with concurrency rather than with one response.

for a senior

Argue the default and the exceptions by response shape, and describe what a streaming endpoint must add in return: work done before the first byte, a completeness contract, abort on failure, and its own failure signal.

for a principal

Own it as platform policy: one threshold, visibility into which endpoints cross it, and an approval path that attaches obligations to streaming instead of leaving the error contract to each handler author.

## The trade in one line Every byte you have not yet flushed is a decision you can still take back. Buffering the whole response keeps the status unchosen until the handler finishes, so the ordinary error machinery — mapper, shared failure body, fallback — works for a failure raised on the last line of the handler exactly as it does for one raised on the first. Streaming spends that option in exchange for bounded memory and an early first byte. This is why the question is a policy question rather than a per-handler one. Left to individual authors, streaming spreads by imitation, and a service ends up with endpoints that cannot report their own failures for no benefit anyone chose. ## What buffering buys - **Failures stay mappable.** A serialisation error, a lazy lookup that throws while the body is being produced, or a permission check that only fires on the last record all still become a proper error status. - **The status can reflect the finished result.** The handler can discover that the result is empty and say so, rather than having already promised a success. - **Retries stay meaningful.** A client receiving a clean error status knows nothing was delivered; a client receiving a truncated success knows almost nothing. - **Observability is free.** The status recorded is the status the client got, so the default dashboards are accurate. ## What buffering costs - **Memory proportional to the body, times concurrency.** This is the real constraint: it is not one large response that hurts but the hundredth concurrent one, and the failure mode is a whole-process memory problem rather than one bad request. - **Time to first byte** grows to the full generation time, which matters for anything a human waits on. - **It caps response size** at whatever the platform is prepared to hold, which is the correct answer for most endpoints and an impossible one for exports. ## A policy grid | Response shape | Default | Why | |---|---|---| | Small bounded document, size known within a limit | Buffer | Failures map normally; memory cost is trivial | | Paginated collection with an enforced page cap | Buffer | The cap is what makes buffering safe; the pagination already bounds it | | Large export or file download | Stream | Buffering it is a memory incident waiting for concurrency | | Long-lived event or progress feed | Stream | There is no finished body to buffer; it is unbounded by definition | | Anything whose size is caller-controlled and uncapped | Cap it first, then buffer | An unbounded parameter is the actual defect; streaming only hides it | ## The hybrid worth standardising Rather than a per-endpoint flag, the platform can set a **buffer threshold**: bytes accumulate in memory, and the response is committed only when the buffer fills. Everything below the threshold is fully replaceable and behaves as if it had been buffered; everything above it streams. That gives three useful properties. It makes the boundary a single tunable number rather than scattered handler decisions. It means most responses in a typical service never commit early at all. And it makes the threshold an honest capacity statement: threshold multiplied by expected concurrency is the memory you have chosen to spend on error fidelity. The trap is that it is silent. An endpoint whose responses grew past the threshold starts streaming without anyone changing a line, which means its late failures stop mapping. A service that leans on the threshold should also surface which endpoints are crossing it. ## What streaming endpoints owe in return Approving an endpoint for streaming should come with obligations, not just permission: 1. **Do the risky work before the first byte** — authorisation, existence checks, parameter validation, opening the upstream source — so the class of failures that can still occur mid-body is as small as possible. 2. **Publish a completeness contract**: an end-of-stream marker, a declared count, or a format that cannot be cut without breaking, so a consumer can tell a short answer from a whole one. 3. **Abort rather than end cleanly** when production fails, so truncation is detectable. 4. **Emit the post-commit failure signal**, since the request's recorded status will be a success. 5. **Release the resources behind the stream** on the abort path. ## The judgment an interviewer is listening for The weak answer is a preference — *streaming scales better* — applied uniformly. The strong answer treats buffering as the default because it preserves the error contract, names the concrete conditions under which that default is unaffordable, and accepts that the exception comes with engineering work rather than just a different call in the handler.

  • What is the risk of relying on a buffer-size threshold to decide which responses stream?
    It changes behaviour silently. An endpoint whose responses grow past the threshold begins committing early, so its late failures stop becoming error statuses and start becoming truncations, with no code change to review. Surface which endpoints cross the threshold rather than only setting it.
  • Why is an uncapped, caller-controlled response size a defect rather than a reason to stream?
    Because the caller then decides how much work and memory the service spends per request. Streaming hides the symptom while leaving the exposure. Cap the result first — with pagination or an explicit limit — and then choose buffering or streaming on the bounded size.
  • If buffering preserves the error contract, why not buffer everything?
    Memory scales with body size times concurrency, and time to first byte grows to the full generation time. For exports and long-lived feeds there may be no finished body to hold at all. Buffering is the right default, not a universal answer.

saying these in an interview costs you the question

  • Picks streaming everywhere because it sounds more scalable
  • Sizes the buffer for one request instead of size times concurrency
  • Treats streaming as a free choice with no obligations attached
  • Streams to work around a response size the caller can grow without limit
  • Ignores that a size threshold silently moves endpoints into streaming
  • Assumes late failures still map to a status once the body is streaming