skip to content

A team ships HTML streaming: locally the page's shell appears while the server is still working, but in production the browser gets nothing until the whole response is ready and the shell never appears early. How do you diagnose where the streaming is being lost?

level: seniorimportance: should knowfreq 38%

answer

  1. compare first byte against last byte
  2. bisect the path, hop by hop
  3. one buffering layer erases it all
  4. Content-Length where you expected chunked
  5. compression that flushes only at close

basics

~20 s

Compare time to first byte with total transfer time for the document at each hop — browser, CDN, reverse proxy, origin. Whichever hop is the first to report them as nearly equal is the one buffering the response.

solid answer

~50 s

Streaming is only preserved if every layer on the path forwards partial bodies, so the diagnosis is a bisection along that path. Start at the browser: in the network panel, a streamed document shows a short wait phase and a long content-download phase, while a buffered one shows the opposite. Then request the origin directly, bypassing the CDN and any reverse proxy, and compare the same two timings — `curl` reporting `time_starttransfer` far below `time_total` means the origin is streaming correctly and something in front of it is collecting the body. The usual culprits are a reverse proxy with response buffering enabled, a CDN that buffers to compute or cache a full response, a compression layer that only emits output when the stream closes, and application middleware that wraps the response to rewrite or measure the body. A `Content-Length` header on a response you expected to be chunked is a strong tell: something upstream had to know the whole body to produce it.

code

bash · 9 lines
bash
# Origin direct — expect ttfb << total when streaming works
curl -N -s -o /dev/null \
  -w 'ttfb=%{time_starttransfer}s total=%{time_total}s\n' \
  https://origin.internal/product/42

# Same route through the public edge — if ttfb jumps to meet total, the edge buffers
curl -N -s -o /dev/null \
  -w 'ttfb=%{time_starttransfer}s total=%{time_total}s\n' \
  https://www.example.com/product/42

go deeper

for a junior

Know that streaming only works if every layer between the server and the browser passes bytes along, and that a proxy or a CDN in the middle can hold the whole response.

for a middle

Be able to name the measurement — time to first byte versus total transfer time — and read a network panel entry to tell a long wait phase from a long download phase.

for a senior

Show a methodical bisection from origin outward with a repeatable command, name the realistic culprits including compression and response-wrapping middleware, and read the headers for tells such as an unexpected Content-Length.

for a principal

Treat streaming as a property that must be asserted, not assumed: put a synthetic guard on it, decide whether contested routes belong behind a buffering edge at all, and weigh the operational complexity against the perceived-speed gain.

## The property that must hold end to end Streaming survives only if *every* participant forwards bytes as they arrive. Origin, application middleware, compression, reverse proxy, CDN, and finally the browser — one buffering layer anywhere collapses the whole thing back to a single delivery, and it looks identical to a server that never streamed at all. Locally there are no intermediaries, which is exactly why the bug shows up only in production. So the diagnosis is a bisection: find the first hop, walking from the origin outward, at which the response stops being incremental. ## The single measurement you repeat at each hop Everything rests on comparing two numbers for the document request: - **time to first byte** — when the first body byte arrives; - **total transfer time** — when the last byte arrives. Streaming: first byte is small, total is large, and the gap is the server's think time. Buffered: the two are close together and both large. ```bash curl -N -s -o /dev/null \ -w 'ttfb=%{time_starttransfer}s total=%{time_total}s\n' \ https://origin.internal/product/42 ``` `-N` disables curl's own output buffering so you are measuring the network, not the tool. Run this against the origin directly, then against the proxy, then against the public CDN hostname. The first URL where `ttfb` jumps up to meet `total` names the buffering hop. ## The suspects, in the order they usually turn up **A buffering reverse proxy.** The most common cause. A proxy configured to read the upstream response into memory before forwarding it will hold the entire body; in nginx this is `proxy_buffering`, which is on by default, and it can be disabled per response by sending the `X-Accel-Buffering: no` header from the application. Whatever the product, the concept is the same: find the response-buffering setting and turn it off for the streaming routes. **A CDN edge that buffers.** Some edges collect a full response before serving it — to cache it, to compute a length, or because streaming is opt-in for that product. Testing origin-direct versus edge isolates this in one step. **Compression that flushes only at the end.** Compressors work on blocks and, left alone, emit nothing until the stream closes, which is a perfectly correct way to produce a smaller file and a perfectly effective way to destroy streaming. Streaming-aware compression flushes at each chunk boundary — a sync flush per chunk — trading a little compression ratio for incremental delivery. A body that is compressed and arrives all at once is a strong hint that this is the layer. **Application middleware.** Anything that wraps the response to rewrite HTML, inject tags, compute an ETag, or record the body for logging has to see the whole body first. These are easy to overlook because they live in your own codebase and look harmless. **The framework's own response handling.** Some server frameworks buffer by default and stream only through a specific response type; returning a plain string from a handler that also supports streams will silently buffer. ## Reading the headers Headers narrow things quickly. A `Content-Length` on a response you designed to stream means some layer computed the full body size, so some layer had the full body. Conversely `Transfer-Encoding: chunked` on HTTP/1.1 shows the response was still open-ended at that hop. Also check whether the connection is HTTP/1.1 or HTTP/2 at each hop, since the mechanism differs even though the observable behaviour you care about is the same. ## Ruling out the application Before blaming infrastructure, confirm the origin really streams. The frequent own-goal is a handler that awaits all its data and then writes: it holds a stream-shaped API but produces the whole document at once, so the first flush happens at the end. A quick check is to write the head, then a deliberate delay, then the body, and watch the timings from a shell on the origin host — no proxies, no CDN, no browser. ## Guarding it after the fix Streaming is a property that silently regresses, because a buffered page is still correct — just slower. Worth adding: a synthetic check that asserts the document's time to first byte is well below its total transfer time on a known-slow route, and a review habit that flags new response-wrapping middleware. Field data helps too — a rise in first-paint times with no change in the total load time is the fingerprint of streaming having quietly turned off somewhere on the path.

  • You confirm the origin streams but the edge does not. What are your options?
    Check whether the edge product has a response-buffering or streaming setting and enable streaming for those routes; some require an explicit opt-in. If it cannot stream, decide whether the route still belongs behind that edge — a cacheable static shell at the edge with the dynamic part fetched separately is often a better fit than fighting a buffering intermediary.
  • Why can compression silently disable streaming, and what is the tradeoff in fixing it?
    A compressor accumulates input to compress it well and, by default, emits output only when the stream ends. Flushing at each chunk boundary restores incremental delivery but produces slightly larger output, because each flush ends a block. The ratio loss is small and the perceived-speed gain is usually far larger, so streaming routes should flush per chunk.
  • How would you stop this regressing again after you fix it?
    Assert it. Add a synthetic check on a deliberately slow route that fails when the document's time to first byte is not well below its total transfer time, and watch for the field fingerprint — first paint degrading while total load time stays flat. Also treat any new middleware that reads or rewrites the response body as a review flag.

saying these in an interview costs you the question

  • Only tests locally, where there are no intermediaries
  • Assumes chunked encoding guarantees incremental delivery to the user
  • Never suspects compression as the buffering layer
  • Ignores a Content-Length header on a supposedly streamed response
  • Blames the browser for waiting on the full document

context