A streamed LLM reply arrives in 4-second bursts only in production — how do you debug it?
answer
- the model does not generate in bursts
- something in the path is accumulating
- bisect hop by hop with timestamps
- proxies and compression buffer by default
- total latency looks fine, timing does not
basics
~20 sSomething in the path is accumulating the body instead of forwarding each chunk. Bisect the chain — app, reverse proxy, CDN, client — by timestamping chunk arrivals at each hop, then disable buffering and compression on the streaming route.
solid answer
~50 sBursty delivery that only appears in production is almost never the model; it is a hop that was designed to buffer whole responses. The method is to bisect: hit the app process directly with an unbuffered client and timestamp each chunk, then repeat through each layer — reverse proxy, load balancer, CDN or WAF — until the timestamps clump. The usual culprits are a reverse proxy buffering responses by default (in nginx, `proxy_buffering on`, fixed with `proxy_buffering off` on that location or by having the app send `X-Accel-Buffering: no`), a compression layer holding output until a block fills, and a CDN or security appliance that inspects the full body. Also check your own side: a framework that does not flush after each write, and a client that awaits a complete parse before rendering, produce the same symptom.
code
nginx · 8 lineslocation /api/chat/stream {
proxy_pass http://app_upstream;
proxy_http_version 1.1;
proxy_buffering off;
proxy_cache off;
gzip off;
proxy_read_timeout 300s;
}go deeper
Know that a streaming response passes through proxies and other layers that may hold bytes, and that identical content can still arrive with very different timing.
Explain the common buffering points — reverse proxy, compression, CDN, framework flush behaviour — and name the concrete settings that turn buffering off on a streaming route.
Demonstrate the bisect method: timestamp chunk arrivals hop by hop until they clump, then fix that hop. Show that you would also instrument the client before blaming infrastructure.
Own the prevention: streaming routes as an explicit edge configuration class, and synthetic checks asserting on chunk timing through the real path, so a perceptual regression fails a check rather than reaching users.
## The symptom and what it rules out Smooth token-by-token output in development, four-second clumps in production, same model and same prompt. That difference isolates the problem to the delivery path, because the only thing that changed is what sits between the process and the browser. A model does not generate in bursts; a network path that accumulates and releases does. ## Why buffering is the default everywhere Almost every layer in a normal HTTP stack was built on the assumption that a response is a finite document. Buffering lets a proxy read the upstream response as fast as the upstream can send it, free the upstream connection, and then dribble the bytes out to a slow client. It lets a compression layer reach a useful block size before emitting. It lets a security appliance scan the whole body before deciding to pass it. All of those behaviours are correct for a web page and fatal for a stream. ## Bisect the path Do not guess; measure. Take one request and record the arrival time of every chunk at successively later points: 1. **The app process directly.** Bypass everything and hit the container or process port with a client that does not buffer — `curl -N` prints as it receives. If chunks arrive evenly here, the app is fine. 2. **Through the reverse proxy.** Same request, proxy port. If the even trickle becomes clumps here, the proxy is the culprit. 3. **Through the load balancer / CDN / WAF.** Repeat outward, one hop at a time. 4. **In the browser.** Log a timestamp per chunk in the client's read loop. If the network tab shows even arrival but the UI updates in bursts, the problem is rendering, not transport. The hop where the timestamps first clump is your answer. This is the whole technique; everything else is knowing what to change once you have found it. ## The usual culprits **Reverse-proxy response buffering.** In nginx, `proxy_buffering` is on by default; on a streaming location you set `proxy_buffering off`, and usually `proxy_cache off` and a generous `proxy_read_timeout` so a long reasoning phase does not look like a dead upstream. An application that cannot change nginx config can send the `X-Accel-Buffering: no` response header, which nginx honours per response. Other proxies have equivalent settings. **Compression.** A gzip or brotli layer that waits for enough bytes to fill a compression block turns a token trickle into periodic emissions. Disable compression on the streaming route, or ensure the layer flushes per chunk. This is a classic cause of "it bursts every few seconds". **CDNs, WAFs and API gateways.** Many inspect or cache full bodies by default. Streaming routes usually need an explicit pass-through or streaming mode, and the same applies to some serverless front doors, which may only deliver a response once the handler returns. **Your own app.** Frameworks differ in whether writing to the response actually flushes. If an internal buffer or a wrapping middleware collects output, you get bursts before any proxy is involved. Check that the streaming handler flushes after each chunk and that no logging, tracing or transformation middleware materialises the body. **The client.** A reader that accumulates until the response completes, a parser that only emits on a well-formed boundary, or a UI that batches state updates on a timer can all convert an even stream into visible bursts. This is why you instrument the browser as well as the wire. ## Preventing the regression The reason this is a production-only bug is that development environments have no proxy chain. Two habits keep it from returning. First, treat streaming routes as a distinct class in your edge configuration, with buffering and compression explicitly disabled and long read timeouts, rather than relying on the default profile. Second, add a synthetic check that exercises the streaming endpoint through the real edge and asserts on chunk timing — for example, that the gap between first and second chunk stays under a threshold. That converts an invisible perceptual regression into a failing check, which matters because nothing about a buffered stream shows up as an error: the response is complete and correct, only the timing is destroyed. Total latency dashboards will look perfectly healthy while the feature feels broken.
- The wire shows chunks arriving evenly but the UI still updates in bursts. What now?That moves the problem to the client. Look for a reader that accumulates before emitting, a parser that only yields on a complete boundary, or a UI framework batching state updates on a timer or animation frame. Log a timestamp in the read loop and another at paint time; whichever gap grows tells you whether it is the consumer or the renderer.
- Why does compression cause this even when the proxy itself is not buffering?A compressor emits output in blocks. If it waits to accumulate enough input to fill a block before flushing, the bytes physically cannot leave until then, regardless of the proxy's buffering setting. The fixes are to disable compression on the streaming route or to use a mode that flushes per chunk, accepting a worse compression ratio.
- Why does this bug rarely show up in local development?Because local development usually talks straight to the app process, with no reverse proxy, CDN, WAF or compression layer in between — exactly the components that buffer. The response is also byte-identical either way, so no test that asserts on content will fail. Only timing-aware checks through the real edge catch it.
saying these in an interview costs you the question
- Blames the model or the provider for bursty output
- Changes proxy settings by guesswork instead of bisecting the path
- Forgets that compression layers buffer independently of the proxy
- Assumes buffering shows up in total-latency dashboards
- Never checks whether the client is accumulating before rendering