skip to content

A Server-Sent Events feed streams smoothly in local tests but arrives in clumps through a reverse proxy — why?

level: middleimportance: must knowfreq 68%

answer

  1. the origin is not the suspect
  2. something between the two ends
  3. released when a buffer fills
  4. clumps track bytes, not seconds
  5. route configuration, not response headers

basics

~20 s

An intermediary is accumulating the response body and releasing it only when its buffer fills. HTTP defines what a message means, not how promptly a proxy must forward part of one, so a response written event by event arrives in blocks.

solid answer

~40 s

The origin is fine; something between the two ends is holding the body. A reverse proxy, an edge tier or an L7 load balancer commonly reads an upstream response into a buffer and forwards when that buffer fills or when the response completes, which is a sensible optimisation for a request/response API and fatal for a response that never ends. The giveaway is that the clumps track a **byte size** rather than an interval: small events arrive in larger groups, large events in smaller ones. Origin-side batching, by contrast, clumps on a clock. The fixes are all downstream of your code -- turn off response buffering for that route, send the vendor-defined per-response opt-out field that some proxies honour, or keep the stream off the buffering path entirely.

code

http · 14 lines
http
GET /lines/3/events HTTP/1.1
Host: plant.example
Accept: text/event-stream

HTTP/1.1 200 OK
Content-Type: text/event-stream
Cache-Control: no-cache, no-transform
Transfer-Encoding: chunked

event: speed
data: 412

event: jam
data: cleared

go deeper

for a junior

Recall that a working stream can be broken by something in the middle, and that the first move is to request the same path from the origin directly and compare.

for a middle

Explain the mechanism: the intermediary accumulates the body and releases it on a size boundary, because nothing obliges it to forward part of a response promptly.

for a senior

Demonstrate the diagnosis - byte-quantised clumps versus time-quantised ones - and name levers you do not own, since the fix is usually configuration in an edge tier another team runs.

for a principal

The tradeoff is whether a live feed should traverse a general-purpose edge tier at all, weighing one more special-cased route against a separate path with its own operational surface.

## The symptom, stated precisely A bottling line publishes jam-and-speed events from the filler and capper as `text/event-stream`. Against a laptop on the plant floor, talking to the origin directly, each event lands the moment the line reports it. Published through the site's edge tier, the same feed goes quiet and then delivers twenty events at once, each carrying a timestamp showing it was produced seconds apart from its neighbours. Nothing is lost. Everything arrives, in order, correctly framed. Only the timing is destroyed -- which is the entire product. This is the failure that never appears in a test and always appears after deploy, because a test has no intermediary in the path. ## Why an intermediary is allowed to do this HTTP tells an intermediary what a message means. It does not oblige it to pass along a piece of one promptly. A proxy that reads an upstream response into a buffer and writes downstream when the buffer fills is doing something entirely reasonable for the traffic it was designed for: fewer and larger writes toward the client, upstream connections released sooner, and a slow reader insulated from the origin. For a response that ends, the only cost is a few milliseconds. For a response that never ends, the buffer becomes the delivery schedule. This is a documented hazard rather than a surprise: RFC 6202, which collects the known issues in long polling and HTTP streaming, records intermediary buffering as one of the central obstacles to streaming over HTTP. ## Bytes, not seconds -- the diagnostic that names the cause One observation separates the two candidate causes: - Clumps that track a **byte size** point downstream. A fixed buffer releases on a size boundary, so a feed of small events clumps in larger counts and a feed of large events clumps in smaller ones, while the elapsed gap between clumps varies with how fast the line is running. - Clumps that track an **interval** regardless of event size point upstream, into the producing code or a scheduler that gathers events before writing them. Compare arrival timestamps at the receiver against the write timestamps at the origin and the two separate in a single pass. ## The levers, in the order worth trying 1. **Request the same path from the origin directly**, bypassing the edge tier. This confirms in one command that the origin writes incrementally and moves the investigation downstream. 2. **Disable response buffering for that route** in the intermediary. This is the real fix, and it is configuration in someone else's tier rather than code in yours -- which is why this failure so often crosses a team boundary. 3. **Send the per-response opt-out field.** Many proxies honour a vendor-defined response header field that turns buffering off for one response. No specification defines such a field, the spelling differs by product, and an intermediary that does not recognise it ignores it silently. Useful, not dependable. 4. **Keep the stream off the buffering path**, on a route or hostname that does not traverse the tier that buffers. 5. **End-to-end TLS through transparent intermediaries.** Something that cannot read the body cannot usefully accumulate or re-frame it, and this was the historical recommendation for HTTP streaming. It does nothing about a reverse proxy in your own edge tier that terminates TLS by design. ## What does not fix it | Attempted fix | Why it does not work | |---|---| | Adding cache directives | `no-cache` and `no-store` govern reuse and storage, not when received bytes are forwarded | | Padding events to fill the buffer | Guesses at a size you do not control, and breaks when the tier is reconfigured | | Switching the response to HTTP/2 or HTTP/3 | An intermediary that buffers a body buffers it under any version | | More flushing at the origin | The origin's writes already left; they are sitting in someone else's buffer | That last row is the one candidates most often get backwards. The emitting side's flush discipline is a real duty and a real bug when it is missing, but it is a **different** bug: when the origin is at fault, a direct request to the origin clumps too. When an intermediary is at fault, the direct request is smooth and only the published path clumps. ## Why the origin's own monitoring is silent There is no error to log. The origin wrote the bytes, the write returned, the response is still open and the client has not disconnected. Every counter the origin owns says the stream is healthy. The stall lives entirely in the gap between two hops, and the only place it is visible is at the receiver's clock -- which is why this failure is usually first reported by a user saying the dashboard "feels laggy" rather than by an alert.

  • Does end-to-end TLS fix a buffering intermediary?
    Only a transparent one. Something that cannot read the body cannot usefully accumulate or re-frame it, and that was the historical recommendation for streaming over HTTP. A reverse proxy in your own edge tier terminates TLS by design, reads the body, and buffers it exactly as before, so TLS changes nothing on the hop that usually matters.
  • Why does the origin's own logging show nothing wrong?
    Because nothing is wrong there. The write succeeded, the flush succeeded, the response is still open and the client is still connected, so no counter the origin owns can register the stall. It is only visible as a gap between the origin's write timestamps and the receiver's arrival timestamps.
  • One route streams fine and another clumps through the same edge tier. What differs?
    Response buffering is usually configured per route or per upstream, and some intermediaries buffer only above a size threshold or only when a content coding is applied. Nothing about `text/event-stream` makes an intermediary special-case it, so a media type that streams on one path can clump on another purely because of that path's configuration.

A buffering intermediary is a mail room that will not despatch a sack until the sack is full. Every letter you wrote arrives, in order and intact, on the schedule of the sack rather than the schedule of your writing.

saying these in an interview costs you the question

  • Blames the origin's flushing when a direct request streams fine
  • Says a cache directive will stop an intermediary buffering
  • Assumes HTTP obliges a proxy to forward a partial response promptly
  • Expects the origin's logs or metrics to show the stall
  • Pads events to a guessed buffer size and calls it fixed
  • Thinks switching HTTP version defeats a buffering intermediary