skip to content

In a Server-Sent Events response, why must a receiver never treat a `chunked` transfer chunk boundary as an event boundary?

level: seniorimportance: nice to knowfreq 26%

answer

  1. two framings, different owners
  2. the coding applies to one hop
  3. any hop may re-draw the pieces
  4. local tests map writes to chunks by accident
  5. dispatch on the media type's boundary

basics

~20 s

A transfer coding is hop-by-hop framing with no application meaning, and any intermediary may re-chunk the body as it forwards it. One chunk can hold two and a half events, or half of one, so chunk edges say nothing about where an event ends.

solid answer

~40 s

Two independent framings sit on the same bytes. The lower one is the transfer coding that carries the body between two hops; the upper one is the event stream's own line grammar. They are chosen by different parties and they do not line up. Every hop is free to re-chunk: combining several small writes into one chunk, or splitting one write across two, and it does so for its own reasons -- buffer sizes, write batching, a change of HTTP version. A receiver that dispatches one event per chunk works perfectly against the origin in a test and fails the first time an intermediary re-frames the body, which is a failure that appears only after deploy and looks like data corruption rather than a parsing bug.

code

http · 12 lines
http
HTTP/1.1 200 OK
Content-Type: text/event-stream
Transfer-Encoding: chunked

14
data: speed=412

dat
e
a: speed=418

0

go deeper

for a junior

Recall that a response body is a stream of bytes and that whatever framing carries it between hops has nothing to do with where an application message ends.

for a middle

Explain that a transfer coding applies to a single hop, so an intermediary re-chunks freely, combining or splitting whatever the previous hop sent.

for a senior

Recognise the symptom - truncated and glued-together payloads that appear only once an intermediary is in the path - and locate it in the receiver's parsing rather than in the producer.

for a principal

The angle is build-versus-adopt on the receiving side: a hand-rolled reader carries this defect until an edge tier changes its write batching, which is an argument for a conforming client over a bespoke one.

## Two framings on the same bytes A bottling-line feed travelling over HTTP/1.1 carries two independent notions of "where something ends": - **The transfer coding.** `Transfer-Encoding: chunked` carries a body of unknown length between two hops by splitting it into sized pieces. The sizes are chosen by whoever is doing the sending on that hop. - **The event stream's own grammar.** Where one event ends inside the body is decided by the media type's line rules, which live in the body's own text and are invisible to anything framing it. Neither party knows about the other. The hop doing the chunking is framing bytes, not events, and has no reason to align its pieces with anything in the payload. ## Who may re-chunk, and why A transfer coding applies to **one hop**. When an intermediary forwards a body it re-applies the coding on its own terms, and it may: - combine several small upstream chunks into one larger downstream chunk, because fewer, larger writes are cheaper; - split one upstream chunk across several downstream ones, because its own buffer filled; - change the coding entirely when the downstream connection uses a different HTTP version. None of that requires permission and none of it is unusual. It is the ordinary behaviour of a tier whose job is to move bytes. ## What the failure looks like A receiving component that reads a chunk and treats it as one event will, against the origin directly, appear flawless: a naive origin that writes one event per socket write produces exactly one chunk per event, so every test passes. Published through an edge tier, the same component starts producing events with truncated payloads and events that contain two payloads glued together. The reports come back as data corruption, and the investigation goes to the producing code, the serialiser, the database -- everywhere except the framing assumption, which is invisible because it was never written down. The correct shape is boring: treat the body as a byte stream, carry an incomplete tail across reads, and dispatch on the boundary the media type defines rather than on the boundary the transport happened to hand you. ## The same rule under a different name HTTP/1.1 is where the word `chunked` appears, but the rule is not about that coding: | Version | How the body is framed underneath | Do those boundaries mean anything to the application? | |---|---|---| | HTTP/1.1 with a transfer coding | Sized chunks chosen by each hop | No | | HTTP/2 and HTTP/3 | `DATA` frames chosen by the sender and by intermediaries | No | There is no version of HTTP in which the transport's framing tells you where an application message ends. Moving a stream from one version to another does not change the discipline; it only changes which boundary a naive receiver will latch onto. ## Why this belongs to the intermediary story This is the quiet sibling of buffering. Both come from the same fact -- an intermediary handles the body on its own terms -- but they present differently. Buffering delays events without damaging them, and is caught by looking at timing. Re-chunking damages the parse without delaying anything, and is caught only by looking at the bytes. A team that has fixed the buffering problem often still has this one latent in a hand-rolled receiver, waiting for the day the edge tier's write batching changes. ## What an interviewer is listening for Three things: 1. That the candidate knows the transfer coding is per hop, not end to end. 2. That they can say why local testing hides it: a naive origin produces a one-to-one mapping between writes and chunks purely by accident. 3. That the fix is a parsing discipline rather than a configuration change, which makes it the one failure in this area you can actually fix inside your own code. A candidate who says "we just never split an event across a write" has missed the point in the most instructive way: they are trying to control a boundary that another party gets to redraw after the bytes leave them.

  • Why does this defect pass every local test?
    Because an origin that writes one event per socket write, talking straight to the receiver, produces one chunk per event by accident. The one-to-one mapping the receiver depends on is real in that setup and is a coincidence of the topology, not a property of the protocol. It survives exactly until a hop with its own buffer sits between the two ends.
  • Does moving the feed to HTTP/2 remove the problem?
    No, it renames it. There is no transfer coding in HTTP/2, but the body arrives in `DATA` frames whose sizes are chosen by the sender and by any intermediary on the path, and those boundaries carry no application meaning either. A receiver still has to buffer an incomplete tail across reads and dispatch on the boundary the media type defines.

saying these in an interview costs you the question

  • Treats one transfer chunk as one application message
  • Thinks the sending side controls the boundaries the receiver sees
  • Assumes a transfer coding is end to end rather than per hop
  • Believes moving to HTTP/2 removes the need to buffer partial reads
  • Concludes truncated payloads mean the producer serialised them wrongly