skip to content

Walk through the wire format of an HTTP/1.1 message that uses Transfer-Encoding: chunked: how a chunk is written, how the body is terminated, and what trailer fields are for.

level: middleimportance: must knowfreq 52%

answer

  1. Hex size, CRLF, data, CRLF
  2. Zero chunk = end
  3. Trailers after the zero chunk
  4. TE: trailers to opt in
  5. Chunk boundaries are not record boundaries

basics

~20 s

Each chunk is its size in hexadecimal, CRLF, that many octets, CRLF. A chunk of size 0 ends the body. After the terminal chunk, optional trailer fields may appear (metadata such as a checksum computed only after streaming), then a final empty line.

solid answer

~50 s

The sender sets `Transfer-Encoding: chunked` and omits Content-Length. The body then reads as: `<size-in-hex>CRLF<size octets>CRLF`, repeated, and finished by a chunk whose size is `0`. After that zero chunk, zero or more **trailer fields** may follow in header syntax, then a final CRLF closes the message. Sizes are hexadecimal, not decimal, and may carry chunk extensions after a semicolon, which almost nobody uses and many parsers reject. Chunked exists so a sender can start writing before it knows the total length: proxied or streamed responses, generated exports, server-sent events. Trailers exist for metadata that can only be computed after the body has been produced, such as a content digest or a gRPC-style status. Trailers are widely under-supported: a client only receives them if it advertises `TE: trailers`, and many proxies and client libraries silently discard them, so do not put anything load-bearing there without controlling both ends.

code

http · 11 lines
http
HTTP/1.1 200 OK
Content-Type: text/plain
Transfer-Encoding: chunked
Trailer: Digest

1a
first line of the export..
9
 and more
0
Digest: sha-256=:qq2ZQ0m8Ph1sJ2Nl:

go deeper

for a junior

Be able to read a chunked body: hex size, data, repeat, zero chunk ends it. Knowing trailers exist is enough.

for a middle

Produce the exact syntax from memory, explain why chunked beats buffering for streams, and state that trailers follow the terminal chunk.

for a senior

Discuss re-framing by intermediaries, why chunk boundaries carry no application meaning, and the practical unreliability of trailers across proxies.

for a principal

Compare the HTTP/1.1 coding with HTTP/2 and HTTP/3 DATA frames plus trailing HEADERS, and judge whether a streaming contract should depend on trailers at all given the deployed proxy population.

## The shape of a chunked body With `Transfer-Encoding: chunked` there is no Content-Length. The body is a sequence of chunks, each written as: ``` <chunk-size in hex>[;extension]CRLF <exactly chunk-size octets>CRLF ``` and the body finishes with a **terminal chunk** whose size is `0`. After the terminal chunk come zero or more trailer fields, and then one empty line that closes the message. Three details trip people up: - **The size is hexadecimal.** `1a` is 26 octets, not 1a-as-decimal. Hand-written chunked writers frequently emit decimal and produce garbage. - **The CRLF after the data does not count** toward the chunk size. The size covers the payload octets only. - **Chunk boundaries are not message semantics.** Chunks are a transport detail; they may be re-split or merged by any intermediary. You cannot use one chunk to mean one JSON record, which is why streaming formats define their own delimiters (newline-delimited JSON, SSE `data:` lines) on top. ## Why it exists Content-Length requires knowing the total size before writing the first byte. That forces buffering the whole response in memory or on disk. Chunked removes that constraint, which matters for: - **Streamed generation:** a report, an export, a log tail, a model response produced token by token. - **Proxying:** an intermediary can forward bytes as they arrive rather than buffering a whole upstream response. - **Long-lived responses:** server-sent events keep a chunked response open indefinitely. The cost is a small per-chunk overhead and the loss of a progress indicator on the receiving side, since the total size is unknown until the end. ## Trailer fields A trailer field is a header-syntax field sent *after* the body instead of before it. It exists for metadata you cannot know until the body is complete. Canonical uses: an integrity digest over the streamed bytes, a per-message status computed after generation (gRPC over HTTP/2 uses trailers exactly this way for `grpc-status`), or timing metadata. Rules that matter in practice: - A sender may list expected trailer names in the `Trailer` header field so the receiver knows what is coming. - Trailers must not carry fields that affect framing, routing, or how the message is processed before the body is read: no Content-Length, no Transfer-Encoding, no Host, no Cache-Control, no Authorization. The receiver has already acted on the message by the time trailers arrive. - A client signals willingness to receive them with the `TE: trailers` request header field. Without that signal, an intermediary is free to drop them, and many do. - A recipient that forwards a chunked message but does not understand trailers may discard them silently, so treat them as best-effort unless you control the whole path. ## Interaction with intermediaries Transfer-Encoding is a hop-by-hop property of one connection. A proxy may receive a chunked response and forward it with Content-Length after buffering, or receive a Content-Length body and re-emit it chunked. This is legal and common, and it is one reason chunk boundaries cannot be trusted end to end. It is also why a proxy must never blindly copy the original framing fields when it changes the framing. ## Versions Chunked transfer coding is an HTTP/1.1 mechanism only. HTTP/2 and HTTP/3 forbid the `Transfer-Encoding` field; streaming is inherent because a body is a series of DATA frames terminated by END_STREAM, and trailers are carried as a second HEADERS frame after the data. Conceptually the model survives, but the syntax on the wire is entirely different. ## What good answers include The exact chunk syntax, that the size is hex, that a zero chunk terminates, that trailers follow the terminal chunk and are best-effort, and at least one real reason to prefer chunked over buffering to compute a length.

  • Why can you not treat each chunk as one logical record in a streaming API?
    Chunk boundaries are a transport detail, and any intermediary may merge or re-split them, or re-frame the message entirely with Content-Length. A receiver may also read a partial chunk from the socket. Streaming formats therefore carry their own delimiters, such as newline-delimited JSON or SSE event blocks.
  • What kinds of header fields must never be sent as trailer fields?
    Anything the receiver needed before reading the body: framing fields such as Content-Length and Transfer-Encoding, routing fields such as Host, request modifiers such as Authorization or Cache-Control, and anything affecting how the message is dispatched. By the time trailers arrive the message has already been routed and largely processed, so such fields would be either ignored or a security hazard.

saying these in an interview costs you the question

  • Writing chunk sizes in decimal instead of hexadecimal
  • Counting the trailing CRLF as part of the chunk size
  • Believing one chunk equals one application-level record
  • Sending Content-Length alongside Transfer-Encoding: chunked
  • Relying on trailers reaching the client without TE: trailers and control of the proxy path

context