skip to content

In HTTP/1.1, how does the receiver of a message know where the body ends, and what exactly does the Content-Length header field count?

level: juniorimportance: must knowfreq 62%

answer

  1. Octets, not characters
  2. Counted after gzip
  3. Blank line, then N bytes
  4. No field: request body = 0
  5. HTTP/2: END_STREAM, not chunks

basics

~20 s

Content-Length states the body size in octets, so the receiver reads exactly that many bytes and stops. If it is absent, HTTP/1.1 can use Transfer-Encoding: chunked, where each chunk announces its own size and a zero-size chunk ends the body.

solid answer

~50 s

HTTP/1.1 reuses one connection for many messages, so every message needs an explicit end marker. There are two framings. **Content-Length: N** says the body is exactly N octets, starting right after the empty line that ends the header block. The receiver counts N bytes, then expects the next message on that connection. **Transfer-Encoding: chunked** streams the body as self-describing chunks, each prefixed by its size in hex and terminated by a zero-size chunk, so the sender never needs to know the total length up front. Content-Length counts octets of the body *as transmitted* (after any content coding such as gzip), not characters and not headers. If it disagrees with reality the connection desynchronises: too small, and leftover bytes are parsed as the start of the next message; too large, and the reader blocks waiting for bytes that never arrive. With neither field, a request has no body; a response body runs until the connection closes.

code

http · 6 lines
http
POST /orders HTTP/1.1
Host: api.example.com
Content-Type: application/json
Content-Length: 27

{"sku":"A-1","quantity":3}

go deeper

for a junior

State the two framings and that Content-Length is a byte count of the body starting after the blank line. Mention that a wrong value breaks the connection.

for a middle

Add the encoded-bytes rule (after gzip), the zero-length default for requests, and close-delimited responses as the fallback.

for a senior

Frame it as connection-state correctness: any mismatch desynchronises a reused connection, which is why proxies must recompute or strip the field when they alter a body.

for a principal

Position framing as the layer everything else on the connection depends on, and note that HTTP/2 and HTTP/3 moved it into the binary layer to remove a whole class of parser disagreements.

## Why framing exists HTTP/1.1 runs over a TCP or TLS byte stream, and that connection is normally persistent: many requests and responses travel over it back to back. TCP delivers an ordered stream with no record boundaries, so the protocol itself must declare where each message ends. The header block is self-delimiting because it ends at the first empty line, but the body is opaque bytes, so its length has to be stated separately. That statement is the message framing. ## Content-Length framing `Content-Length: 348` means: the body is exactly 348 octets (bytes), beginning immediately after the blank line that terminates the headers. Three details matter. 1. **Octets, not characters.** A UTF-8 body of 200 characters may be 260 octets. Computing Content-Length from a string's character count is a classic bug. 2. **After content coding.** If you send `Content-Encoding: gzip`, Content-Length is the size of the gzipped bytes on the wire, not the original. The order is: build the representation, apply content coding, count what you are about to write. 3. **It is a promise the connection depends on.** If the declared value is smaller than the bytes written, the extra bytes are read as the beginning of the next message on that connection, corrupting it. If it is larger, the receiver waits for bytes that never come and eventually times out. Frameworks compute it for you precisely because getting it wrong is destructive. Content-Length can also legitimately appear with no body: a `HEAD` response advertises the size the equivalent `GET` would return, and a `304 Not Modified` may echo it. In both cases no body is actually sent, and the framing rules say so explicitly, which is why a parser must special-case them rather than trusting the number. ## Chunked framing When the sender does not know the total size in advance (a streamed report, a proxied response, a slow database export), it sends `Transfer-Encoding: chunked` and omits Content-Length. The body then becomes a sequence of chunks, each starting with its own size in hexadecimal on its own line, followed by that many octets. A chunk of size zero marks the end. The receiver therefore learns lengths incrementally and still knows exactly where the message stops. ## The default when neither is present The rules are asymmetric: - **Request with neither field:** the body length is zero. A server must not wait for a body just because the method usually has one. This is why an HTTP client that forgets Content-Length on a POST often sees the server treat it as an empty body. - **Response with neither field:** the body runs until the server closes the connection ("close-delimited"). This works, but it is HTTP/1.0-style: the connection cannot be reused, and the client cannot distinguish a complete body from a truncated one caused by a network failure. It is a last resort, not a design choice. ## Newer versions HTTP/2 and HTTP/3 do not use either mechanism for framing. They carry the body in DATA frames, and the end of the body is signalled by an END_STREAM flag on the stream. Content-Length may still be sent as metadata, and endpoints are required to reject a message where the declared length does not match the DATA actually delivered, but the framing itself is done by the binary layer. Chunked transfer coding is forbidden in HTTP/2 and HTTP/3 entirely. ## What interviewers are checking That you understand HTTP messages are delimited explicitly rather than by connection close, that Content-Length is a byte count of the encoded body, and that a wrong value is not a cosmetic error but a connection-level correctness bug.

  • If the server sends Content-Length: 100 but writes 120 bytes, what does the client see?
    The client reads the first 100 bytes as the body and treats byte 101 onward as the start of the next response on that persistent connection. Parsing then fails, or worse, succeeds against attacker-chosen bytes. Well-behaved clients drop the connection; the practical symptom is intermittent malformed responses under keep-alive that disappear when connection reuse is turned off.
  • Does Content-Length describe the compressed or uncompressed size when Content-Encoding: gzip is used?
    The compressed size, because Content-Length frames the octets actually transmitted. The uncompressed size is not carried in any standard field. This is also why a proxy that decompresses a body must recompute or remove Content-Length before forwarding.

saying these in an interview costs you the question

  • Saying Content-Length counts characters rather than octets
  • Claiming it includes the header block
  • Saying the body ends when the connection closes, as though that were the normal case in HTTP/1.1
  • Thinking Content-Length must be recomputed to the pre-gzip size
  • Assuming a POST without Content-Length still delivers a body

context