skip to content

Inside a gRPC call's DATA frames, what precedes each message, and why can't one frame be read as one message?

level: middleimportance: should knowfreq 48%

answer

  1. five bytes in front of every message
  2. one flag byte, four length bytes
  3. big-endian, and length excludes the prefix
  4. frames and messages are unrelated
  5. read by declared length, not by frame

basics

~10 s

Each message carries a five-byte prefix: one Compressed-Flag byte, then a four-byte big-endian Message-Length. HTTP/2 DATA frame boundaries are unrelated to message boundaries, so a receiver reads by that declared length, never by frame.

solid answer

~50 s

The body of a gRPC call is a stream of **Length-Prefixed-Message** records, not a blob. Each record is one byte of **Compressed-Flag** — `0` for an uncompressed message, `1` for one compressed with whatever `grpc-encoding` named — followed by a four-byte **Message-Length** as an unsigned big-endian integer, followed by exactly that many bytes of encoded message. A receiver reads five bytes, learns the length, reads that many bytes, and repeats. DATA frame boundaries are a separate concern entirely: a large message may span many frames and several small messages may share one, so a reader that treats a frame payload as a message breaks the moment either happens. This framing is also what makes a streaming call possible — the same stream carries an unbounded series of records with no delimiter and no terminator beyond the lengths themselves.

code

http · 10 lines
http
DATA (stream 5)
00 00 00 00 05 0a 03 41 42 43

Compressed-Flag = 00              (message is not compressed)
Message-Length  = 00 00 00 05     (5 bytes, big endian)
Message         = 0a 03 41 42 43  (field 1, a 3-byte string)

DATA (stream 5, END_STREAM)
00 00 00 01 20 ...               (the next message, 288 bytes,
                                  which may continue in later frames)

go deeper

for a junior

Know that the body is a series of length-prefixed records rather than one payload, and that five bytes — a flag and a length — sit in front of each message. That is enough to read a capture.

for a middle

Be able to state the receive loop and why a declared length is the only workable option when the body's total size is unknown and the transport reframes freely.

for a senior

Recognise the failure signature: a reader built on a frame-per-message assumption that passes locally and breaks against a peer that coalesces, and a truncated final record that must fail the call.

for a principal

The relevant lever at scale is the receive limit, not the wire ceiling. Where you set it decides whether large documents travel as one message or are chunked by the schema, and that choice propagates into every client.

## Two framings, stacked There are two independent framings in play on a gRPC call, and conflating them is the mistake this question exists to catch. - **HTTP/2 framing** chops the stream into `DATA` frames sized for the transport — bounded by the peer's maximum frame size and by how much the sender happens to flush at once. - **gRPC framing** chops the *byte stream those frames reconstruct* into messages, using a five-byte prefix on each one. Neither knows about the other. Concatenate every `DATA` payload on a stream, in order, and you get one continuous byte sequence; the prefixes tell you where the messages inside it start and stop. ## The five bytes | bytes | name | meaning | |---|---|---| | 1 | `Compressed-Flag` | `0` = message as-is; `1` = compressed with the method named by `grpc-encoding` | | 4 | `Message-Length` | unsigned, **big-endian**, the length of the message bytes that follow | Three details do real work: 1. **The flag is per message, not per call.** One stream can carry a compressed message and an uncompressed one back to back, which is how an implementation skips compressing a payload too small to benefit. 2. **The length covers the message only** — not the prefix itself. A five-byte message occupies ten bytes on the wire. 3. **Big-endian, four bytes**, so the ceiling is an unsigned 32-bit count. In practice implementations impose a far lower receive limit of their own, and a message over that limit fails the call rather than being truncated. ## Why the length is not optional A reader needs to know where a message ends, and none of the alternatives work here: - A **content length** for the whole body is meaningless: a streaming call's body length is not known when the response headers are written. - A **delimiter byte** would have to be escaped, since protobuf messages are arbitrary binary. - The **frame boundary** cannot serve, because the sender does not control it: flushing, coalescing and the peer's frame-size setting all move it. A declared length is the only mechanism that works for a message of any size, in a stream of unknown total length, over a transport that reframes freely. ## Reading the stream The receiving loop is small and worth being able to state: 1. Buffer incoming `DATA` payloads into one byte sequence. 2. Wait until at least **five** bytes are available; read the flag and the length. 3. Wait until **Message-Length** further bytes are available; that is one message. 4. Decompress it if the flag says so, hand it to the application, and go back to step 2. 5. When the direction closes with `END_STREAM`, a **partial** record — fewer bytes than the prefix promised — is a broken call, not an empty one. Step 5 is the operational tell. A truncated final record means the stream ended mid-message, which is a transport failure to be reported as such; it must never be delivered as a short message. ## What this looks like when it goes wrong - A reader written against a **frame-per-message** assumption works perfectly in a local test, where every small message happens to land in its own frame, and fails against a real peer that coalesces them. - A hop that **rewrites the body** — recompressing, or reframing without preserving byte order and completeness — corrupts the length stream, and the failure surfaces as a nonsense length several messages later rather than at the point of damage. - Enabling a **transport-level** compression over a body that is already per-message compressed gains nothing and costs CPU on every hop; the flag byte exists so compression is decided message by message inside the call. ## Where the prefix sits relative to everything else It is worth placing the five bytes precisely, because three layers are stacked here and each has its own notion of size: - The **header block** carries the call's definition fields and never carries a message. - The **`DATA` frames** carry bytes, sized by the sender's flushing and the peer's maximum frame size. - The **length-prefixed records** live inside those bytes and are the only layer that knows what a message is. - The **trailer section** arrives after the last record and carries the outcome, never a message. So "how large may a message be" and "how large may a frame be" are different questions with different answers, and neither is answered by the other. A receive limit rejects an oversized message; a frame-size setting only changes how many frames that message is spread across. For a filing gateway carrying one large declaration document per call, the practical reading is that message size is a *receive limit* question, not a frame question: the transport will happily spread one message across as many frames as it needs.

  • The stream ends with three bytes left in the buffer. What has happened?
    The call was truncated mid-record: a prefix promised bytes that never arrived. That is a transport failure, and the client must fail the call rather than deliver a short message or silently ignore the remainder.
  • What does the Compressed-Flag being per message let an implementation do?
    Decide compression message by message — typically skipping it for payloads too small to win, or for one already-compressed attachment inside an otherwise compressible stream — without renegotiating anything for the call.
  • Does the four-byte length mean a message can be four gigabytes?
    The field would allow it, but implementations apply a much lower configured receive limit and fail a call whose declared length exceeds it. Treat the wire ceiling as theoretical and the configured limit as the real one.

saying these in an interview costs you the question

  • Thinks one DATA frame always holds exactly one message
  • Believes the four-byte length includes the five-byte prefix
  • Reads the length as little-endian
  • Expects a delimiter byte or terminator between messages
  • Delivers a truncated final record as a short message