skip to content

A team adds several large gRPC metadata keys to every call, and some calls now fail before the handler runs. Why?

level: seniorimportance: nice to knowfreq 30%

answer

  1. the envelope has a size limit
  2. two limits, one field block
  3. counted uncompressed, so compression cannot help
  4. base64 adds about a third
  5. rejected before the handler is entered

basics

~20 s

Metadata rides in a size-bounded field block. The gRPC specification suggests a default limit of 8 KiB on request metadata and trailers, and the HTTP/2 peer advertises its own limit, so an oversized block is rejected by the transport before any handler sees the call.

solid answer

~40 s

Custom metadata is not free space. It travels in the call's request field block, and that block is bounded twice over: the gRPC specification suggests a default limit of **8 KiB** on request metadata and trailers, and an HTTP/2 peer advertises the limit it will accept with **`SETTINGS_MAX_HEADER_LIST_SIZE`**. Cross either and the call is rejected at the transport, before the handler is entered — so the failure names no application cause and points at nothing in the business logic. Two details make it arrive sooner than people expect: the limit is computed on **uncompressed** name and value sizes plus a per-field overhead, so compression buys no headroom, and a `-bin` value is base64 on the wire, about **a third larger** than the bytes you handed over.

code

http · 8 lines
http
:path /pharmacy.v1.Eligibility/CheckPrescription
te: trailers
content-type: application/grpc+proto
grpc-timeout: 800m
counter-id: STORE-4417
dispensing-pharmacist: 4417-KL
request-trace-bin: AAAAAAAAAAAAAAAAAAAAAQ==
formulary-snapshot-bin: <about 8 KiB of base64 from a 6 KiB value>

go deeper

for a junior

Know that gRPC metadata is for small facts such as identifiers, and that putting large data there is not simply a slower version of putting it in the request message.

for a middle

Explain the two limits over one field block, that the count is taken on uncompressed sizes, and that a -bin value is base64 on the wire and therefore about a third larger than its raw bytes.

for a senior

Recognise the failure signature: data-correlated failures that never reach a handler, no application status, nothing in the service's logs, and possibly a new limit advertised after a deployment on the receiving side.

for a principal

Metadata conventions spread across teams faster than anyone budgets for. Decide who owns the total per-call budget and how a new organisation-wide key is approved before every service pays for it.

## Metadata is a budget, not a bag It is easy to treat gRPC metadata as somewhere convenient to put things — it needs no schema change, no version negotiation, no conversation with the team that owns the service. That convenience is exactly why this failure shows up: the block it lands in is bounded, and nothing in the calling code says so. Every call at the pharmacy counter carries a counter identifier, a dispensing identifier and a correlation key. Add one more key holding a snapshot of formulary state 'because it was easier than changing the message', and a few kilobytes later some calls stop working. ## Two limits over one block 1. **The gRPC specification's suggested default.** It suggests a default limit of **8 KiB** on the request metadata block and on trailers. It is a suggestion to implementers, not a constant you can rely on being identical everywhere — which is itself worth knowing, because the same oversized call may pass in one environment and fail in another. 2. **The HTTP/2 peer's advertised limit.** A peer states the largest field block it is prepared to accept using **`SETTINGS_MAX_HEADER_LIST_SIZE`**. The setting is advisory in the sense that it tells the other side what to expect rather than physically preventing a send, and a sender that ignores it should expect the call to be rejected. The number counted is **not** the number of bytes on the wire. It is the sum, over every field, of the **uncompressed** name length plus the value length plus a small fixed per-field overhead. ## Why compression does not rescue you HTTP/2 does compress field blocks, and this is the point where reasoning goes wrong. The limit is applied to the **uncompressed** sizes, so a block that compresses beautifully is measured as though it had not been compressed at all. You cannot buy headroom by making the values more repetitive, and you cannot buy it by shortening only the key names when the payload is in the values. ## Base64 is a third of the problem A metadata key ending in `-bin` carries its value as base64, which encodes every three bytes as four characters. A 6 KiB binary value therefore occupies roughly **8 KiB** in the field block — about a third more than the data you started with. Anyone sizing metadata against the raw bytes will be surprised by where the limit lands. ## What the failure looks like, and why it is hard to read The crucial property is **where** the rejection happens. The block is refused by the transport while the call is being accepted, which means: - the handler is **never entered**, so nothing in the service's own logging mentions the call; - there is no application-chosen status code describing the real cause, because no application code chose one — commonly the caller sees `RESOURCE_EXHAUSTED (8)`, and in some stacks a transport-level stream reset instead; - the failure is **correlated with data**, not with load: calls whose metadata happens to be longer fail while shorter ones succeed, which looks like flakiness until someone measures the block size; - it may appear only **after a deployment on the receiving side**, if that side's advertised limit is lower than the sender assumed. ## Staying inside the budget - Keep metadata to **small, fixed-shape facts**: identifiers, correlation keys, a locale. Metadata is the envelope. - Put anything that grows with the data in the **request message**, where the schema describes it, versioning covers it and the size limits are the message's rather than the field block's. - Size `-bin` values against their **base64** length, not their raw length. - Remember trailers share the same discipline: a verbose `grpc-message` or a large details payload competes for the same bounded block on the way back. - Treat the suggested 8 KiB as a **planning number**, and confirm what the receiving side actually advertises rather than assuming the two agree.

  • Why does the failure from an oversized gRPC metadata block carry no useful application status?
    Because the block is rejected while the call is being accepted, before the handler runs. No application code chose a status, so the caller usually sees `RESOURCE_EXHAUSTED (8)` or a transport-level reset, and the service's own logs mention nothing at all.
  • If a gRPC metadata block compresses to a fraction of its size, does that help it fit?
    No. The limit is computed on uncompressed name and value lengths plus a per-field overhead, so a highly compressible block is measured as if it were not compressed. Headroom comes from sending less, not from sending more repetitive data.

saying these in an interview costs you the question

  • Thinks metadata size is limited only by available memory
  • Believes header compression makes the size limit moot
  • Expects a normal application error when the block is too large
  • Forgets base64 grows a binary value by about a third
  • Moves bulk into metadata to avoid changing the schema