skip to content

A gRPC call through a new reverse proxy returns HTTP 200 but its status never arrives — what broke, and how do you confirm it?

level: seniorimportance: should knowfreq 52%

answer

  1. the verdict travels at the end
  2. a hop re-originates what it forwards
  3. fast rejections survive, late failures do not
  4. capture both sides of the hop
  5. a downgrade to HTTP/1.1 is the suspect

basics

~20 s

An in-path hop terminated the call and did not forward the trailer section, so grpc-status was erased. Confirm by capturing at both sides of the hop: the status is present leaving the server, absent arriving at the client.

solid answer

~40 s

The outcome of a gRPC call travels in trailing metadata, so any hop that terminates the connection and does not forward a trailer section deletes it. The client then sees a complete `:status 200` response that ends with no `grpc-status` at all, and has to synthesise an outcome — typically an unknown or internal failure with no useful message. The diagnostic tell is *which* failures survive: a call that failed on sight uses a **Trailers-Only** response, whose status sits in the response headers and therefore passes through, while a failure decided after the first message is delivered in a genuine trailer section and vanishes. Confirm it with a capture on both sides of the suspect hop, and check whether that hop downgrades to HTTP/1.1 anywhere, since trailing fields are far more commonly dropped there.

code

http · 16 lines
http
--- captured at the server's egress ---
HEADERS (stream 7, END_HEADERS)
:status = 200
content-type = application/grpc+proto

HEADERS (stream 7, END_STREAM, END_HEADERS)
grpc-status = 7
grpc-message = declaration%20locked%20by%20another%20filing

--- captured at the client's ingress, same call ---
HEADERS (stream 4, END_HEADERS)
:status = 200
content-type = application/grpc+proto

DATA (stream 4, END_STREAM)
<empty: stream ends, no trailer section>

go deeper

for a junior

The fact to carry away is that a gRPC response with no grpc-status is a broken call, not a quiet success. Anything between client and server can be the reason.

for a middle

Explain the mechanism: a hop that terminates the connection re-originates the response and may omit the trailing header block, and a downgrade to HTTP/1.1 makes that far more likely.

for a senior

Show the diagnosis, not the guess: capture both sides of the suspect hop for one call, and use the fast-rejection-versus-late-failure asymmetry to separate a trailer problem from a service problem.

for a principal

The durable lesson is that putting the verdict in trailing metadata makes every terminating hop a correctness dependency. Treat trailer forwarding as a standing acceptance criterion for anything admitted into the path.

## The symptom, stated precisely A broker's client reports that failures from the adjudicating service come back with no reason. The service's own logs show the calls failing normally with a status and a message. Between them sits a newly introduced reverse proxy. On the wire the client sees: - an initial `HEADERS` block with `:status = 200`, - possibly some `DATA`, - the stream ending — with **no `grpc-status` anywhere**. That is not a successful call and it is not a timeout. The arrival of the trailer section *is* the completion signal, so a response that ends without one is a broken call. The client library has no choice but to invent an outcome, and what it invents carries none of the server's diagnosis. ## The mechanism Any hop that terminates HTTP/2 rather than passing bytes through is re-originating the response, and it forwards exactly the fields it knows how to forward. Three variants produce this symptom: 1. **Trailing fields not forwarded.** The hop reconstructs the response from the initial header block and the body and simply omits the second header block. 2. **A protocol downgrade in the path.** Somewhere the call becomes HTTP/1.1, where trailing fields ride on a chunked body and are far more often dropped, then becomes HTTP/2 again. The status does not survive the round trip. 3. **Buffer-then-forward with a truncated tail.** The hop buffers the body, forwards it, and closes the stream without ever emitting what came after it. All three look identical from the client. This is precisely the failure `te: trailers` was invented to expose: a client announces trailer support on every request so a hop that cannot honour it has something to refuse rather than silently swallow. ## The tell that names the hop The distinguishing observation, and the one worth carrying into an interview: | how the call failed | where the status sits | survives a trailer-stripping hop? | |---|---|---| | rejected on sight (unknown method, bad credential) | Trailers-Only — in the response **headers** | usually yes | | failed after processing started | genuine **trailer section** | no | | succeeded | genuine trailer section | no — the success status is lost too | So the report "our fast rejections are legible and our real errors are blank" is close to a fingerprint. Note the third row: a trailer-stripping hop breaks *successful* calls as well, because success is also delivered as `grpc-status = 0`. If successes still work, the hop is probably dropping only the second header block on some path rather than all of them, which is worth pinning down before blaming it. ## Confirming it 1. **Capture on both sides of the suspect hop** for the same call. Server egress shows the trailing `HEADERS` with `grpc-status`; client ingress does not. That single comparison is conclusive and nothing else is needed. 2. **Compare a fast rejection with a late failure** through the same path. If one carries a status and the other does not, the difference is Trailers-Only versus a trailer section, which points at trailer handling rather than at the service. 3. **Establish where HTTP/2 is terminated** and whether any segment runs HTTP/1.1. Every termination point is a candidate. 4. **Bypass the hop** — call the service directly from inside the same network boundary. If statuses come back, the hop is confirmed without any further instrumentation. ## The related failure it is often confused with A hop can also fail the call *itself*, returning its own HTTP error — a 502 when it cannot reach the service, a 504 when it gives up waiting. That response has no `grpc-status` either, but the shape is different: the HTTP status is not 200. The specification defines a mapping the client applies in that case, so the caller still gets something actionable: - `400` → `INTERNAL` - `401` → `UNAUTHENTICATED` - `403` → `PERMISSION_DENIED` - `404` → `UNIMPLEMENTED` - `429`, `502`, `503`, `504` → `UNAVAILABLE` - everything else → `UNKNOWN` The two cases are told apart by the HTTP status alone: **200 with no trailer section** is a stripped status, and **a non-200** is the hop failing the call on its own account. They need completely different fixes, so establish which one you have before changing anything.

  • The hop returns its own 502 instead, with no grpc-status at all. What does the caller see?
    The client applies the specification's HTTP-status mapping: 502, like 429, 503 and 504, becomes `UNAVAILABLE`. 400 becomes `INTERNAL`, 401 `UNAUTHENTICATED`, 403 `PERMISSION_DENIED`, 404 `UNIMPLEMENTED`, and anything unmapped becomes `UNKNOWN`. The caller gets an actionable status even though no gRPC response ever existed.
  • Would successful calls also break through a trailer-stripping hop?
    Yes. Success is delivered the same way, as `grpc-status = 0` in the trailer section, so it is erased too and the call ends with no outcome. If successes work while failures do not, the hop is dropping trailers only on some path, and that difference is worth locating.
  • Why does te: trailers not simply prevent this?
    It is a declaration, not an enforcement. It gives a hop that knows it cannot forward a trailer section grounds to reject the request outright, which is better than silence — but a hop that drops the section without noticing sends nothing back to say so.

saying these in an interview costs you the question

  • Reads a 200 with no grpc-status as a successful call
  • Blames the server, whose own logs show a normal failure
  • Assumes only failing calls are affected, not successful ones
  • Confuses a stripped trailer section with the hop's own 502
  • Expects te: trailers to enforce forwarding rather than declare support