skip to content

Curators debug a gRPC catalogue fleet through a reverse proxy mid-rollout and reflection hands them two definitions of one message. What happened?

level: seniorimportance: nice to knowfreq 30%

answer

  1. the answer describes one instance
  2. descriptors are not fleet-wide facts
  3. a new call may reach a new build
  4. one stream, one server, one answer set
  5. pin the session, do not cache across it

basics

~20 s

A reflection answer describes the instance that produced it. During a rollout the pool holds two builds with different descriptors, and separate calls through a reverse proxy can reach different instances, so two lookups return two versions of the same message.

solid answer

~50 s

Descriptors are not a fleet-wide fact. Each server answers reflection from the schema compiled into **that build**, so a pool that is halfway through a rollout genuinely holds two different answers for the same message. Within one reflection call the answers are consistent, because one call is one stream on one connection and therefore one server — but a second call, or a reconnect, can be routed by the reverse proxy to an instance running the other build. The damage is downstream: a caller that caches a descriptor from build A and then sends a request to build B is composing a message against a contract the receiving instance does not share. The fix is affinity, not retries — address a single instance for the session and do the lookups and the calls on one connection, treating any reconnect as a new session.

go deeper

for a junior

The takeaway to hold on to is that a gRPC server describes itself, so two servers running different builds honestly give different answers. Nothing is malfunctioning when that happens.

for a middle

Explain that one reflection call is one stream on one connection and therefore one server, which is why answers are consistent inside a call and not across calls, and why a reconnect changes which build replies.

for a senior

Walk the whole failure: cached descriptor from the newer build, request routed to the older one, call succeeds, work silently not done. Then give the fix as affinity — one addressed instance, one connection, re-fetch after any reconnect.

for a principal

The question you own is whether production instances answer reflection at all, weighing operability during an incident against publishing an API's complete shape to anything that can reach the port, and making that a stated default rather than an accident.

## The symptom A curator opens a generic client against the catalogue's pooled address, looks up the accession message, and gets a definition with eight fields. Ten minutes later the same lookup returns nine. Nobody deployed anything in between — the rollout that added the ninth field started an hour ago and is still half finished. Nothing is broken. The tool is reporting the truth about two different servers and presenting it as though it were the truth about one service. ## Why a descriptor is a fact about one instance Reflection is answered from the descriptors compiled into the running binary. There is no shared registry behind it and no cluster-wide store; each server describes **itself**. That makes a reflection answer a fact of the form "the instance that answered this call was built from this schema", which is a narrower claim than "the catalogue service looks like this". During a rollout, both claims are true simultaneously and of different machines. Outside a rollout the distinction is invisible, which is exactly why it surprises people the first time it matters. ## Two things scatter a session across builds - **A new call may reach a new instance.** A reverse proxy in front of a pool routes what it is given; nothing binds a second reflection call to the instance that answered the first. The same is true of a client-side balancer choosing a connection. - **A reconnect is a new routing decision.** A dropped connection, an idle timeout that cuts a long-lived one, or a client that simply opens a fresh connection for the next operation all hand the routing layer another chance to pick differently. What *is* stable is one reflection call: it is a single stream on a single connection, so everything asked and answered inside it comes from one build. Consistency within a call is guaranteed by the call shape. Consistency across calls is guaranteed by nothing at all. ## What actually breaks downstream The inconsistent descriptor is rarely the real damage; it is the warning. The failure that costs time looks like this: 1. The client fetches a descriptor from an instance running the newer build and caches it. 2. It composes a request using a field that only the newer build knows about. 3. The request is routed to an instance running the older build, whose schema has no such field. 4. The call succeeds. The field is simply not understood by the receiving build, so the work the curator asked for silently does not happen. A silent no-op is far worse than an error, and the diagnosis is genuinely hard because every individual component behaved correctly. Whether two schema versions can safely exchange messages at all is a schema-compatibility question and belongs elsewhere; the point here is that reflection made it easy to *believe* both instances shared one contract. ## Keeping a debugging session honest 1. **Pin the session to one server.** Address a single instance rather than the pooled name, so that the lookups and the calls you make from them are answered by the same build. 2. **Do it on one connection.** Reflection's own affinity only holds within a call; keeping the whole session on one connection extends the same property to the calls you compose. 3. **Treat a reconnect as a new session.** Re-fetch descriptors rather than carrying a cached one across a routing decision you did not make. 4. **Prefer a compiled contract when you need stability.** Reflection is a debugging and tooling surface. When correctness depends on the contract, the descriptors you built against are the ones to trust. ## The judgement underneath There is a second decision hiding in this scenario: whether reflection should be answered by production instances at all. It is genuinely useful during an incident, and it also publishes the complete shape of an API — every service, method and message — to anything that can reach the port. Teams reasonably land in different places: registered everywhere for operability, registered only in non-production, or registered on a listener that only trusted callers reach. What is not defensible is arriving at the answer by never asking the question, and then being surprised by what an unauthenticated peer can enumerate.

  • How do you keep a reflection-driven debugging session consistent against a rolling fleet?
    Give the session one server and keep it there: address a single instance rather than the pooled name, and do the lookups and the calls on one connection so the descriptors and the requests are handled by the same build. Treat a reconnect as a new session and re-fetch, rather than carrying a cached descriptor across a routing decision you did not make.
  • Should a server answer reflection outside a debugging context?
    Treat it as a decision rather than a default. Reflection publishes the full shape of an API — every service, method and message — to anything that can reach the port, which is exactly what someone probing it wants. Common positions are to register it only in non-production builds, or to keep it on a listener that only trusted callers can reach, while accepting that this costs you a tool during an incident.
  • Why is a silent no-op the likely outcome rather than an error?
    The caller built a valid message against the newer build's descriptor and sent it to an older build. From the receiving server's point of view the request decodes and the field it does not know is simply not part of its schema, so the call completes normally and the requested behaviour never happens. The call's outcome tells the curator nothing, which is why this is diagnosed late.

saying these in an interview costs you the question

  • Assumes a reflection answer describes the service rather than the instance that answered.
  • Caches a descriptor from one session and reuses it against a different build.
  • Assumes every new call in a debugging session reaches the same instance.
  • Blames the schema when the real cause is two builds serving at once.
  • Believes reflection is safe to answer on any reachable port without a decision.
  • Expects a mismatched request to fail loudly rather than quietly do nothing.