skip to content

A gRPC caller invokes an rpc method the deployed server never implemented — what comes back, and why is it not a transport error?

level: middleimportance: should knowfreq 50%

answer

  1. the server answered, so not transport
  2. one method fails, siblings succeed
  3. two builds, two schema revisions
  4. status twelve
  5. deploy the implementing side first

basics

~20 s

The call reaches the server and comes back with gRPC status UNIMPLEMENTED (12). The connection was healthy and the server answered — it simply has no handler for the method named, which signals contract skew rather than a connectivity problem.

solid answer

~40 s

The caller gets the gRPC status `UNIMPLEMENTED (12)`. Everything below the contract worked: the connection was established, the call was delivered, the server responded. What failed is the agreement about *which methods exist* — the caller was generated from a schema revision that declares the method and the deployed server was not. That is why it is not a transport error and why retrying is pointless: a healthy server will answer the same way every time. It is also why the symptom is per-method rather than global; the server's other methods keep succeeding throughout, which is the clue that separates this from `UNAVAILABLE (14)`, where the caller never reached a serving backend at all.

code

protobuf · 9 lines
protobuf
service PermitAdjudication {
  rpc SubmitApplication(SubmitApplicationRequest) returns (ApplicationReceipt);
  rpc GetApplication(GetApplicationRequest) returns (Application);

  // Added in this release. The deployed server predates it, so
  // a caller generated from this revision is answered
  // UNIMPLEMENTED (12) for this method only.
  rpc AmendApplication(AmendApplicationRequest) returns (ApplicationReceipt);
}

go deeper

for a junior

Know that this status means the server answered and does not have the method you named, so the network is not the thing to investigate.

for a middle

Explain the skew that produces it: two sides generated from different schema revisions, discovered at call time because nothing compares contracts beforehand.

for a senior

Diagnose from per-method error rates, spot the mixed-fleet variant where it appears intermittently, and fix it by ordering the deploys rather than by retrying.

for a principal

Make the ordering a rule rather than a habit — additive contract changes ship server-first, removals only after measured zero traffic — so no release depends on someone remembering.

## What the caller actually sees The call does not hang, does not time out, and does not report a broken connection. It completes, promptly, carrying the gRPC status `UNIMPLEMENTED (12)`. That is the framework saying: *I reached a server, and it does not provide the method you named.* Everything underneath the contract behaved. A connection existed, the call was delivered, a server processed it far enough to look the method up and answer. The only thing that failed is the shared assumption about which methods exist — and that assumption lives in the `.proto`, on two sides that were generated at different times. ## Why it is not a transport error, and how to tell The three statuses this gets confused with describe three different stages, and distinguishing them is most of the diagnostic value: | Status | What actually happened | Retry helps? | |---|---|---| | `UNAVAILABLE (14)` | the caller could not reach a serving backend at all | often, yes | | `UNIMPLEMENTED (12)` | a server answered and does not declare that method | no | | `INTERNAL (13)` | a handler that exists ran and failed | sometimes | The observable difference is **scope**. A connectivity fault hits every method on the service at once. Contract skew hits *exactly* the methods one side knows about and the other does not, while the rest of the service keeps serving normally. If three of four methods succeed and the fourth fails on every attempt, you are not looking at the network. ## How the situation arises Almost always through deploy order. The sequence is mundane: 1. Someone adds an `rpc` line to the shared schema and publishes the revision. 2. A consuming team rebuilds; their stub now has a member for the new method. 3. That team deploys. 4. The implementing team has not deployed yet — or has, but an older instance is still serving. Between steps 3 and 4, every invocation of the new method is answered `UNIMPLEMENTED (12)`. The same thing happens during a rollback, when instances at two revisions serve simultaneously, and during a partial rollout, where the status appears intermittently on some fraction of calls — which is how it acquires the reputation of being "flaky" when it is nothing of the kind. ## Reading it in production - **Per-method error rates are the instrument.** One method at 100% failure with its siblings at 0% is contract skew, not infrastructure. - **An intermittent version of the same pattern** means a mixed fleet: some instances carry the new implementation, some do not. - **Every method failing** points somewhere else — a renamed service or proto package, which changes the identity of all of them at once. - **Retries will not rescue it.** A retry policy aimed at transient failures should not treat this status as retryable; it burns the caller's deadline to reach the same answer. ## The fix is deploy order, not code When a release adds a method, the implementing side deploys first. Once the server declares and implements the method, a caller that has not upgraded simply never invokes it, and a caller that has upgraded finds it. Deploying the calling side first guarantees a window in which every call to the new method fails — a self-inflicted outage whose length is whatever the gap between the two deploys turns out to be. The reverse rule applies to removals: stop calling a method everywhere first, prove the traffic is zero, and only then take it out of the contract. ## Mistakes to avoid - Reading the status as a network or DNS problem and escalating to the wrong team. - Retrying, or worse, retrying with backoff, so the failure is slow as well as certain. - Expecting a not-found-flavoured status; the specification's mapping sends an HTTP `404` to `UNIMPLEMENTED` precisely because the method, not a resource, is what is missing. - Assuming the generated server base type ships a working default for methods nobody wrote. - Shipping the caller first and calling the resulting window flakiness.

  • How do you tell this apart from the server being down?
    By scope. A server that cannot be reached produces `UNAVAILABLE (14)` for every method on it. A server that is reachable but older produces `UNIMPLEMENTED (12)` for exactly the methods it does not declare, while every other method on the same service keeps succeeding. That per-method split is the giveaway.
  • Which side should deploy first when a release adds an rpc method?
    The implementing side. Once the method exists there, a caller that has not upgraded never invokes it and a caller that has upgraded finds it, so no window of failure exists. Deploying the calling side first guarantees one, lasting exactly as long as the gap between the two deploys.
  • Does the same thing happen if the service is renamed rather than a method removed?
    Yes, with a wider blast radius. The server declares nothing matching the name the caller sent, so the outcome is the same `UNIMPLEMENTED (12)` — but because the rename changes the identity of every method under that service at once, the failure is total rather than confined to one method.

saying these in an interview costs you the question

  • Reads UNIMPLEMENTED as a network or connectivity failure
  • Retries the call, expecting a healthy server to change its answer
  • Assumes a missing method comes back as a not-found style status
  • Thinks the generated base type supplies a working default implementation
  • Deploys the calling side first and calls the resulting errors flaky