skip to content

In a gRPC service config, how do retryPolicy and hedgingPolicy differ, and what keeps either from amplifying a backend outage?

level: seniorimportance: should knowfreq 34%

answer

  1. one policy per method, never both
  2. wait for failure versus don't wait
  3. a token budget per server name
  4. half of maxTokens stops retries
  5. response headers commit the call

basics

~20 s

A gRPC method config holds either a retryPolicy, which replays a call after a retryable status with backoff, or a hedgingPolicy, which sends staggered copies without waiting. retryThrottling, a per-server-name token budget, halts both when failures pile up.

solid answer

~50 s

Both sit in a `methodConfig` entry of the service config, and a method gets one or the other, never both. `retryPolicy` waits for failure: on a status in `retryableStatusCodes` the client library replays the call after exponential backoff (`initialBackoff`, `backoffMultiplier`, `maxBackoff`), up to `maxAttempts` counting the original. `hedgingPolicy` does not wait: it sends the original, then another copy every `hedgingDelay` up to `maxAttempts`, keeps the first OK and cancels the rest — copies can run on several backends, so only idempotent methods should be hedged. The brake is `retryThrottling`, kept per server name: counted failures take a token, successes add `tokenRatio`, and at or below `maxTokens / 2` no retry or further hedge is sent. `maxAttempts` (clamped to 5 by default), one deadline across all attempts and server pushback bound it too, and a call that has received response headers is committed and never retried.

code

json · 26 lines
json
{
  "loadBalancingConfig": [ { "round_robin": {} } ],
  "methodConfig": [
    {
      "name": [ { "service": "shop.inventory.v1.StockService", "method": "GetStock" } ],
      "timeout": "2s",
      "retryPolicy": {
        "maxAttempts": 4,
        "initialBackoff": "0.1s",
        "maxBackoff": "1s",
        "backoffMultiplier": 2,
        "retryableStatusCodes": [ "UNAVAILABLE" ]
      }
    },
    {
      "name": [ { "service": "shop.inventory.v1.StockService", "method": "ListWarehouses" } ],
      "timeout": "1s",
      "hedgingPolicy": {
        "maxAttempts": 3,
        "hedgingDelay": "0.05s",
        "nonFatalStatusCodes": [ "UNAVAILABLE", "INTERNAL" ]
      }
    }
  ],
  "retryThrottling": { "maxTokens": 10, "tokenRatio": 0.1 }
}

go deeper

for a junior

Recall that gRPC retries are declared in the service config, not coded by hand, and that a method uses either retrying or hedging, never both.

for a middle

Explain the fields: maxAttempts including the original, the backoff trio, retryableStatusCodes, and how hedgingDelay and nonFatalStatusCodes drive staggered copies and cancellation.

for a senior

Show you can stop retries making an outage worse: retryThrottling per server name, one deadline across attempts, pushback, and why early response headers silently disable retries.

for a principal

Weigh who owns the policy: service owners publish it to every client, so a careless retryable code or hedge on a non-idempotent method multiplies load across the whole fleet.

## Where the policy lives gRPC can re-send a failed call **inside the client library**, without the application writing a loop. What it does is not decided by the caller's code: it is declared by the service owner in the **service config**, the JSON document a gRPC channel receives from its name resolver alongside the address list (or that a client sets as a default when the resolver returns none). Each entry of `methodConfig` names the methods it applies to through `name` (a service and method, or a whole service) and may carry a `timeout` and **exactly one** of two policies: - **`retryPolicy`** — wait for an attempt to fail, then replay it. - **`hedgingPolicy`** — do not wait for a failure; send staggered copies and keep the first good answer. If neither is set, no configured retrying or hedging happens. The application cannot override the policy, though a client can switch retries off entirely. ## retryPolicy: wait, then replay | Field | Rule | Meaning | |---|---|---| | `maxAttempts` | required, greater than 1 | total attempts **including the original** | | `initialBackoff` | required, Duration string > 0 | delay before the first retry | | `backoffMultiplier` | required, number > 0 | growth factor per retry | | `maxBackoff` | required, Duration string > 0 | ceiling on the computed delay | | `retryableStatusCodes` | required, non-empty | codes that allow a retry, e.g. `["UNAVAILABLE"]` or `[14]` | When an attempt ends with a non-OK status, the client checks it against `retryableStatusCodes`. On a match it waits `initialBackoff × random(0.8, 1.2)` before the first retry and `min(initialBackoff × backoffMultiplier^(n-1), maxBackoff) × random(0.8, 1.2)` before the n-th — the general backoff-with-jitter pattern, fixed here as concrete fields. Any other status goes straight back to the application. Two bounds matter more than the curve. The call's **deadline covers every attempt**: a retry never gets a fresh one. And the retry logic sits **between the channel and the load-balancing policy**, so each attempt gets a new pick and may land on a different subchannel than the one that failed. ## hedgingPolicy: do not wait `hedgingPolicy` takes `maxAttempts` (required, greater than 1), an optional `hedgingDelay` and an optional `nonFatalStatusCodes`. The original is sent at once; if no successful response has arrived after `hedgingDelay`, another copy goes out, and so on up to `maxAttempts`. A `hedgingDelay` of `"0s"`, or none at all, sends every copy immediately. 1. The first **OK** response wins; every other outstanding copy is cancelled. 2. A status in `nonFatalStatusCodes` sends the next pending copy **immediately**, skipping its delay. 3. Any other status cancels all outstanding copies and returns that error. 4. If every copy fails, nothing further is attempted. Because copies run concurrently, typically on **different backends**, a hedged method may execute more than once. That is why hedging belongs only on methods that are safe to run several times — gRPC's retry design offers no way to mark a method idempotent, so the service owner carries that judgment (whether hedging pays off at all is a general tail-latency question). | | `retryPolicy` | `hedgingPolicy` | |---|---|---| | sends the next attempt | after a failure plus backoff | after `hedgingDelay`, failure or not | | attempts in flight | one at a time | up to `maxAttempts` at once | | status list | `retryableStatusCodes` (required) | `nonFatalStatusCodes` (optional) | | extra load when healthy | none | up to `maxAttempts - 1` copies per slow call | ## The brakes on amplification A retry storm turns a struggling backend into a dead one. The service config's brake is **`retryThrottling`**, set once **per server name** — never per method or service: 1. A `token_count` starts at `maxTokens` (range 0 exclusive to 1000) and stays between 0 and `maxTokens`. 2. A call that fails with a retryable or non-fatal status, or with a pushback saying not to retry, takes one token. A status outside those lists, such as `INVALID_ARGUMENT`, does not count. 3. Each successful call adds `tokenRatio` (greater than 0). 4. While `token_count` is **at or below `maxTokens / 2`**, no retry is made and no further hedge is sent. Nothing queues: the failure is returned to the application. The first attempt of every call still goes out. With `maxTokens: 10` and `tokenRatio: 0.1`, retries stop once failures run above roughly one per ten successes. The other bounds are `maxAttempts` — clamped by the client to 5 by default, so a config fetched from a resolver cannot demand a hundred attempts — the shared deadline, and **server pushback**: a response carrying `grpc-retry-pushback-ms` makes the client wait exactly that long before the retry, and a negative or unparseable value means do not retry. Each retry also carries `grpc-previous-rpc-attempts`, so the server can see it is not a first try. ## When a call can no longer be retried A call becomes **committed**, and then gets no further attempts, in two cases: - the client receives the server's **response headers** (initial metadata), which the application may already have acted on; - the outgoing messages **outgrow the client's retry buffer**, as a long client stream can, so they cannot be replayed. So servers should hold response headers until the first response message; an error raised before that travels as Trailers-Only and stays retryable. Separately, calls that never reached the server's application logic — one that never left the client, or a stream refused with `REFUSED_STREAM` — are retried **transparently** whatever the policy says, and those retries count against neither `maxAttempts` nor the throttle.

  • A gRPC server sends response headers the moment a streaming call opens. Why are its later UNAVAILABLE failures never retried?
    Receiving response headers commits the call: the initial metadata has reached the client and may have changed its state, so the client library makes no further attempts even if `UNAVAILABLE` is retryable and attempts remain. Servers should hold headers until the first response message; an error raised before that goes out as Trailers-Only and stays retryable.
  • With retryThrottling set to maxTokens 10 and tokenRatio 0.1, when do retries stop and how do they come back?
    The count starts at 10. Each failure with a retryable or non-fatal status takes one token, each success adds 0.1. At or below 5, retries and further hedges are simply not sent and the failure goes to the application. Successes push the count back above 5 and retries resume — in steady state, once failures exceed about one per ten successes.
  • Why does a gRPC retry often reach a different backend than the attempt that failed?
    The retry logic sits between the channel and the load-balancing policy, so every attempt is picked afresh. Under `round_robin` the retry normally goes to the next `READY` subchannel in the rotation; under `pick_first` it goes to whichever single address the policy is connected to at that moment.

saying these in an interview costs you the question

  • Application code must catch UNAVAILABLE and loop to get gRPC retries.
  • One method can carry both a retryPolicy and a hedgingPolicy for extra safety.
  • maxAttempts counts only the retries, not the original call.
  • Each retry attempt gets a fresh deadline of its own.
  • Hedging suits any method, since only the first response is kept.
  • A call that already received response headers can still be retried on UNAVAILABLE.