skip to content

Load Balancing and Discovery

Why one long-lived connection defeats naive balancing, and the two fixes: a resolver-driven client policy, or a sidecar. Interviewers ask because it is the first real operational surprise teams hit.

part ofAPI stylesoverview, primer and where to startread it →
on this pageshow

questions

7

In gRPC, what does a pick_first policy do with the resolved address list, and how does round_robin differ?

level: middleimportance: must knowfreq 70%

answer

  1. balanced once, at connect time
  2. one connection or one per address
  3. the unit has to be the call
  4. a managed connection with its own state
  5. rotation skips anything not ready

basics

~20 s

pick_first connects to the resolved addresses in order and sends every call over the first connection that comes up. round_robin opens a subchannel to each address and rotates calls across the ready ones, so a pool actually shares work.

solid answer

~40 s

`pick_first` is the default. It walks the resolved address list in order, keeps the first connection that establishes, and every call on that channel rides it — the other addresses are fallbacks, not peers. `round_robin` instead creates a **subchannel** per address, connects them all, and hands successive calls to the subchannels currently in `READY`, skipping the rest. The distinction matters because gRPC multiplexes many concurrent calls onto one long-lived HTTP/2 connection: if you balance per connection you balance once, at connect time, and a scaled-out pool stays unevenly loaded. Balancing has to happen **per call**, which is precisely what a subchannel-per-address policy gives you. Which policy runs is chosen by `loadBalancingConfig` in the service config the resolver supplies.

code

json · 5 lines
json
{
  "loadBalancingConfig": [
    { "round_robin": {} }
  ]
}

go deeper

for a junior

Remember that one gRPC channel normally means one connection, and that a connection sticks. More backends alone does not mean more of them get work.

for a middle

Explain both policies concretely: connect in order and keep the first, versus a subchannel per address with calls rotated across the ready ones. Say where the choice is configured.

for a senior

Demonstrate the cost side — N connections per caller, and rotation that is blind to per-call cost — and say when you would push the decision out of the caller entirely.

for a principal

The judgment is whether uniform per-call cost holds across your services. Positional rotation is a fair share only under that assumption, and standardising on it quietly bets the fleet on it.

## Why connection-level balancing quietly fails The usual way to spread load is to balance connections: a caller opens a connection, something in the middle picks a backend, and because callers open connections constantly the picks average out. gRPC breaks that assumption in one specific way. Every call is one HTTP/2 stream, many streams ride one connection at once, and that connection is opened once and held for the life of the process. A caller that made one connection an hour ago makes no further balancing decisions at all. The symptom is unmistakable the first time a pool is scaled out. Add four workers to a fingerprint-matching pool, watch the dashboards, and the original worker is still at full CPU while the new ones idle. Nothing is broken; every caller simply resolved a name once, connected once, and has been pinned ever since. The fix is to move the decision from **per connection** to **per call**, and in gRPC the component that does that is the client-side balancing policy. ## pick_first `pick_first` is the default policy and it is exactly what its name says: 1. Take the address list the resolver produced, in the order given. 2. Attempt to connect to them in that order. 3. Keep the first connection that establishes, and route every call over it. 4. On failure, re-resolve and start again. It is the right default for a single server, for a local socket, and for the very common case where something *else* — a data plane in front of the pool — is already doing the spreading. It is the wrong default the moment the address list genuinely contains peers you wanted to share work. ## round_robin and subchannels `round_robin` introduces the **subchannel**: one managed connection to one address, with its own connectivity state. The policy creates a subchannel per resolved address, asks them all to connect, and then, for each call, picks the next subchannel in rotation from the ones currently in `READY`. Subchannels that are connecting or failed are skipped and rejoin the rotation when they become ready. | | `pick_first` | `round_robin` | |---|---|---| | Connections opened | one, to the first address that answers | one per resolved address | | Balancing granularity | none after connect | per call | | A new address appears | ignored until the connection fails | a subchannel is added for it | | One endpoint dies | all calls fail until reconnect | that subchannel leaves rotation, others carry on | | Cost | one connection per caller | N connections per caller | That last row is the honest trade. A hundred callers against forty workers means four thousand connections instead of a hundred, each with its own transport state, its own keepalive and its own memory on both sides. ## Where the policy comes from The application does not usually choose the policy imperatively. The name resolver returns a **service config**, and its `loadBalancingConfig` field names the policy for the channel. That indirection is the point: policy can be changed centrally for every caller of a service without any of them being recompiled. There was an earlier design, `grpclb`, in which the client asked an external **look-aside balancer** over its own connection for the current address list plus load reports, then connected directly to the addresses it was given. It is deprecated; policy configured through the resolver's service config replaced it. ## What round_robin does not give you Rotation is positional, not informed: - it does not know how long a call will take, so a mix of cheap and expensive work still skews; - it does not know a worker's queue depth or CPU, so a slow worker gets the same share as a fast one; - it does not know whether the process behind a `READY` subchannel can actually serve — a connection accepting bytes says nothing about the application's health, which is why a balancer that wants to eject a sick endpoint has to consume a health signal separately; - it does not rebalance an established set: new addresses only arrive with a re-resolution. For a fingerprint-matching pool this is usually acceptable, because match latency is roughly uniform, and that uniformity is what makes positional rotation approximate a fair share. Where per-call cost varies by an order of magnitude, positional rotation is closer to random assignment than to balancing, and the next honest step is a policy that consumes a load signal or a data plane that does.

  • Why is round_robin blind to how expensive each call is?
    It rotates positionally across the subchannels in READY. Nothing feeds back per-call cost, queue depth or CPU, so a worker handed a run of expensive calls receives exactly the same share as one handed cheap ones. Rotation approximates fairness only when per-call cost is roughly uniform.
  • What was the deprecated look-aside balancer, and what did it add?
    grpclb. The client opened a connection to an external balancer service, received the current backend address list plus load reports, and then connected directly to those backends. It is deprecated; policy supplied through the resolver's service config took its place.
  • A load-generating client points one channel at a pool and reports one worker saturated. What is wrong?
    Nothing is broken. The generator holds a single channel under the default policy, so every synthetic call rides one connection to one worker no matter how much concurrency the tool configures. Select a per-call policy, or run several channels, before reading the numbers as a pool measurement.

round_robin is a bank with four open tellers and one queue that feeds them in turn. pick_first is a bank where everyone who walks in all day is sent to whichever teller happened to unlock the door first.

saying these in an interview costs you the question

  • Says adding backends spreads gRPC calls automatically, as stateless HTTP requests would
  • Thinks round_robin opens a fresh connection for every call
  • Believes pick_first probes all addresses and keeps the fastest
  • Assumes round_robin routes by worker load rather than by position
  • Thinks the balancing policy is a server-side setting
open as a page

One worker in a round_robin gRPC channel stops accepting connections — what happens to its subchannel and to the channel's own state?

level: seniorimportance: must knowfreq 55%

basics

~10 s

That worker's subchannel moves to TRANSIENT_FAILURE and the picker drops it from rotation until it reconnects and reports READY. The channel itself stays READY while at least one other subchannel is READY.

open as a page

In gRPC, what does the scheme at the front of a channel target string select, and which port applies when none is given?

level: juniorimportance: should knowfreq 45%

basics

~10 s

The scheme selects the name resolver: dns: resolves a name to addresses, ipv4: and ipv6: take literal addresses, unix: points at a local socket. A dns: target that names no port defaults to 443.

open as a page

A gRPC server starts closing a caller's connections with the debug data too_many_pings — what did the caller's keepalive settings do wrong?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Keepalive pings were sent more often than the server tolerates, typically on connections with no active calls. The server counted strikes and, past its limit, sent a GOAWAY carrying ENHANCE_YOUR_CALM and the debug data too_many_pings.

open as a page

In a gRPC service config, how do retryPolicy and hedgingPolicy differ, and what keeps either from amplifying a backend outage?

level: seniorimportance: should knowfreq 34%

basics

~20 s

A gRPC method config holds either a retryPolicy, which replays a call after a retryable status with backoff, or a hedgingPolicy, which sends staggered copies without waiting. retryThrottling, a per-server-name token budget, halts both when failures pile up.

open as a page

Across an organisation, would you put gRPC balancing in every client library or terminate every connection in a sidecar data plane?

level: principalimportance: should knowfreq 42%

basics

~20 s

Client-side balancing costs no extra hop but duplicates the policy in every language and ships only when callers redeploy. A sidecar centralises the policy at the price of a hop, a process per workload, and a new failure domain.

open as a page

A gRPC caller's channel was created hours ago, so a freshly doubled worker pool receives nothing from it — what moves those calls?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Nothing in gRPC rebalances an established connection on its own. New workers are picked up only when the caller re-resolves, which happens after a subchannel fails or after the server bounds connection lifetime and closes gracefully.

open as a page