A gRPC caller's channel was created hours ago, so a freshly doubled worker pool receives nothing from it — what moves those calls?
answer
- the snapshot is as old as the channel
- lookups happen on events, not clocks
- healthy connections never ask again
- someone has to end the connection
- not all at the same instant
basics
~20 sNothing in gRPC rebalances an established connection on its own. New workers are picked up only when the caller re-resolves, which happens after a subchannel fails or after the server bounds connection lifetime and closes gracefully.
solid answer
~50 sRe-resolution is **event-driven**, not periodic: a channel asks its resolver again when a subchannel or the channel enters failure, or when its policy requests it — not on a timer. So a caller whose connections are all healthy never learns about the new addresses. Three levers exist. Under a per-call policy, a re-resolution adds subchannels for the new addresses, so anything that triggers one spreads the load; under the default single-connection policy even that is not enough, because the established connection keeps carrying every call. The reliable server-side lever is to **bound connection lifetime**: after a bounded period the server sends a `GOAWAY`, existing calls finish, the caller reconnects and re-resolves and now sees the full pool. Stagger it, or the whole fleet reconnects at once. The third option is to stop balancing in the caller at all and terminate in a sidecar.
code
pseudocode · 7 linesfor each accepted connection:
lifetime = base_lifetime + random_spread // stagger, so the fleet does not return at once
at (connection.opened_at + lifetime):
send GOAWAY // no new streams; running calls continue
wait until in_flight_streams == 0 or grace_deadline
close transport
// the caller reconnects, resolves again, and now sees every worker in the poolgo deeper
Remember that a gRPC connection is long-lived, so a caller keeps talking to whichever workers existed when it connected.
Explain that re-resolution is triggered by events rather than by a clock, and say which events those are.
Name the lever: bound connection lifetime and close gracefully so callers reconnect and resolve again, with the closes staggered so the fleet does not return in one burst.
Decide who owns the event that refreshes the fleet's view. A platform that never ends a connection has no rebalancing story, and capacity changes then depend on failures to take effect.
## The property that makes this happen Two gRPC behaviours combine into the surprise. The first is that a connection is long-lived and multiplexes every concurrent call from that caller. The second is that **re-resolution is event-driven**: a channel consults its name resolver when it starts, and again when a connection or subchannel fails and the balancing policy asks for a fresh answer. There is no background timer re-reading the address list for the life of the process. Put those together and a healthy caller is, by design, working from a stale snapshot. Adding workers changes the pool; nothing about adding them reaches a caller whose connections are all fine. ## Why the policy alone may not save you It helps to be precise about what a per-call policy does and does not fix here: - Under the default single-connection policy, even a fresh address list changes nothing while the established connection is healthy — the policy is still routing every call over the connection it already has. - Under a per-call policy with one subchannel per address, a re-resolution genuinely does add subchannels for the new workers and they enter the rotation. The problem is narrower but real: something still has to *cause* that re-resolution. So the question is not only "which policy" but "what event". In a pool that is scaled out deliberately, on a healthy day, there is no failure to supply one. ## The server-side lever The reliable answer is to stop treating a connection as permanent. A server can bound how long it will keep any one connection and then close it gracefully: 1. The server decides this connection has lived long enough. 2. It sends a connection-closing frame, which stops new streams from starting on that connection while the streams already running are allowed to finish. 3. When the last call completes, or a grace period expires, the transport closes. 4. The caller reconnects — and because reconnecting means resolving again, it now sees every address in the pool, including the new workers. That is the whole mechanism by which long-lived connections rebalance at all. Without something that ends connections, a client-side balancer's picture of the pool is frozen at whatever it was when it last had a problem. Two details make the difference between this working and hurting: - **Stagger the closes.** If every connection is bounded by the same fixed lifetime, connections opened during the same deploy expire together and the whole fleet reconnects in one burst — the load spike lands precisely on the pool you were trying to relieve. A spread of lifetimes fixes it. - **Close gracefully, not abruptly.** Ending the connection with a frame that lets running calls finish costs nothing, while dropping the transport fails every in-flight call for no reason. ## The third option The other honest answer is to move the problem. If callers hold a connection to a local sidecar data plane, and that process holds the connections to the workers and balances per call, then scale-out is picked up by the sidecar and no caller has to re-resolve anything. What you have bought is a component that can be updated without touching any caller; what you have paid is an extra hop, an extra process per workload, and another thing whose own failures you now diagnose. ## What not to reach for A few instincts make things worse: - **Restarting callers.** It works, and it is a deploy to fix a capacity change. - **Shortening resolution caching in the hope of a timer.** The channel is not re-resolving on a timer in the first place, so a lower cache time changes nothing while connections are healthy. - **Cutting connections abruptly from the server.** This does rebalance, at the price of failing every call in flight. - **Treating the pool as balanced because a dashboard shows even CPU across long windows.** Average it over an hour and a badly skewed pool can look flat; the signal that matters is share of calls per endpoint from each caller. ## The rule to carry away Balancing in gRPC is only ever as fresh as the last event that made a caller look again. If the design has no event that fires on a normal day — no bounded connection lifetime, no intermediary that re-decides, no periodic churn — then scaling a pool out will not move traffic, and the fix is to create that event deliberately rather than to wait for a failure to supply one.
- Why is re-resolution not enough on its own?Because it is event-driven. A fresh address list is only fetched when a connection or subchannel fails or the policy asks, and on a healthy day a deliberate scale-out supplies no such event. Under the default single-connection policy a fresh list would not move calls anyway.
- What does moving the decision into a sidecar change here?The caller keeps one connection to a local process, and that process holds connections to every worker and balances per call, so scale-out is absorbed without the caller re-resolving anything. The price is an extra hop, a process per workload, and one more failure domain to diagnose.
saying these in an interview costs you the question
- Assumes a client re-resolves its target on a timer regardless of connection state
- Thinks adding workers rebalances existing connections by itself
- Believes closing every caller's connection at the same moment is harmless
- Confuses a stream reset, which ends one call, with closing a connection
- Assumes a caller's address list is always current