Ten replicas of a gRPC service sit behind a connection-level (layer 4) load balancer, but two replicas serve almost all the traffic and replicas added by autoscaling stay idle. Why does connection-level balancing produce this, and what fixes it?
answer
- balancer counts connections, not calls
- one decision, hours ago
- new replicas never get dialled
- least-connections counts the wrong thing
- make connections mortal
basics
~20 sA connection-level balancer picks a backend once per connection, and a gRPC client keeps one long-lived multiplexed connection carrying every call. Load follows connections, not requests, so a handful of pinned connections concentrate traffic and new replicas never receive one.
solid answer
~50 sgRPC runs over HTTP/2, where a client holds one long-lived connection and multiplexes thousands of calls onto it. A layer 4 balancer makes exactly one backend choice, at connect time, so every call on that connection lands on the same replica for as long as it stays open. With a small number of client processes you get a small number of connections, spread by luck rather than by load, and a replica that appears after those connections were established gets nothing at all. Three fixes: put a request-level, HTTP/2-aware proxy in the path so it balances per RPC; or move to client-side balancing, where the client resolves every endpoint and picks per call; or, if the layer must stay connection-level, bound connection lifetime so servers periodically close connections gracefully and clients re-dial into the current pool. Switching the algorithm to least-connections does not help, because it still counts connections.
go deeper
Know that a connection-level balancer picks a backend when the connection opens, and that a client which keeps one connection open therefore keeps hitting one backend.
Explain multiplexing: gRPC over HTTP/2 carries many calls on one connection, so balancing connections stops approximating balancing load, and new replicas get nothing until new connections are made.
Show the diagnosis — per-replica request rate against per-replica connection count — and weigh the three fixes, including why an autoscaler keeps adding capacity that receives no traffic.
Own the general rule that a balancing layer is only as good as the correlation between the unit it counts and the work being done, and decide where per-call balancing lives platform-wide: shared proxy, mesh sidecar, or client library.
## The mechanism The imbalance is not a bug in the balancer — it is the defining property of the altitude it operates at. A connection-level balancer sees connections. It applies its algorithm once, when a connection arrives, and after that it is a byte pipe for the life of that connection. For classic browser traffic this is nearly invisible: connections are numerous, short-lived and constantly re-established, so "balance the connections" approximates "balance the load" well enough. Multiplexed protocols destroy that approximation. gRPC (and any HTTP/2 client with a long-lived pool) opens one connection to an endpoint and runs every call as a stream on it. Ten client processes might produce ten connections total, and those ten are spread across your replicas by whatever the algorithm decided at connect time — hours ago. So the load a replica receives is proportional to how many long-lived connections happen to be pinned to it, multiplied by how busy those particular clients are. Two replicas holding the connections of your two chattiest callers will carry most of the traffic no matter how even the connection count looks. ## Why autoscaling makes it worse Scaling out only helps if new capacity receives traffic. New replicas receive traffic only when new connections are made. If the client fleet is stable and its connections are already established, a new replica sits at zero utilisation while the balancer reports a perfectly healthy pool — and the autoscaler, watching average CPU across the pool, may keep adding replicas that also do nothing. Scale-in then removes idle replicas rather than the overloaded ones. It is an unusually frustrating incident because every dashboard at the pool level looks fine. ## Diagnosing it Look at **per-replica** request rate and concurrency, not pool totals, and put them next to per-replica established-connection counts. The signature is a small integer number of connections per replica, wildly different request rates, and a strong correlation between the two. If a replica's request rate is flat from the moment it started, it never received a connection. ## Fix 1: move the balancing decision up a layer Put an HTTP/2-aware request-level proxy in front of the service. It terminates the client connection, sees each stream as an individual request, and chooses a backend per call over its own pooled upstream connections. This is the direct fix: it changes the unit of balancing from connection to request. The price is the usual layer 7 price — the proxy must speak HTTP/2 end to end, must not downgrade the protocol, and becomes another hop with its own CPU, state and timeouts. ## Fix 2: move the decision to the client With client-side balancing the client resolves the full set of endpoints, keeps a connection to each, and picks one per call. There is then no intermediary making the wrong-granularity decision at all, and latency is lowest. The costs are real: every client needs the balancing implementation, they need a way to learn endpoint changes promptly, and connection count becomes clients times endpoints, which is exactly why a service mesh sidecar exists — it gives per-call balancing without every client library implementing it. ## Fix 3: stop connections living forever If you must keep the connection-level tier, take away the assumption it depends on. Servers can cap connection age and shut a connection down gracefully — for HTTP/2, by sending a GOAWAY that lets in-flight streams finish while the client re-dials. Add jitter so the whole fleet does not reconnect at the same instant. Rebalancing then happens continuously instead of never. This is mitigation rather than a fix: distribution is still per connection, just resampled often enough to track the pool. ## The trap: changing the algorithm The reflex answer is to switch the balancer from round-robin to least-connections. It does not help. Least-connections counts *connections*, and every one of these counts as exactly one whether it is carrying one call per minute or a thousand per second. Hash-based or affinity settings make it strictly worse by pinning harder. Nothing at this altitude can measure request load, because nothing at this altitude can see requests. ## The general lesson Any protocol that multiplexes many logical operations onto one long-lived connection — gRPC, HTTP/2 and HTTP/3 APIs, some message-broker and database clients — breaks connection-level balancing in the same way. When you meet imbalance, ask what unit the balancing layer is actually counting, and whether that unit still correlates with work.
- Why does switching the balancer to a least-connections algorithm not fix this?Because it still counts connections. A connection carrying one call a minute and one carrying a thousand a second both count as one, so the metric it optimises has no relationship to load. Hash or affinity settings are worse still — they pin connections harder. No algorithm at the connection level can see request volume.
- What does client-side balancing cost compared with putting an HTTP/2-aware proxy in front?It removes a hop and gives the lowest latency, but every client must implement the balancing and endpoint-discovery logic, keep it consistent across languages, and hold a connection per endpoint — so connection count grows as clients times endpoints. A mesh sidecar is the middle ground: per-call balancing without changing client libraries.
- If you cap connection lifetime to force rebalancing, what must you be careful about?Shut down gracefully — signal the client that no new streams should start, let in-flight ones finish, then close — and jitter the cap so the fleet does not reconnect in lockstep. Too short a cap turns handshake cost and reconnect churn into a new problem; too long and it never rebalances.
saying these in an interview costs you the question
- Blaming the balancing algorithm rather than the balancing unit
- Believing least-connections tracks in-flight requests
- Adding replicas and expecting them to pick up traffic
- Enabling stickiness to 'stabilise' the distribution
- Assuming HTTP/2 clients open a connection per call