A round-robin load balancer sends each backend the same number of requests, yet one backend sits at 90% CPU while another is nearly idle. Explain how that happens, and what least-connection and weighted balancing each change.
answer
- counts requests, not work
- request cost is not uniform
- long-lived connections never re-balance
- in-flight count as a load signal
- fast failures look idle
basics
~20 sRound-robin equalises request counts, not work. When request costs or backend capacities differ, equal counts produce unequal load. Least-connection routes by in-flight requests so busy backends are skipped; weighting fixes only known, static capacity differences.
solid answer
~50 sRound-robin's fairness is counted in requests, and requests are not equal. One backend can draw the expensive report queries while another gets health pings; a slow request occupies a worker for seconds while a cached one returns in a millisecond; long-lived connections such as WebSockets or gRPC streams accumulate permanently on whichever backend happened to accept them. Heterogeneous hardware or a noisy neighbour produces the same skew from the other direction. Least-connection changes the signal: the balancer tracks in-flight requests per backend and sends the next one to the smallest count, so a backend still chewing on thirty slow requests stops receiving new work automatically. That reacts to dynamic imbalance, which round-robin cannot. Weighted round-robin only encodes a static ratio you supply — useful when one backend genuinely has twice the cores, useless when the imbalance comes from the traffic mix. Least-connection has its own trap: a backend failing instantly has the fewest in-flight requests and attracts the most traffic.
go deeper
Know that round-robin simply takes turns and that taking turns is not the same as sharing work evenly. Be able to name least-connection as the alternative that looks at how busy each backend currently is.
Explain concretely why equal counts give unequal load — cost variance, long-lived connections, unequal hardware — and what signal least-connection substitutes for the counter.
Show the diagnosis: compare per-backend counts against CPU and latency, and recognise the black-hole case where a fast-failing backend attracts traffic under least-connection.
Own the defaults for the fleet: which algorithm is standard, whether static weights are allowed at all given how they rot, and what evidence a team must bring before overriding the platform default.
## What round-robin actually promises Round-robin walks the backend list in order and hands each new request to the next entry. Its guarantee is exact and narrow: over time every backend receives the same *number* of requests. Nothing in the algorithm looks at CPU, memory, queue depth, response time, or whether the previous request to that backend has even finished. If every request cost the same and every backend were identical, count-fairness would equal load-fairness. Neither assumption survives contact with production. ## Four ways equal counts become unequal load **Request cost variance.** A single endpoint mix can span four orders of magnitude — a cached lookup at one millisecond, a report query at ten seconds. Round-robin deals them out like cards, so over a short window one backend can accumulate several expensive requests purely by arrival order. The larger the variance in cost, the longer the window over which the imbalance persists. **Long-lived connections.** At the connection level — WebSockets, gRPC streams, server-sent events, or plain HTTP keep-alive when the balancer routes per connection rather than per request — the balancing decision is made once and then lasts for the connection's lifetime. Backends restarted most recently hold the fewest long-lived connections, and a backend that has been up longest slowly accretes them. This is why a fleet can look badly skewed after a rolling restart even under perfect round-robin. **Heterogeneous capacity.** Mixed instance types, mixed generations of hardware, or containers with different CPU limits mean equal request counts land on unequal machines. **Noisy neighbours and local state.** A backend sharing a host with a busy tenant, or one whose cache is cold after a restart, does the same work more slowly. Its own throughput drops while its share of arrivals does not. ## What least-connection changes Least-connection (sometimes least-request) picks the backend with the fewest requests currently outstanding from this balancer. In-flight count is a proxy for how busy a backend is: a backend that is slow to finish work accumulates outstanding requests and is automatically passed over until it drains. That makes it self-correcting in exactly the cases round-robin is blind to — variable request cost, temporary slowness, cold caches, uneven long-lived connections. The signal has two important limits. First, the count is *local to this balancer*. With several proxies in front of one pool, each sees only its own share, so the global picture can still be lopsided. Second, and more dangerous, **fast failure looks like idleness**: a backend that immediately returns 500s or refuses to do work completes requests instantly, so its in-flight count stays at zero and least-connection preferentially feeds it. This is the classic black-hole pattern, and the mitigation is not in the algorithm — it lives in the health-checking and outlier-ejection layer that removes such a backend from the pool. ## What weighting changes, and what it does not Weighted round-robin (or weighted least-connection) attaches a static integer to each backend and deals proportionally: weight 3 receives three requests for every one that weight 1 receives. This is the right tool for *known, stable* differences — a pool where half the machines have twice the cores, or a deliberate traffic split while you validate a new instance size. It is the wrong tool for anything dynamic, because the weight is a number you decided in advance and traffic changes hourly. Weights also age badly: they are typically set once during an incident and then quietly outlive the reason for them. ## How to choose For homogeneous backends serving short, uniform-cost requests, round-robin is entirely adequate and has the cheapest implementation (no per-backend state, no coordination). Reach for least-connection when request durations vary widely, when a backend can be temporarily slow rather than down, or when connections are long-lived. Use weights only for capacity ratios you can state and would notice going stale. And regardless of algorithm, remember what balancing cannot do: it distributes work among backends that are in the pool, so a backend that should not be receiving traffic at all is a health-check problem, not a balancing one. ## Confirming the diagnosis Do not argue from the algorithm — measure. Compare per-backend request counts against per-backend CPU and per-backend p99 latency. Equal counts with skewed CPU is the cost-variance or capacity story. Skewed counts under round-robin points at something else entirely: connection-level pinning, affinity, or backends that were out of the pool for part of the window.
- Your traffic is mostly long-lived WebSocket connections. Which algorithm helps, and what still does not?Least-connection helps at accept time, since a backend already holding many connections is skipped for new ones. Nothing re-balances an established connection, though — a backend restarted mid-day will stay under-loaded until clients reconnect. Bounding connection lifetime, or draining and forcing reconnects, is what actually redistributes them.
- When is plain round-robin the better choice despite these weaknesses?When backends are homogeneous and requests are short and uniform in cost. Round-robin needs no per-backend state, so it is trivially correct across multiple proxies, cheaper per request, and perfectly reproducible — which also makes traffic patterns easier to reason about during an incident.
- Why can least-connection send more traffic to a broken backend than to healthy ones?Because it measures outstanding requests, not successful ones. A backend that fails instantly finishes every request immediately, so its in-flight count stays lowest and it keeps winning the selection. Only removing it from the pool — via health checking or outlier ejection — stops that.
saying these in an interview costs you the question
- Claims round-robin distributes load evenly by definition
- Treats least-connection as always superior with no downside
- Sets weights to fix an imbalance whose cause is dynamic
- Confuses connection count with CPU load for long-lived streams
- Ignores that a fast-failing backend looks least loaded