skip to content

questions

6

What problem does a load balancer solve, and what is the practical difference between the round-robin and least-connections algorithms for distributing requests across backend servers?

level: juniorimportance: must knowfreq 85%

answer

  1. turns vs live counters
  2. capacity + availability, two jobs
  3. round-robin blind to cost
  4. least-connections self-corrects
  5. weighted variants for uneven fleets

basics

~20 s

A load balancer spreads incoming requests across many servers so no single one gets overloaded and the app stays up if one server dies. Round-robin just takes turns in order; least-connections sends the next request to whichever server currently has the fewest requests in flight.

solid answer

~40 s

A load balancer sits in front of a pool of identical backend instances and decides, per request, which instance handles it, giving you horizontal scalability (add instances for more capacity) and availability (route around a dead instance). Round-robin cycles through the pool in fixed order, which is simple and cheap and works fine when requests are roughly equal cost and instances are equally sized. Least-connections tracks each instance's current in-flight request count and always routes to the lowest, which adapts automatically when request cost varies (some calls are cheap, some do heavy DB work) or instances have different capacity. In production, least-connections or a weighted variant is usually the safer default; plain round-robin can pile slow requests onto an already-busy server.

go deeper

for a junior

Should know a load balancer spreads traffic and that round-robin cycles in order while least-connections picks the least-busy server; no need to discuss weighting or warm-up.

for a middle

Should articulate why least-connections handles variable request cost better and identify the connection-count bookkeeping it requires.

for a senior

Should discuss weighted variants, the slow-start problem for newly added instances, and choosing an algorithm based on the workload's cost variance.

for a principal

Should reason about algorithm choice at fleet scale - coordination cost across multiple balancer nodes, heterogeneous instance types in autoscaling groups, and how algorithm choice interacts with rolling deploys.

## What a load balancer is and why it exists A load balancer is the component that sits between clients and a fleet of backend servers, accepting incoming connections or requests and deciding, per request, which backend instance actually handles it. It exists to solve two problems that appear the moment an application outgrows a single machine. 1. **The first is capacity:** a single server has a ceiling on CPU, memory, and network throughput, so once traffic exceeds that ceiling you must run several copies of the service and spread traffic across them. 2. **The second is availability:** if a client talks directly to one server and that server crashes or is redeployed, every one of that client's requests fails; a load balancer can detect a dead instance and simply stop sending it traffic, so the failure of one box does not become an outage for the whole service. Underneath, a load balancer maintains a **pool** (sometimes called an upstream or target group) of backend addresses and applies an algorithm to pick one for each new request or connection. ## Round-robin: taking turns Round-robin is the simplest algorithm: the balancer keeps a pointer into the pool and advances it after every request, so `server1` gets request 1, `server2` gets request 2, and so on, wrapping back to `server1` after the last entry. - **Cheap, and even in request count.** It requires no runtime state about the servers themselves beyond the pool membership, which makes it cheap to implement and easy to reason about, and it delivers a genuinely even distribution of request *count* over time. - **Blind to actual load.** Its weakness is that it distributes requests, not work. If request costs vary, one server can end up doing far more CPU-seconds of work than another purely by bad luck of ordering, and if one server is already bogged down handling a slow request, round-robin will still send it the next one in strict turn. ## Least-connections: deciding on live load Least-connections fixes exactly that blind spot by making the decision load-aware: the balancer tracks how many requests (or, for L4 balancers, how many open connections) each backend currently has in flight and always forwards the next request to whichever backend has the fewest. This means: - a backend that is stuck processing a slow query naturally receives fewer new requests until it catches up; - a backend that just finished a batch of fast requests gets more work sooner. This **self-correcting** property makes least-connections noticeably better than round-robin whenever request duration is heterogeneous - which describes almost all real web and API traffic, where some endpoints are a cache hit and others trigger a multi-table join or an outbound call to a third party. The cost is a small amount of bookkeeping: the balancer (or, in a distributed setup, each balancer node) must maintain a live counter per backend and update it on request start and completion, which is trivial for a single load balancer process but requires coordination or approximation when load is split across many balancer instances that do not share state. ## Weighted variants for uneven fleets A closely related refinement is weighted variants of both algorithms. | Variant | How the weight is applied | |---|---| | **weighted round-robin** | gives faster or bigger instances a proportionally larger share of the turns | | **weighted least-connections** | divides the connection count by a capacity weight before comparing, so a server rated at twice the capacity is allowed roughly twice the concurrent connections before it stops receiving new ones | This matters in heterogeneous fleets - for example during a rolling deploy where new instances are still warming up caches and JIT-compiling hot paths, or in mixed-instance-type autoscaling groups on AWS where a `c5.xlarge` and a `c5.2xlarge` sit in the same target group. Real load balancers such as **NGINX**, **HAProxy**, and AWS's Application/Network Load Balancers default to round-robin or a randomized-choice variant for simplicity but expose least-connections and weighted modes as configuration, and HAProxy in particular also supports a hybrid 'first' policy and consistent-hashing-based balancing for cache-affinity use cases. The engineering decision in practice is rarely 'round-robin vs least-connections' in the abstract; it is choosing the algorithm that matches the traffic's variance in per-request cost, and re-checking that choice whenever the workload mix changes, for instance when a new expensive endpoint is added to a service that was previously uniform. ## A subtler option: least-response-time There is also a subtler algorithm worth naming: least-response-time (sometimes called weighted-response-time), which factors in not just the current connection count but each backend's recent observed latency, favoring instances that are both lightly loaded and currently fast. This catches a case plain least-connections misses: - a backend can have a low connection count simply because it is slow and requests are backing up elsewhere; - conversely a backend can look 'busy' by connection count while actually being extremely fast per request. Response-time-aware balancing needs the load balancer to track a rolling latency metric per backend, which is more bookkeeping than a raw counter but produces routing decisions that track real-world performance more closely. ## Every algorithm has edge cases None of these algorithms are free of edge cases: any algorithm that reacts to a live signal (connection count, latency) can overreact to noise, and any algorithm that ignores load (plain round-robin) can systematically starve or flood a backend under skewed traffic patterns, which is why production teams typically pick the simplest algorithm that actually matches their traffic's characteristics rather than reaching for the most sophisticated option by default.

  • When would round-robin actually be preferable to least-connections?
    When requests are near-uniform in cost and you want the absolute simplest, statistically most predictable distribution with zero per-backend bookkeeping - for example a fleet of identical stateless workers each doing the same fixed-cost transformation. It also avoids the coordination overhead least-connections needs when multiple load balancer nodes must share connection counts.
  • How does least-connections behave right after you add a brand-new, empty backend to the pool?
    Because the new instance starts at zero in-flight requests, least-connections will slam it with a disproportionate burst of new traffic until its count catches up to the rest of the pool, which can overwhelm an instance that is still warming up caches or JIT compilation. Production setups mitigate this with 'slow start' / connection ramp-up, where the balancer artificially caps how much traffic a newly added backend can receive for the first minute or two.
  • Does least-connections work the same way for an L7 HTTP balancer as for an L4 TCP balancer?
    Not quite - an L4 balancer only sees connections, so 'least connections' literally means fewest open TCP connections, which is a poor proxy for load if connections are long-lived and multiplexed (e.g. HTTP/2). An L7 balancer can count actual in-flight requests within those connections, which is a much more accurate load signal for HTTP traffic.

Round-robin is like a host seating customers at whichever table is next in rotation, even if the last party at that table ordered a giant multi-course meal that will keep the table busy for an hour. Least-connections is a host who actually looks at which tables are still occupied and seats new customers at the emptiest one.

saying these in an interview costs you the question

  • Says load balancers only exist for scaling and never mentions failover/availability
  • Thinks round-robin adapts to server load
  • Can't explain why request-cost variance matters to the algorithm choice
  • Confuses least-connections with least-response-time
  • Assumes least-connections requires no additional state or coordination

context

open as a page

What is the difference between an L4 (transport-layer) and an L7 (application-layer) load balancer, and what routing or scaling decisions can an L7 balancer make that an L4 balancer cannot?

level: middleimportance: must knowfreq 80%

basics

~20 s

An L4 load balancer only looks at IP addresses and TCP/UDP ports and just forwards packets, without knowing what's inside. An L7 load balancer reads the actual HTTP request (URL, headers, cookies) and can route based on that content, but doing so costs more CPU and adds latency.

open as a page

What is session affinity (sticky sessions) in a load balancer, how is it typically implemented, and what breaks when the backend a client is pinned to becomes unavailable?

level: middleimportance: must knowfreq 70%

basics

~20 s

Session affinity means the load balancer sends all of one client's requests to the same backend server, usually to keep server-stored session data (like a login session) working. It's typically done with a cookie identifying the server. If that server goes down, the client loses that session state unless it's shared elsewhere.

open as a page

How do active and passive health checks work in a load balancer, and what production failure modes can occur if health-check thresholds are configured too aggressively or too loosely?

level: seniorimportance: must knowfreq 75%

basics

~20 s

Active health checks are the load balancer regularly pinging each server with a test request to see if it's alive. Passive health checks watch real traffic and mark a server unhealthy if its actual responses start failing. Bad thresholds can either kick out healthy servers too fast or take too long to notice a broken one.

open as a page

Why is consistent hashing used as a load-distribution strategy instead of plain modulo hashing, and what problem does it specifically solve when backend servers are added or removed?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Plain modulo hashing (key % number-of-servers) reshuffles almost every key's assigned server whenever a server is added or removed, which is disastrous for caches. Consistent hashing arranges servers and keys on a ring so that adding or removing one server only reassigns a small fraction of keys, not nearly all of them.

open as a page

When a fleet behind a load balancer autoscales horizontally in and out with traffic, what mechanisms keep the load balancer and backend pool consistent, and what can go wrong if connection draining or backend registration is handled naively?

level: principalimportance: should knowfreq 55%

basics

~20 s

When autoscaling adds or removes servers, the load balancer needs to know about it and give departing servers time to finish in-flight requests instead of cutting them off mid-response. Doing this poorly causes dropped requests during every scale-in and traffic slamming into unready new servers during every scale-out.

open as a page