When a fleet behind a load balancer autoscales horizontally in and out with traffic, what mechanisms keep the load balancer and backend pool consistent, and what can go wrong if connection draining or backend registration is handled naively?
answer
- registration gated on health check, not just process start
- slow start ramps a cold new instance
- connection draining = deregistration delay before kill
- naive scale-in cuts in-flight requests
- reactive autoscaling can thrash; balancer itself needs headroom
basics
~20 sWhen autoscaling adds or removes servers, the load balancer needs to know about it and give departing servers time to finish in-flight requests instead of cutting them off mid-response. Doing this poorly causes dropped requests during every scale-in and traffic slamming into unready new servers during every scale-out.
solid answer
~60 sHorizontal scaling behind a load balancer requires two coordinated mechanisms: registration on scale-out (a new instance must pass its health checks and be added to the pool before receiving traffic, and often needs 'slow start' to ramp up gradually rather than being hit at full share immediately) and connection draining on scale-in (a terminating instance is first removed from new-request routing but allowed a grace period, often called a deregistration delay, to finish requests already in flight before the process is actually killed). Getting this wrong causes concrete production symptoms: without slow start, a fresh instance with cold caches and un-JIT'd code gets its full share of traffic immediately and can itself trigger a health-check failure or high error rate; without connection draining, scale-in events (very common with reactive autoscaling policies) abruptly kill instances mid-request, causing a spike of client-visible errors or resets timed exactly with every scale-down. At larger scale this also intersects with autoscaling policy tuning itself - reactive threshold-based scaling can thrash (rapid scale-out/scale-in cycles) under bursty traffic, and the load balancer's own capacity (e.g. its pre-warming for expected traffic spikes) must be considered, since the balancer itself is not infinitely elastic instantaneously.
go deeper
Should have a basic sense that new servers need to be added to rotation and old ones shouldn't just vanish mid-request.
Should name connection draining and describe why abrupt termination causes visible errors during scale-in.
Should explain slow start for scale-out and connection draining for scale-in as a coordinated pair, with concrete timeout trade-offs.
Should reason about the interaction between autoscaling policy stability (thrashing, cooldowns) and the registration/draining machinery's overhead, plus the load balancer's own capacity limits during extreme traffic events.
## The two moments that need coordination Horizontal scaling means changing the number of backend instances behind a load balancer in response to load - adding instances under an autoscaling group, Kubernetes HorizontalPodAutoscaler, or similar mechanism when demand rises, and removing them when demand falls to save cost. This sounds simple in principle - just change how many servers are in the pool - but making it safe in practice requires the load balancer and the orchestration layer to coordinate carefully around two moments: 1. an instance joining the pool (**scale-out**); 2. an instance leaving it (**scale-in**). ## Scale-out: health-gated registration is not warm-up On scale-out, a newly launched instance is not immediately ready to receive production traffic even once its process has started: it needs to pass its health checks first (see the separate discussion of active/passive health checking), and orchestration systems correctly gate registration on that. - Kubernetes will not route Service traffic to a Pod until its readiness probe succeeds. - Cloud auto-scaling groups similarly wait for target-group health checks to pass before counting an instance as InService. But passing a health check only proves the process can respond to a shallow or deep probe; it does not mean the instance is warmed up in the sense that matters for real traffic. All of these take some time under real load to reach steady-state performance: - JIT-compiled hot paths, - populated in-memory caches, - established connection pools to the database. If the load balancer immediately gives this brand-new instance its full statistical share of traffic (which, as covered under least-connections, it's especially prone to do since a fresh instance starts with the lowest in-flight count), the instance can be hit disproportionately hard while still cold, producing elevated latency or even errors that can, ironically, cause it to fail its own health check and get pulled right back out - a self-defeating oscillation. The standard fix is **'slow start'** or 'connection ramp-up': the load balancer artificially caps how much traffic a newly healthy instance receives for a configured warm-up period (e.g. linearly ramping from 10% to 100% of its fair share over 60-300 seconds), giving it time to warm up under a gentler load before facing full production traffic. ## Scale-in: connection draining On scale-in, the naive failure mode is worse and more visible to users: if an instance is simply terminated the moment the autoscaler decides to remove it, any request already in flight to that instance - a request the load balancer sent it milliseconds before termination, or a longer-running request still being processed - gets abruptly cut off, typically surfacing to the client as a connection reset or a 502/504 error. Because reactive autoscaling (scaling based on current CPU/request-rate thresholds) by nature triggers scale-in events fairly often as load fluctuates, a naive setup turns every single scale-in into a small burst of client-visible errors, which is a real and recurring reliability cost, not a rare edge case. The fix is **connection draining**, sometimes called deregistration delay: when an instance is marked for removal, the load balancer first stops sending it any new requests (removes it from the routing pool immediately) but does not kill the underlying process yet; instead it waits a grace period - AWS ALB defaults to 300 seconds, tunable - during which any requests already in flight to that instance are allowed to complete normally, and only after the drain period (or once all connections have naturally closed, whichever comes first) is the instance actually terminated. This requires coordination between the orchestrator (which wants to reclaim the instance's compute) and the load balancer (which needs time to safely stop using it), and getting the drain timeout wrong in either direction causes a problem: - **too short**, and legitimately long-running requests still get cut off; - **too long**, and scale-in events take a long time to actually free capacity, which matters if you're scaling in specifically to relieve resource pressure elsewhere. ## How autoscaling policy tuning multiplies the cost At larger scale, this whole registration/draining machinery interacts with the tuning of the autoscaling policy itself. Reactive, threshold-based autoscaling (scale out when CPU > 70%, scale in when CPU < 30%) is prone to **thrashing** under bursty or noisy traffic - the fleet scales out, the new capacity brings average CPU back down, which triggers scale-in, which then raises CPU again, cycling repeatedly - and every cycle incurs the cold-start cost of slow-starting new instances and the churn cost of draining old ones, so instability in the autoscaling policy directly multiplies the operational cost of the registration/draining machinery. Cooldown periods (a minimum time between consecutive scaling actions) and using a smoothed or predictive metric rather than instantaneous CPU are standard mitigations. ## The load balancer is not instantly elastic either Finally, it's worth remembering the load balancer itself is not infinitely and instantaneously elastic either - a sudden, very large traffic spike can require the load balancer's own capacity (e.g. an ALB's internal node count, or an NGINX fleet's own instance count) to scale up, which for some managed load balancers involves **pre-warming**: providers like AWS explicitly recommend contacting support ahead of an expected massive traffic event (a product launch, a major sale) so the ALB's own infrastructure is provisioned ahead of time, because the load balancer scaling reactively to sudden load is itself a real, bounded process, not an instantaneous one.
- Why can't a load balancer just rely on the readiness/health check alone to decide when a new instance is fully ready for production-level traffic share?A health check typically only verifies the process can respond correctly to a probe, which is a binary readiness signal, not a measure of performance readiness - it says nothing about whether caches are warm, connection pools are established, or JIT compilation has kicked in for hot code paths. An instance can pass its health check the moment it starts while still performing far worse than steady-state under real load, which is exactly why slow-start ramping is a separate mechanism layered on top of health-check-gated registration.
- What is a concrete symptom in monitoring that suggests connection draining is misconfigured or missing?A recurring pattern where 502/504 errors or connection resets spike in tight correlation with autoscaling scale-in events (visible by overlaying the autoscaling group's activity history with the error-rate graph), rather than being randomly distributed - that correlation is the signature of in-flight requests being cut off by abrupt instance termination rather than a general reliability issue.
- How does thrashing in a reactive autoscaling policy make the registration/draining overhead worse, not just the scaling itself?Every scale-out event pays the cold-start cost of slow-starting a new instance before it's fully useful, and every scale-in event pays the drain-delay cost of waiting for in-flight requests before capacity is actually freed; if the policy oscillates rapidly between scaling out and in, the fleet spends a disproportionate amount of time with instances that are either still warming up or already draining, rather than serving at full efficiency, which can make the fleet's effective capacity noticeably worse than its instance count would suggest.
Scaling out a fleet without slow start is like opening a new checkout lane at a packed store and immediately sending it the longest queue before the cashier has even found the barcode scanner. Scaling in without connection draining is like flipping the lane's light off and walking the cashier away mid-transaction, leaving the customer's cart half-scanned.
saying these in an interview costs you the question
- Thinks a passed health check means an instance is fully warmed up for production traffic
- Doesn't know what connection draining/deregistration delay is
- Assumes scale-in can just terminate instances immediately with no consequence
- Can't explain why reactive autoscaling can thrash
- Believes the load balancer itself has infinite, instant capacity regardless of traffic spikes