A backend instance that a load balancer had ejected is now passing its health checks again and is about to rejoin the pool. Why can handing it an immediate equal share of traffic break it a second time, and what does a slow-start or warm-up ramp do about that?
answer
- passing a probe is not being ready
- cold caches, empty pools, unoptimised code
- an idle host looks most attractive
- ramp the weight, not the health flag
- same problem for freshly started instances
basics
~20 sA returning instance is cold: empty caches, no warm connection pools, unoptimised runtime. Worse, load-aware algorithms see an idle host as the most attractive target and send it a burst above its fair share. A slow-start ramp raises its weight gradually so it warms under partial load.
solid answer
~60 sPassing a health check proves an instance can answer a cheap probe; it does not prove the instance can absorb a full share of production traffic. A host that has been out of rotation has cold application and local caches, an empty or drained connection pool to its own dependencies, and — on runtimes that optimise at run time — code that has fallen back to an unoptimised state. Readmitting it at equal weight sends it a full share into all of that at once, and it is worse than equal share with load-aware balancing: an idle host has the fewest active requests and the lowest observed latency, so least-request and latency-weighted algorithms actively prefer it and pile on a burst. It slows down, fails, and is ejected again — the flap. Slow start fixes it by ramping the host's effective weight from near zero to full over a window, so it warms up under a fraction of the load. Two companions matter: growing the ejection duration on repeated ejection, and remembering that the same problem applies to any freshly started instance, not only a returning one.
go deeper
Understand that a backend just back in service is cold — its caches and connections are empty — so its first requests cost far more than normal even though its health check passes.
Explain the mechanism of slow start: the balancer ramps the host's weight over a window instead of flipping it to full share, and explain what is warming during that window.
Bring up the load-aware balancing interaction unprompted, size the ramp from measurement, and connect it to escalating ejection duration and to cold instances from deploys and autoscaling.
Decide where warm-up belongs across the platform — proxy-side ramping, application-side pre-warming at startup, or gating readiness until warm — and what you standardise so every team is not solving cold start differently.
## Passing a probe is not the same as being ready for load The health check answers a cheap question: can you respond? Production traffic asks an expensive one: can you respond to a few hundred requests per second right now? Between the two sits everything that a running instance accumulates and that an idle one has lost. **What is cold on a returning instance:** - **Application-level caches** — an in-process cache that normally absorbs most reads is empty, so nearly every request goes to the database. The instance's per-request cost is several times its steady-state cost, and its dependency's load spikes at the same time. - **Connection pools.** Pools to databases, caches and downstream services have been idle-reaped or were never established. The first burst of requests all try to open connections simultaneously, each paying a handshake — and TLS handshakes are expensive enough to matter — while queueing behind a pool whose maximum size throttles them. - **Runtime optimisation state.** On runtimes that compile or optimise while running, code that has not executed recently runs in a slower mode until it is re-optimised, and the optimisation itself costs CPU concurrently with the traffic that triggered it. - **Anything lazily initialised** — class loading, schema or metadata fetches, first-call configuration, a client library that resolves discovery on first use. Each of these makes the first minute of traffic far more expensive per request than the steady state. ## The load-aware balancing trap Here is the part that surprises people. You might expect a returning instance to receive its fair share — one of N. With a load-aware balancing algorithm it receives *more*. Algorithms that pick the backend with the fewest outstanding requests, or the best recently-observed latency, evaluate an idle host as the most attractive host in the pool: zero active requests, no recent slow responses. So they steer a disproportionate burst of new requests to precisely the instance least able to handle them, and they keep doing so until its latency and queue rise enough to make it look ordinary. This is a known, ugly interaction, and it is the strongest single argument for a ramp: the balancer's own heuristics need to be overridden while a host is cold. ## What slow start actually does Slow start (also called warm-up) makes the host's effective weight a function of how long it has been in the pool: near zero at readmission, rising — usually linearly — to its configured weight over a window of tens of seconds to a few minutes. Two properties matter: - It is **time-based, not request-based**, which is what you want: caches and pools warm as a function of both traffic and elapsed time, and a time ramp is predictable and cheap to reason about. - It **overrides the load-aware preference** during the window, so the idle-host burst cannot happen. Sizing the window is empirical: measure how long a cold instance takes to reach steady-state latency and set it there. A JVM service with large caches may need minutes; a small stateless process may need seconds and gain nothing from a ramp at all. ## The two companions **Escalating ejection duration.** If an instance is ejected, readmitted, and ejected again, readmitting it just as quickly the second time produces a flap that perturbs the whole pool. Growing the exclusion period with each consecutive ejection — the ejection duration multiplied by the ejection count is the usual shape — turns a flapping host into a quietly excluded one, and it recovers automatically once the host stays healthy. **The same problem on cold starts.** Nothing about slow start is specific to ejection. A newly started instance, an instance added by autoscaling, and an instance replaced during a rolling update are all cold in exactly the same way. If a deploy shows a latency spike on every new instance, the cause is usually this, and the ramp is the same fix. ## The failure mode you are preventing Without it, the sequence is: eject, recover, readmit at full weight, receive an outsized burst because it looks idle, fall over under cold-start cost, get ejected again — while each cycle redistributes load across the pool and clears whatever locality the balancer had. Slow start converts an oscillation into a monotone recovery, which is why it is a small feature that shows up disproportionately in incident write-ups.
- Why does a least-request or latency-aware balancer make readmission worse rather than better?Because it optimises on exactly the metrics an idle host trivially wins. Zero in-flight requests and no recent slow responses make the cold instance look like the best backend available, so it receives a burst well above a fair share at the moment it can least absorb one. The heuristic is right in steady state and wrong for a host with no history.
- How would you choose the length of the warm-up window?Measure it rather than guess. Restart one instance under representative load and watch how long its latency, cache hit rate and CPU take to reach the levels of its peers; set the window to about that. Services with large in-process caches or heavy runtime optimisation can need minutes, while a small stateless process may need almost nothing and gains only complexity from a ramp.
- What should happen if the same instance is ejected repeatedly?Extend the exclusion each time — the usual shape multiplies the base ejection duration by the number of consecutive ejections, so a persistently bad host is retried at widening intervals instead of flapping every few seconds. It self-heals: once the host stays healthy through a full period, the counter resets and it returns to normal rotation without an operator.
saying these in an interview costs you the question
- Assumes a passing health check means full-load readiness
- Expects a returning instance to get exactly one Nth of traffic
- Thinks slow start changes the health status rather than weight
- Readmits a repeatedly flapping instance at the same interval
- Believes only ejected instances are cold, not newly started ones