skip to content

You are taking a backend out of a load balancer's pool for maintenance. Why is "stop sending it new requests" not the same thing as draining it, and what does the proxy have to do to remove it without failing work that is already in flight?

level: seniorimportance: should knowfreq 50%

answer

  1. removal is about future routing only
  2. in-flight work is still attached
  3. needs a third state, not "unhealthy"
  4. bound the wait with a deadline
  5. streams never end on their own

basics

~20 s

Marking a backend down only stops new requests. Draining also waits out requests already dispatched and connections already pooled to it, bounded by a deadline after which the proxy force-closes. Long-lived connections never end on their own and must be cut.

solid answer

~50 s

Removing a backend from the rotation is instantaneous for *future* routing decisions and irrelevant to everything already in progress. At the moment you mark it down there are requests mid-flight on it, idle pooled connections the proxy may have already handed a request to, and — if you use affinity — clients that expect to keep landing there. A drain state says "no new requests, but keep the existing ones alive", and it needs three things: an in-flight counter you can watch go to zero, a deadline sized from the p99 request duration so a stuck request cannot block the maintenance forever, and a forced close when the deadline expires. Streaming traffic such as WebSockets or gRPC streams has no natural end, so for those the deadline is not a safety net — it is the mechanism.

go deeper

for a junior

Know that removing a server from a load balancer only stops new requests, and that requests already sent to it are still running and need time to finish.

for a middle

Be able to describe drain as a distinct state with an in-flight counter and a timeout, and explain why a failing health check is a different and blunter instruction.

for a senior

Demonstrate that you size the drain deadline from measured p99 request duration, watch in-flight counts to confirm zero, and know why the proxy must not replay dispatched requests elsewhere.

for a principal

Own drain as a platform guarantee: a standard deadline and shutdown sequence every tier inherits, capacity headroom for the overlap window, and a reconnect contract for streaming clients.

## Two different questions "Is this backend eligible for the next routing decision?" and "is this backend free of work?" are separate questions, and only the first one is answered instantly when you take a member out of the pool. Draining is the process of getting from the first answer to the second without turning in-flight requests into errors. ## What is still attached when you mark it down At the instant the backend leaves the rotation, the proxy is still holding several kinds of attachment to it: - **Requests already dispatched.** A request forwarded three seconds ago is still executing. The response has to come back over the connection it left on. - **Idle pooled connections.** If the proxy pools upstream connections, it holds open sockets to that backend. Worse, there is a small window where a worker selected the backend just before the state change and is about to write a request onto one of those sockets. - **Sticky clients.** With cookie or hash affinity, some clients are pinned to that member. Drain has to decide what happens to them: usually new sessions go elsewhere while existing ones ride out the drain window. - **Streams.** WebSocket upgrades, gRPC server streams, server-sent events and long-polling connections stay open by design. They will not finish because you stopped routing. None of those are addressed by a routing change alone. ## What a real drain does A usable drain has four parts. **A distinct state.** Not "healthy" and not "failed" — a third state meaning *no new work, finish what you have*. This matters because the blunt substitute people reach for, deliberately failing the health check, is a different instruction: many proxies treat a failed member as a candidate for immediate ejection and may reset connections to it rather than let them complete. **An in-flight counter.** You need to see the number of active requests and connections on the draining member fall toward zero; otherwise you are guessing when it is safe to stop the process. Most proxies expose this in their stats or admin interface. **A deadline.** Pick it from the distribution of request durations, not from a round number: roughly p99 request duration plus margin covers ordinary traffic. Without a deadline, one hung request holds the maintenance open indefinitely. HAProxy's `hard-stop-after` is one expression of this idea at the process level. **A forced close at the deadline.** When the timer expires the remaining connections are closed. For streaming traffic this is the normal path, not the exception, and it means clients must be able to reconnect — which is a client-side property you have to have designed for in advance. ## The proxy's own shutdown is the same problem The symmetric case is replacing the proxy process itself during a config reload or a deploy. Graceful shutdown means the process stops accepting new connections, keeps serving established ones, and exits when they finish or the deadline hits. nginx does exactly this on SIGQUIT (old workers linger while new ones take over); HAProxy soft-stops on SIGUSR1. In both cases the old and new processes serve simultaneously for the drain window, which is a capacity fact your sizing has to accommodate. ## Why the load balancer cannot just re-send the work A tempting shortcut: when the backend goes away, replay its in-flight requests on another member. The proxy generally must not. Once the request was forwarded, the backend may have already committed a side effect — charged a card, written a row — and the proxy has no way to know. Replaying is only safe when the request provably never reached the application, which is why proxy-level retry conditions are restricted to connection-establishment failures rather than "anything that did not come back". ## Failure signatures When draining is missing or misconfigured, the symptoms are recognisable: - A burst of 502s at the exact moment of a scale-in or deploy, with a duration equal to the tail of your request-latency distribution. - Errors that stop as soon as you add a manual pause between removing the member and stopping the process — the classic sign that the pause *is* the missing drain. - Streaming clients dropping en masse at deploy time, because nothing ever asked them to reconnect gracefully. ## The one-line version Removal is about the future; drain is about the present. Give the draining member a state of its own, watch its in-flight count, bound the wait with a deadline you derived from real latency data, and accept that anything long-lived will be cut rather than finished.

  • How would you choose the drain deadline?
    From the request-duration distribution of that service, not a default. Roughly p99 plus a margin covers ordinary requests; anything beyond that is a stuck request you do not want to wait for. Then sanity-check it against the deploy process: if the deadline exceeds the time your orchestration allows before killing the process, the deadline is fiction and requests will be cut anyway.
  • What happens to sticky sessions during a drain?
    New sessions must be steered to other members while existing pinned sessions ride out the drain window. That is the point of a separate drain state — affinity would otherwise keep routing returning clients to a member you are trying to empty. When the deadline expires the pinned clients get rebalanced like everyone else, which is exactly the moment affinity-dependent state shows up as broken.
  • Why can't the load balancer just retry the in-flight requests on another backend?
    Because it cannot tell whether the backend already applied the request's effect before it went away. Once bytes were forwarded, a charge may already be committed. Replay is only defensible for failures that provably happened before the application saw the request — a refused connection, for instance — which is why proxy retry conditions are deliberately narrow.

saying these in an interview costs you the question

  • Assumes taking a backend out of rotation ends its work
  • Uses a failing health check as a drain signal
  • Drains with no deadline, so one hung request blocks maintenance
  • Expects the proxy to replay in-flight requests on another backend
  • Forgets that WebSocket and gRPC streams must be force-closed

context