skip to content

Every rolling deploy behind an Application Load Balancer produces a burst of 5xx responses and reset connections as old instances are replaced. Which target-group mechanism is supposed to prevent that, and how do you make it work?

level: seniorimportance: should knowfreq 52%

answer

  1. draining, not instant removal
  2. deregistration_delay.timeout_seconds
  3. order: deregister, wait, then stop
  4. lifecycle hook holds the instance
  5. long-lived connections get cut anyway

basics

~20 s

Connection draining, controlled by the target group attribute deregistration_delay.timeout_seconds. On deregistration the target enters the draining state and stops receiving new requests while in-flight ones finish. Errors usually mean the process is killed before draining completes, not that the delay is wrong.

solid answer

~50 s

The mechanism is connection draining, configured per target group by `deregistration_delay.timeout_seconds` (default 300 seconds). When a target is deregistered it moves to the `draining` state: the load balancer sends it no new requests but lets in-flight ones finish, and only closes what remains when the delay expires. Bursts of 5xx almost always mean the sequencing is wrong rather than the value — the instance or task is stopped at the moment of deregistration instead of after draining. Fix it by ordering the shutdown: deregister (or let an Auto Scaling lifecycle hook hold the instance in `Terminating:Wait`), wait out the drain, then send the process its termination signal, and make the application shut down gracefully. Then tune the delay to real request duration; 300 seconds makes every deploy crawl, and a few seconds is plenty for short requests but far too short for long-lived streams, which the load balancer simply cuts when the timer expires.

code

bash · 4 lines
bash
aws elbv2 modify-target-group-attributes \
  --target-group-arn "$TG_ARN" \
  --attributes Key=deregistration_delay.timeout_seconds,Value=30 \
               Key=slow_start.duration_seconds,Value=60

go deeper

for a junior

Know the term: deregistration delay, also called connection draining, keeps a removed target serving requests already in flight while new ones stop arriving.

for a middle

Explain the draining state, the deregistration_delay.timeout_seconds attribute and its default, and why an in-flight request survives deregistration but not the expiry of the timer.

for a senior

Demonstrate that the burst of errors is a sequencing bug: deregister, drain, then terminate — enforced with an Auto Scaling lifecycle hook or a container stop timeout — plus graceful shutdown in the application and a delay derived from p99 request duration.

for a principal

Own the deployment-safety contract end to end: drain windows versus deploy duration at fleet scale, minimum healthy capacity during a rolling update, warm-up via slow start, and an explicit stance that long-lived connections require client reconnect rather than an ever-longer delay.

## What draining actually is Deregistering a target is not instantaneous by design. The moment the deregistration is accepted, the target's state becomes `draining` (`TargetHealth.Reason` = `Target.DeregistrationInProgress`) and two things are true at once: **no new requests are routed to it**, and **requests already in flight are allowed to complete**. When `deregistration_delay.timeout_seconds` elapses, the load balancer closes whatever is still open and the target leaves the target group. The attribute is per target group, defaults to 300 seconds, and accepts values from 0 to 3600. ```bash aws elbv2 modify-target-group-attributes --target-group-arn "$TG_ARN" \ --attributes Key=deregistration_delay.timeout_seconds,Value=30 ``` ## Why deploys still throw errors Draining only helps if nothing kills the backend during the drain. The failure modes, roughly in order of how often they are the cause: **1. Termination races deregistration.** The deployment tool stops the process, the container, or the instance at the same moment it deregisters. The load balancer dutifully drains a backend that is already gone, and every in-flight request resets. The fix is ordering: deregister first, wait, then terminate. In an Auto Scaling group, an instance-terminating **lifecycle hook** holds the instance in `Terminating:Wait` while the target group drains, and your script completes the hook when the application has shut down. ECS provides the same shape through the container's stop timeout, giving the process a window between `SIGTERM` and `SIGKILL`. **2. The application does not shut down gracefully.** Even with perfect ordering, a process that exits immediately on `SIGTERM` drops its own in-flight requests. Graceful shutdown means: stop accepting new connections, finish what is running, then exit — and the drain window must be at least as long as that takes. **3. Keep-alive connection races.** The ALB holds persistent connections to targets and reuses them. If the target's own keep-alive idle timeout is shorter than the load balancer's, the target can close a connection at the exact moment the load balancer dispatches a request onto it, and the client sees a 502. The standard remedy is to make the application's keep-alive timeout comfortably longer than the load balancer's idle timeout. **4. Nothing was ever healthy to take over.** If new targets are registered but still in `initial` while old ones drain, capacity dips. Minimum-healthy-percentage settings in a rolling update exist to prevent this — and the Auto Scaling **health-check grace period** must be long enough for a booting instance to start passing checks, or the group will kill and replace instances that were merely still starting. ## Choosing the value The delay is a direct deployment-latency cost: with 300 seconds and a one-at-a-time rolling update, twenty instances take at least an hour and a half. Set it from the p99 (not the mean) duration of a real request, plus your application's graceful-shutdown window, plus a small margin. Short JSON APIs are usually fine in the tens of seconds. Endpoints that stream large downloads, hold WebSocket connections, or run long-poll requests are the ones to think about, because when the timer expires the load balancer **closes those connections regardless** — draining does not wait forever, and clients of long-lived connections need reconnect logic no matter what value you pick. ## The NLB variant A Network Load Balancer drains differently because it forwards flows rather than requests. By default, established flows to a deregistering target are *not* torn down when the delay expires — the target group attribute `deregistration_delay.connection_termination.enabled` (default `false`) controls whether the NLB terminates them. If you need a hard cutover on an NLB, that attribute is the switch, and forgetting it explains connections that survive long after the target was supposedly removed. ## The mirror-image setting The opposite end of the lifecycle is `slow_start.duration_seconds`. Without it, a newly healthy target receives its full share of traffic on the first request — fine for a stateless process, unpleasant for a runtime that has not warmed up or a service with a cold cache. Slow start ramps a new target's share over a configured window (as of 2025, roughly 30 to 900 seconds; `0` disables it), giving it time to warm instead of taking a full share of load, timing out, and failing its health check straight back out of rotation.

  • How would you pick a value for deregistration_delay.timeout_seconds?
    From the p99 duration of a real request plus the application's graceful-shutdown window, plus a margin — not from the 300-second default, which makes rolling deploys crawl. Long-lived connections are a separate matter: they will be cut when the timer expires whatever value you choose, so those clients need reconnect logic rather than a longer delay.
  • An ALB throws sporadic 502s under normal traffic, unrelated to deploys. What target-side setting is a common cause?
    The application's keep-alive idle timeout being shorter than the load balancer's. The ALB reuses persistent connections to targets, so if the target closes one just as a request is dispatched onto it, the client sees a 502. Making the application's keep-alive timeout comfortably longer than the load balancer's idle timeout removes the race.
  • What does slow start do, and when is it worth enabling?
    slow_start.duration_seconds ramps a newly healthy target's share of traffic over a window instead of giving it a full share immediately. It is worth enabling when a fresh process is genuinely slower at first — an unwarmed runtime, or a service with a cold local cache — because otherwise the new target can be overwhelmed, time out, and fail straight back out of rotation.
  • How does draining differ on a Network Load Balancer?
    An NLB forwards flows, and by default it leaves established flows to a deregistering target intact rather than terminating them when the delay expires. The target group attribute deregistration_delay.connection_termination.enabled controls that, and leaving it at its default explains connections that persist long after the target was removed.

saying these in an interview costs you the question

  • Raising the deregistration delay when the process is killed immediately
  • Thinking draining kills in-flight requests instantly
  • Assuming a long delay protects WebSocket connections forever
  • Ignoring graceful shutdown in the application itself
  • Leaving the 300-second default and blaming slow deploys on the load balancer

context