skip to content

How does an external load balancer's health check, such as an AWS ALB target group health check, differ semantically from a Kubernetes readiness probe, and what production problem can arise from treating them as interchangeable during a rolling deploy or scale-down?

level: seniorimportance: must knowfreq 60%

answer

  1. two independent control planes, no handshake
  2. internal Endpoints vs external target group
  3. preStop sleep bridges the detection gap
  4. terminationGracePeriodSeconds must cover drain time
  5. L4 TCP checks vs L7 HTTP checks can disagree

basics

~20 s

The load balancer's health check and Kubernetes' internal readiness check run on separate schedules and separate systems, so there's a lag between Kubernetes deciding a pod is gone and the load balancer noticing. If a pod shuts down before the load balancer catches up, some requests get sent to a pod that no longer exists, causing failed requests during deploys or scale-downs.

solid answer

~50 s

A Kubernetes readiness probe controls membership in the Service's internal Endpoints list, used by kube-proxy or the CNI for in-cluster routing. An external load balancer's health check is a completely separate system polling the pod directly or via a node port, on its own interval, deciding target-group membership independently. These two control planes aren't synchronized: when a pod is deleted, Kubernetes sends SIGTERM almost immediately, but the load balancer may keep sending it traffic for several more of its own check intervals, often ten to thirty-plus seconds, until it notices via its own failing checks. Treating them as interchangeable causes dropped connections when the pod terminates before the external load balancer has deregistered it. The fix is a preStop hook that sleeps long enough for the slower external check to catch up, while the app keeps draining in-flight connections during that window.

go deeper

for a junior

Should recognize there are two separate systems involved, the cluster and the external load balancer, even without precise terminology.

for a middle

Should know that a preStop hook or sleep is the standard mitigation and that termination isn't instantaneous from the outside world's perspective.

for a senior

Should explain the mechanism precisely — internal Endpoints update vs external polling interval — and correctly size terminationGracePeriodSeconds against preStop sleep and deregistration delay.

for a principal

Should generalize this as a distributed-control-plane synchronization problem beyond Kubernetes/ALB specifically, and weigh cost/complexity trade-offs of tuning drain windows at fleet scale.

## Two control planes, one pod A Kubernetes readiness probe governs one specific, internal thing: membership in the Service's `Endpoints` or `EndpointSlice` object, which `kube-proxy` (or an Ingress controller watching Endpoints directly) uses to build the routing rules that direct traffic within the cluster. An external load balancer sitting in front of the cluster — an AWS ALB or NLB target group, a hardware load balancer, or any self-managed edge proxy — runs a completely separate health-check system with its own configuration: - a check path; - a polling interval; - healthy and unhealthy thresholds; - and a deregistration delay, all independent of anything Kubernetes knows about. Each system polls or watches on its own schedule and has no built-in awareness of the other's current state. ## Why they are separate This separation exists because the two systems solve different layers of the same problem. The orchestrator manages pod scheduling, restarts, and in-cluster routing; the external load balancer handles ingress from outside the cluster, often across multiple nodes, availability zones, or even clusters, and frequently also does TLS termination and cross-region traffic distribution. They are typically built by different teams or vendors for different scopes, so there is no automatic handshake between 'Kubernetes has decided this pod is gone' and 'the external load balancer has stopped sending it traffic.' ## What goes wrong during termination The practical problem surfaces during any pod termination — a rolling update, a scale-down, or a node drain. | Side | What it does at termination | |---|---| | **Kubernetes** | Marks the pod Terminating and removes it from the internal `Endpoints` list essentially immediately, while concurrently sending `SIGTERM` to the container; if the process exits quickly, it's fully gone within a second or two. | | **The external load balancer** | However, is still relying on the result of its last health check, which can be stale by up to its interval times its unhealthy threshold — for a typical ALB configuration of a 30-second interval and 2 consecutive failures to mark unhealthy, that's up to roughly 60 seconds of lag. | During that window it keeps routing new connections to a target that no longer exists, producing connection-refused errors or 502s at the edge. ## The standard mitigation The standard mitigation bridges that detection gap deliberately: a `preStop` hook, typically just a sleep for some duration matched to the external check's worst-case detection time, delays the actual SIGTERM-driven shutdown. During the sleep, the pod has already been removed from Kubernetes' own internal routing (so in-cluster traffic stops immediately), but the process keeps its listening socket open and continues answering any requests still arriving from the load balancer, which hasn't caught up yet. Only after the preStop sleep elapses does the container actually begin shutting down, so `terminationGracePeriodSeconds` must be sized large enough to cover the preStop sleep plus whatever real shutdown work — like finishing in-flight requests — the app still needs to do; if it's too short, Kubernetes SIGKILLs the process before draining completes, defeating the entire purpose. ## Related failure modes Several related failure modes show up in production around this gap. 1. **An under-provisioned `terminationGracePeriodSeconds`** relative to the preStop sleep causes exactly the premature-kill problem the preStop hook was meant to solve. 2. **A deregistration delay on the load-balancer side that's mismatched to the app's real drain time** causes either dropped connections, if too short, or wasted resources, if a pod lingers serving nothing new purely to satisfy an overly conservative delay. 3. **A subtler semantic mismatch.** There's also a subtler semantic mismatch: many external load balancers default to L4 (TCP) health checks, which only confirm a port accepts a handshake, versus Kubernetes readiness probes, which are almost always L7 (HTTP) and inspect an actual response. A process that's hung inside its request-handling logic but whose listening socket is still open can pass an L4 check indefinitely while correctly failing an HTTP-based readiness probe — the two systems can disagree even outside of any transition, not just during shutdown. ## A concrete pattern A concrete, well-documented pattern involves a service behind an AWS ALB in EKS with a typical HTTP health check on a 30-second interval requiring two consecutive failures, giving up to roughly 60 seconds of detection lag. During a rolling deploy without a preStop hook, old pods receive SIGTERM the instant Kubernetes replaces them, but the ALB keeps routing to them for up to a minute, producing a visible burst of 502s in the ALB's access logs that correlates exactly with deploy timestamps. The standard remediation, consistent with AWS's own EKS best-practice guidance, is a preStop hook sleeping roughly as long as the target group's deregistration delay, combined with enabling connection draining, so old targets keep answering both in-flight and newly arriving but stale traffic gracefully until the load balancer has fully deregistered them.

  • Why does removing a preStop hook and relying only on SIGTERM handling inside the app usually still leave a gap with an external cloud load balancer?
    SIGTERM handling inside the app can gracefully stop accepting new local connections and finish in-flight requests almost immediately, but that's purely an internal concern — it tells the external load balancer nothing. The load balancer only learns the target is gone through its own health-check polling cycle, which runs on its own timer independent of the app's shutdown speed, so even a perfectly graceful app can still receive a burst of doomed connections during that detection window.
  • What's the difference between an L4 and an L7 load balancer health check, and why might they disagree with a Kubernetes readiness probe?
    An L4 check only verifies the TCP handshake succeeds; something is listening on the port. An L7 check sends an actual HTTP request and inspects the response status or body. A process hung inside its request-handling logic but whose socket is still listening passes an L4 check while failing an L7 check or a Kubernetes readiness probe, which is almost always HTTP-based, so a load balancer using only L4 checks can keep routing to a functionally dead pod Kubernetes has already correctly marked not ready.

It's like a restaurant hostess instantly crossing a table off her seating chart the moment it closes out, while a third-party reservation app still shows that table as available for another twenty minutes because it only refreshes its listing periodically — a party can get seated at a table that's already gone.

saying these in an interview costs you the question

  • Assumes the external load balancer immediately knows when Kubernetes marks a pod not ready
  • Doesn't mention preStop hooks or connection draining when discussing graceful termination
  • Confuses terminationGracePeriodSeconds with the load balancer's deregistration delay
  • Treats L4 and L7 health checks as equivalent
  • Has no explanation for why rolling deploys produce a burst of errors at the edge specifically

context