During a rolling update of a Kubernetes Deployment, clients see a burst of connection-refused and reset errors even though the app implements graceful shutdown on SIGTERM. Walk through why this happens and how you would eliminate it.
answer
- SIGTERM and endpoint removal race, unordered
- preStop runs before SIGTERM - sleep there
- Grace period > preStop + drain time
- Readiness failure has the same propagation lag
- terminating+serving = availability floor, not the fix
basics
~20 sPod deletion and endpoint removal are concurrent, not ordered. The kubelet can send SIGTERM and the app can stop listening before every node's kube-proxy has removed that pod from its rules, so in-flight traffic still arrives. Fix it by making the container keep serving briefly after deletion starts - a preStop sleep or a delayed shutdown - so proxies converge before the socket closes.
solid answer
~60 sDeleting a pod fans out **in parallel**: the API server marks it terminating, which simultaneously (a) tells the kubelet to run preStop and send SIGTERM and (b) tells the EndpointSlice controller to flip the endpoint, which then has to reach kube-proxy on **every** node and reprogram kernel rules. Nothing sequences those. If the app reacts to SIGTERM by closing its listener immediately, it wins the race and traffic still being DNATed to it gets refused. The reliable fix is to make the container **outlive** the propagation window while still serving: - add a `preStop` hook that sleeps a few seconds (SIGTERM is only sent after preStop returns), or have the app ignore SIGTERM for a short grace window, then drain; - set `terminationGracePeriodSeconds` comfortably larger than preStop sleep plus real drain time; - fail the readiness probe first if you want removal to start sooner, but do not rely on that alone - it travels the same propagation path. EndpointSlice's `terminating` and `serving` conditions let modern kube-proxy keep sending traffic to terminating-but-serving pods when a Service would otherwise have none, which prevents blackholes but does not remove the need for the drain window.
code
yaml · 12 linesspec:
terminationGracePeriodSeconds: 45
containers:
- name: api
image: registry.example.com/api:1.4.2
lifecycle:
preStop:
sleep:
seconds: 8 # keep serving while proxies converge
readinessProbe:
httpGet: { path: /readyz, port: 8080 }
periodSeconds: 2go deeper
Be able to state that pod shutdown and endpoint removal happen at the same time, so a pod must keep serving briefly after being told to stop.
Explain the two concurrent chains, and configure the fix correctly: preStop sleep, grace period sized above preStop plus drain, readiness as a signal not a barrier.
Diagnose it end to end - reproduce with load during a rollout, read EndpointSlice conditions, and cover the second-order cases: keep-alive connections, external LB health-check lag, externalTrafficPolicy Local.
Generalise to eventual consistency in the dataplane: every control-plane-mediated change has propagation delay, so workloads need overlap windows, and this belongs in a platform-wide pod template baseline rather than per-team folklore.
## The race When a pod is deleted (directly, or because a Deployment rollout replaces it), the API server sets `deletionTimestamp`. From that single event, two independent chains start at the same time: **Chain A - stop the workload.** The kubelet watching that pod runs the `preStop` hook if present, then sends SIGTERM to PID 1 in each container, then waits up to `terminationGracePeriodSeconds` before SIGKILL. **Chain B - stop sending it traffic.** The EndpointSlice controller in kube-controller-manager notices the terminating pod and updates the slice. That write goes to etcd, then out over watches to kube-proxy on **every node in the cluster**, plus any ingress controller or mesh control plane. Each kube-proxy then reprograms its kernel rules. Chain A is local and fast - often single-digit milliseconds. Chain B crosses the control plane and fans out to every node; under load or in large clusters it can take hundreds of milliseconds to seconds. There is **no happens-before relationship** between them. A well-behaved app that closes its listening socket the instant it sees SIGTERM therefore stops accepting connections while some nodes are still DNATing new connections to it. Those SYNs hit a closed port and the kernel answers RST - the client sees connection refused, or a reset mid-request. This surprises people precisely because the app *is* doing graceful shutdown. Graceful shutdown handles requests already accepted; it does nothing about connections the dataplane is still about to send. ## Making the pod outlive propagation The cure is to keep the pod accepting traffic for a window longer than worst-case endpoint propagation: 1. **`preStop` sleep.** The kubelet runs preStop *before* SIGTERM, so a sleep-style hook delays the shutdown signal entirely while the app keeps serving normally. Since v1.29 there is a native `sleep` hook action, so you no longer need a shell in the image. This is the standard fix because it requires no application change. 2. **Application-side grace window.** Alternatively, on SIGTERM the app keeps its listener open and keeps passing health checks for a few seconds, then stops accepting, finishes in-flight requests, and exits. This is strictly better for apps you control because it can also flip readiness and close keep-alive connections cleanly with `Connection: close`. 3. **Size the grace period.** `terminationGracePeriodSeconds` (default 30) must exceed preStop sleep plus longest in-flight request, or SIGKILL truncates the drain and you have swapped one error class for another. 4. **Do not rely on readiness alone.** Failing the readiness probe removes the endpoint through exactly the same controller to watch to kube-proxy path, so it has the same propagation delay. It is a useful signal, not a synchronisation primitive. ## What the terminating conditions do and do not do EndpointSlice carries `ready`, `serving`, and `terminating` per endpoint. A terminating pod that still passes its probe is `ready: false, serving: true, terminating: true`. With ProxyTerminatingEndpoints (GA in v1.28) kube-proxy will fall back to terminating-but-serving endpoints when a Service has **no** ready endpoints - most valuable for `externalTrafficPolicy: Local`, where a node might otherwise blackhole traffic during a rollout. That is an availability floor, not a fix for the race: in the normal case, where ready endpoints exist, terminating pods stop receiving new connections and you still need the drain window. ## Other contributors to the same symptom - **Long-lived keep-alive connections.** Even perfect endpoint handling does not move an established connection. The app should send `Connection: close` (or an HTTP/2 GOAWAY) during drain so clients reconnect and get re-balanced. - **External load balancers.** Cloud LBs have their own health-check intervals, typically far slower than in-cluster propagation. Node draining needs the same trick with bigger numbers. - **Rollout pacing.** `maxUnavailable: 0` plus a PodDisruptionBudget keeps capacity up but does not order deletion against endpoint removal - people often reach for these and are puzzled the errors persist. - **Conntrack for UDP.** Stale conntrack entries can pin UDP flows to a dead backend; this is a separate, known failure mode from the TCP race. ## How to verify Run a steady request load during a rollout and count non-2xx responses. Add the preStop sleep and re-run: the errors should go to zero. If they do not, check whether the grace period is truncating drain (look for SIGKILL and exit code 137) and whether clients are holding keep-alive connections.
- Why does setting maxUnavailable to 0 with a PodDisruptionBudget not solve this?Those controls govern how much capacity may be missing at once; they say nothing about the ordering between a pod stopping its listener and every node's proxy learning about it. The rollout still deletes pods, and each deletion still races endpoint propagation. You will keep the fleet at full size and keep seeing per-pod connection errors until you add a drain window.
- Your app already closes its listener on SIGTERM and you cannot change the code. What is the least-invasive fix?Add a preStop hook that sleeps for several seconds. The kubelet runs preStop before delivering SIGTERM, so the container keeps serving normally throughout the sleep while the endpoint removal propagates. Then raise terminationGracePeriodSeconds so the sleep plus the app's own drain still fits inside the grace period before SIGKILL.
- After fixing new-connection errors you still see resets on long-lived HTTP keep-alive connections. Why?Endpoint changes only affect where new connections are DNATed; an established connection stays pinned via conntrack to the pod it was assigned. When that pod finally exits, the connection dies. The pod must actively shed clients during drain - respond with Connection: close on HTTP/1.1 or send an HTTP/2 GOAWAY - so clients reconnect and land on live backends before shutdown completes.
It is like taking your name off a phone directory while calls are already dialling: the directory update takes time to reach everyone, so you must keep answering the phone for a while after you asked to be delisted.
saying these in an interview costs you the question
- Assuming Kubernetes removes the endpoint before sending SIGTERM - the two are concurrent
- Believing a readiness probe failure removes traffic instantly
- Setting a preStop sleep longer than terminationGracePeriodSeconds, so SIGKILL cuts the drain short
- Thinking PodDisruptionBudgets or maxUnavailable ordering fixes the race
- Claiming the terminating and serving conditions make drain windows unnecessary