Why can a terminating Kubernetes pod, already not ready in its Service's EndpointSlice, keep receiving requests over keep-alive connections, and how do you drain them?
answer
- new connections versus new requests
- conntrack pins established TCP flows
- pooled upstream connections
- Connection: close and GOAWAY
- proxy idle timeout below app's
basics
~20 sEndpointSlice readiness only steers new connections; established keep-alive, HTTP/2 and gRPC connections keep reaching the same pod. The application must drain them after SIGTERM: stop accepting, signal close or GOAWAY, finish in-flight work, then exit.
solid answer
~40 sMarking the endpoint `ready: false` changes where **new** connections go. kube-proxy picks a backend when a connection opens and conntrack keeps established TCP flows pinned to that pod; it only clears stale conntrack entries for UDP. An ingress controller may still reuse pooled connections to the pod, and client libraries keep their connections to the Service IP. So the pod keeps getting requests until those connections close. To drain: keep serving through a `preStop` delay, then on SIGTERM stop accepting, answer with `Connection: close` or send HTTP/2 `GOAWAY`, finish in-flight requests, close idle connections and exit, all within `terminationGracePeriodSeconds`. Also make the proxy's upstream idle timeout shorter than the app's keep-alive timeout, and make requests idempotent so retries are safe.
go deeper
Remember that readiness changes affect new connections only, and that a pod must close its existing connections itself before it exits.
Explain conntrack pinning, pooled proxy connections and client channels, and walk through Connection: close, GOAWAY and the idle-close race.
Show how you would set proxy idle timeouts, server connection age and idempotent retries so a rolling update drops nothing, and how you would prove it with rollout metrics.
Weigh long-lived connections against rollout speed and load balance: shorter connection lifetimes spread load and ease drains but add handshakes and latency across the fleet.
## The short version Marking an endpoint not ready in an **EndpointSlice** changes where **new** connections go. It does nothing to a TCP connection that is already open. HTTP keep-alive, HTTP/2 and gRPC clients reuse one connection for many requests, so a pod that no consumer will pick for a new connection can keep receiving requests over the connections it already has. Draining them is the application's job. ## Why the endpoint change does not reach open connections When the loyalty-points accrual pod gets a `deletionTimestamp`, the EndpointSlice controller in `kube-controller-manager` sets its endpoint to `ready: false` and `terminating: true`. Each consumer then reacts in its own way: - **kube-proxy** rewrites the Service rules on every node. In iptables and nftables mode, the Service rules pick a backend when a connection is first opened. After that, the kernel's **conntrack** entry keeps sending the connection's packets to the same pod IP. kube-proxy clears stale conntrack entries for **UDP** Services only; established TCP flows keep their pinned pod. - **An ingress controller** that proxies straight to pod IPs read from the EndpointSlice stops choosing the pod for new upstream connections, but it may still hold idle pooled connections to it and send the next request down one of them. - **Client-side pools** in other services, such as HTTP clients or gRPC channels, connect to the Service's virtual IP once and then reuse that connection. As far as they know, nothing has changed. So "removed from the ready set" means **no new connections**, not **no new requests**. ## Where the requests fail A pod that receives requests over a live connection is not itself a problem. Failures come from how that connection ends: 1. The process exits, or is killed after the grace period, while connections are still open. Clients see a connection reset, or a proxy returns a 502. 2. The server closes an idle keep-alive connection at the same moment a client sends a new request on it. The request is lost and, if it is not idempotent, cannot be retried safely. A loyalty-points **accrual** is exactly that kind of request: retrying it blindly could credit points twice. 3. The pod stops accepting connections too early, before every consumer has stopped choosing it, and new connections are refused. ## How to drain connections properly The pattern that works has three phases, and all of them must fit inside the pod's `terminationGracePeriodSeconds`: | Phase | What the pod does | Why | |---|---|---| | Delay | keeps serving normally (for example a `preStop` sleep) | lets every consumer stop opening new connections | | Drain | stops accepting new connections, finishes in-flight requests, closes idle ones | ends existing connections cleanly | | Exit | exits once the connection count reaches zero | frees the pod before the kubelet's SIGKILL | In the **drain** phase, the server should tell clients to go away instead of just dropping the socket: - For **HTTP/1.1**, send `Connection: close` on the next response on each connection, so the client opens a fresh connection that now lands on a ready pod. - For **HTTP/2 and gRPC**, send a `GOAWAY` frame. The client finishes its current streams and opens new ones elsewhere. - Close connections that are **idle** from the server side. This leaves a small window where a client's request can race the close, so the client or proxy should retry idempotent requests on a connection that was just closed. Most web frameworks and gRPC servers have a graceful-shutdown call that does most of this; the work is to call it on SIGTERM rather than exit. ## Settings on the other side The application cannot fix everything alone: - Keep the **proxy's idle timeout for upstream connections shorter** than the application's own keep-alive timeout, so the proxy, not the app, is the side that closes idle connections. - Give long-lived gRPC clients a **maximum connection age** on the server side, so connections are recycled over time rather than living until the pod dies. - Make accrual requests **idempotent**, for example with a request ID, so a retry after a reset is safe. ## A worked example The loyalty-points accrual Deployment runs behind a Service and an ingress controller. During a `RollingUpdate`, old pod `points-accrual-7c9f` is marked terminating. Its preStop sleep runs for 25 seconds while ingress and kube-proxy update. SIGTERM then arrives. The server stops listening, answers its next request on each HTTP/1.1 connection with `Connection: close`, sends `GOAWAY` on HTTP/2 connections, waits up to 8.2 seconds for in-flight accruals, and exits. No connection is still open when the process ends.
- Why should an ingress controller's upstream idle timeout be shorter than the application's keep-alive timeout?If the application closes an idle connection first, the proxy may send a request on it at the same moment and get a reset, which surfaces as a 502. When the proxy's idle timeout is shorter, the proxy is always the side that retires idle connections, so it never picks a connection the server has just closed.
- Does kube-proxy close existing TCP connections to a pod when its endpoint becomes not ready?No. kube-proxy updates the rules that choose a backend for new connections. Established TCP connections keep following their conntrack entry to the original pod. kube-proxy's stale-conntrack cleanup applies to UDP Services only. The connections end when the client or the pod closes them, or when the pod process dies.
Closing a shop's front door stops new customers, but the people already inside keep shopping; someone still has to announce closing time and walk them out before the lights go off.
saying these in an interview costs you the question
- Once the endpoint is not ready, kube-proxy cuts that pod's TCP connections.
- A preStop sleep alone drains keep-alive connections.
- Removing a pod from the ready set stops all requests to it.
- Closing idle keep-alive connections from the server is always safe.
- gRPC clients rebalance on their own when an endpoint turns not ready.