On a 16-node Kubernetes cluster, how do you decide how long a pod's preStop sleep must be so every endpoint consumer has stopped routing to it, and fit that inside terminationGracePeriodSeconds?
answer
- slowest consumer sets the window
- kube-proxy on every node, not one
- ingress and load balancer paths
- p99 under rollout churn
- sleep plus drain within grace
basics
~20 sMeasure, under rollout load, how long each EndpointSlice consumer takes to stop routing to a deleted pod: kube-proxy on every node, the ingress controller, any cloud load balancer. Sleep longer than the slowest p99, then fit sleep plus drain inside terminationGracePeriodSeconds.
solid answer
~50 sPod deletion starts the kubelet's shutdown and the EndpointSlice update at the same time, and each consumer catches up on its own: kube-proxy on all 16 nodes (rate-limited by `minSyncPeriod`, one second by default), an ingress controller that proxies to pod IPs directly, and an external load balancer that must deregister the target. The window lasts until the **slowest** path has updated, so I measure each path's p99 during real rollouts and set the `preStop` sleep above the worst one with headroom; for our accrual service the load balancer took 14.5 s, so we use 25 s. The hook runs inside `terminationGracePeriodSeconds`, and the app still needs its drain time after SIGTERM: 25 s plus 8.2 s is more than the default 30 s, so we set 40. Also keep the load balancer's deregistration delay consistent with how long the pod can actually serve.
code
bash · 5 lineskubectl get endpointslices -n loyalty \
-l kubernetes.io/service-name=points-accrual -w \
-o custom-columns='VERSION:.metadata.resourceVersion,ENDPOINTS:.endpoints[*].conditions.ready'
kubectl rollout restart deployment/points-accrual -n loyaltygo deeper
Know that the pod keeps serving during the preStop sleep while other components catch up, and that the sleep counts toward terminationGracePeriodSeconds.
List every consumer of the EndpointSlice and explain why each updates on its own timing, including kube-proxy's minSyncPeriod.
Show the measurement plan and the arithmetic: slowest p99 plus headroom for the sleep, then sleep plus drain inside the grace period, rechecked when the ingress or load balancer changes.
Discuss the cost of a long sleep on rollout, scale-down and node-drain speed across the platform, and when it is worth reducing propagation lag at the source instead.
## The problem being budgeted When a pod is deleted, two things start **at the same time**: the kubelet begins shutting the pod down, and the EndpointSlice controller in `kube-controller-manager` marks the pod's endpoint `ready: false`, `terminating: true`. Nothing makes the kubelet wait for the endpoint change to reach anyone. Each consumer of the EndpointSlice updates on its own schedule, and until the **slowest** one has updated, some traffic can still be sent to the pod. A `preStop` sleep keeps the application serving normally through that window. The question is how long the window really is. ## Every consumer that must catch up On the 16-node regulated-workload cluster running the loyalty-points accrual service, the consumers of the Service's EndpointSlice are: - **kube-proxy on all 16 nodes.** Each instance watches EndpointSlices and reprograms its node's Service rules. In iptables and nftables mode its `minSyncPeriod` defaults to one second, so updates are rate-limited, and the watch itself lags when the API server is busy. The window closes only when the **last** node has synced, not the first. - **The ingress controller.** Many controllers read EndpointSlices and proxy straight to pod IPs, bypassing kube-proxy entirely, so their lag is a separate path with its own timing. - **An external cloud load balancer.** If it targets pods directly, a controller must deregister the pod and the load balancer must stop sending new connections to it. That can take much longer than the in-cluster paths, and it is set by the provider's configuration. - **Clients that cache addresses**, such as those using a headless Service's DNS records, which update only after the records change and their cache expires. The controller can also add delay before any of this: `kube-controller-manager` has an `--endpointslice-updates-batch-period` flag that holds pod changes briefly to batch them. ## Measuring instead of guessing Size the sleep from measurements taken during real rollouts, not from defaults: 1. Record when the pod's `deletionTimestamp` was set. 2. Record when the EndpointSlice changed, for example by timestamping each event from a `kubectl get endpointslices -w` stream. 3. Track kube-proxy's sync latency metrics on every node, the ingress controller's own update or reload timing, and the load balancer's target-state changes. 4. Take the **p99 of the slowest path**, measured under the churn of a large rollout, not an idle cluster. ## The arithmetic For the accrual service, measurements over several rollouts gave these p99 figures after `deletionTimestamp`: | Path | p99 lag | |---|---| | EndpointSlice written | 0.4 s | | kube-proxy synced on all 16 nodes | 3.8 s | | ingress controller routing updated | 6.3 s | | cloud load balancer stopped new connections | 14.5 s | The slowest path is 14.5 s. The team set a **25-second preStop budget**, leaving about 10 seconds of headroom for API-server slowness during large rollouts. The sleep does not replace draining. After the preStop hook finishes, the kubelet sends SIGTERM, and the application still needs time to finish in-flight accruals and close keep-alive connections; their p99.9 was 8.2 s. The preStop hook runs **inside** `terminationGracePeriodSeconds`, so the grace period must cover both parts: - 25 s sleep + 8.2 s drain = 33.2 s, which is more than the default of 30 s. With the default, the kubelet would send SIGKILL at 30 s, only 5 s after the sleep ends, cutting the drain short. - The team set `terminationGracePeriodSeconds: 40`, which covers 33.2 s with a few seconds to spare. ```yaml apiVersion: apps/v1 kind: Deployment metadata: name: points-accrual namespace: loyalty spec: replicas: 6 selector: matchLabels: app: points-accrual template: metadata: labels: app: points-accrual spec: terminationGracePeriodSeconds: 40 containers: - name: app image: registry.example.internal/loyalty/points-accrual:4.12.3 lifecycle: preStop: sleep: seconds: 25 ``` ## What the load balancer adds A load balancer's **deregistration delay** (often called connection draining) is how long it keeps **existing** connections to a removed target open. It is not the same as how long it takes to stop sending **new** ones. The preStop sleep must cover the time until new connections stop. The deregistration delay should be no longer than the time the pod will actually remain able to serve, or the load balancer will keep connections open to a pod that has already exited. ## Tradeoffs to state - **A longer sleep slows every rollout**, and it applies to every pod deletion, including scale-down and node drains. - **A sleep that is too long relative to the grace period** leaves no time to drain, and SIGKILL does the rest. - **The sleep is a timer, not a signal.** It does not confirm that consumers have updated, so re-measure after changing the ingress controller, the load balancer setup or the cluster size. - A `sleep` handler needs no shell or `sleep` binary in the image, unlike an `exec` hook.
- Why is sizing the preStop sleep from kube-proxy's sync time alone not enough?kube-proxy is one path among several. Many ingress controllers read EndpointSlices and send traffic straight to pod IPs, and an external load balancer that targets pods has its own deregistration path, often the slowest. The sleep has to cover whichever path is slowest, measured under the load of a large rollout.
- What goes wrong if the preStop sleep is longer than terminationGracePeriodSeconds allows?The preStop hook counts against the grace period. If the sleep uses most of it, the application gets SIGTERM with too little time left to finish in-flight requests, and the kubelet kills it when the grace period ends. With the default 30 seconds and a 25-second sleep, only 5 seconds remain for the drain.
- How does a load balancer's deregistration delay relate to the preStop sleep?They cover different things. The sleep must last until the load balancer stops sending new connections to the pod. The deregistration delay is how long the load balancer keeps existing connections open to it afterwards. That delay should not outlast the time the pod can actually serve, or clients hold connections to a process that has exited.
saying these in an interview costs you the question
- The kubelet waits for the EndpointSlice update before running preStop.
- The preStop sleep runs on top of terminationGracePeriodSeconds, not inside it.
- One kube-proxy sync interval is the whole propagation delay.
- A load balancer's deregistration delay is when it stops new connections.
- A long enough preStop sleep makes application draining unnecessary.
- The grace period default of 30 seconds suits any preStop sleep length.