skip to content

Graceful Shutdown

Deleting a pod starts two things at once: endpoints are withdrawn and the container gets SIGTERM, with SIGKILL after the grace period. Interviewers reach for it whenever a candidate calls rolling updates zero-downtime - deploy-time 502s live here.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

3

When a Kubernetes pod is deleted, why can it still receive new requests after SIGTERM, and how does a preStop sleep prevent the resulting errors?

level: middleimportance: must knowfreq 66%

answer

  1. one timestamp, two watchers
  2. kubelet local, routing cluster-wide
  3. delay the signal, not traffic
  4. hook time spends the grace budget
  5. no sleep binary in distroless

basics

~10 s

Deletion starts two unordered paths: the kubelet stops the container while endpoint controllers, kube-proxy and ingress controllers withdraw the pod from routing. A preStop sleep delays SIGTERM until routing has caught up.

solid answer

~50 s

Setting `deletionTimestamp` is observed in parallel by the kubelet on the pod's node and by the EndpointSlice controller. The kubelet runs the `preStop` hook, then has the runtime send the stop signal (SIGTERM by default). Meanwhile the endpoint is marked terminating and not ready, and every node's kube-proxy plus any ingress controller must see that and reprogram. Nothing orders the two, so an app that stops accepting connections on SIGTERM can refuse traffic that routing rules still send it - the classic deploy-time 502 or connection reset. The fix is to keep the container serving normally for a few seconds first: `lifecycle.preStop.sleep.seconds: 5` (or an exec `sleep` on images that ship the binary) delays SIGTERM while routing converges. The sleep counts against `terminationGracePeriodSeconds`, so size the grace period to cover the sleep plus the application's own drain.

code

yaml · 24 lines
yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: payments-authz
spec:
  replicas: 3
  selector:
    matchLabels:
      app: payments-authz
  template:
    metadata:
      labels:
        app: payments-authz
    spec:
      terminationGracePeriodSeconds: 37
      containers:
        - name: api
          image: registry.example.com/payments-authz:1.8.3
          ports:
            - containerPort: 8443
          lifecycle:
            preStop:
              sleep:
                seconds: 8

go deeper

for a junior

Recall that deleting a pod both removes it from load balancing and signals the container, and that these two things are not ordered.

for a middle

Explain both paths - kubelet with preStop then SIGTERM, EndpointSlice then kube-proxy and ingress - and how a preStop sleep and the grace budget fit together.

for a senior

Show you size the sleep from measured propagation, check the image can run the hook, and verify with deploy-time error rates rather than trusting the manifest.

for a principal

Discuss making the sleep a platform default versus per-service tuning, and the cost of longer terminations on rollout speed and node drains.

## Two paths start from one field When a Pod is deleted, the API server sets `metadata.deletionTimestamp`. That single change is watched by several independent components, and **none of them waits for the others**: - **Path A - the kubelet** on the pod's node begins termination: run each container's `preStop` hook, then ask the container runtime to stop the container, which sends the image's stop signal (SIGTERM unless the image says otherwise), and finally SIGKILL whatever is left when the grace period runs out. - **Path B - traffic withdrawal**: the EndpointSlice controller in kube-controller-manager rewrites the pod's endpoint with `terminating: true` and `ready: false`. kube-proxy on **every** node must receive that update and rewrite its rules; an ingress controller or external load balancer must do the same for its own upstream list. Path A is local and fast - the kubelet can send SIGTERM within milliseconds. Path B crosses the API server, a controller, a watch fan-out and a rule rewrite on each node, and routinely takes from under a second to several seconds, longer under API server load. ## What goes wrong Many applications react to SIGTERM by closing their listening socket and finishing in-flight work. That is correct behaviour, but in the window where Path A has finished and Path B has not: 1. A client request is load-balanced to the terminating pod's IP by a node whose kube-proxy rules are still old. 2. The pod's listener is closed, so the connection is refused or reset. 3. The client, or the ingress controller in front of it, reports a **502/503 or connection error**. This is why a rolling update that is "zero-downtime" on paper still produces a burst of errors on every deploy. The application's shutdown code is not the bug; the ordering is. ## The preStop sleep workaround The `preStop` hook runs **before** the stop signal is sent, and the kubelet waits for it to finish (up to the grace period). A hook that simply sleeps turns it into a delay for Path A: ```yaml lifecycle: preStop: sleep: seconds: 5 ``` During those seconds the container keeps serving normally, including any new requests that stale routing still sends it, while Path B completes. When the sleep ends, SIGTERM arrives at a pod that is no longer receiving new traffic, and the application's normal drain only has to finish what is already in flight. | Option | Needs a binary in the image | Checked against grace period at admission | |---|---|---| | `preStop.sleep.seconds` | no | yes - must not exceed `terminationGracePeriodSeconds` | | `preStop.exec` running `sleep` | yes | no | | Application delays its own shutdown after SIGTERM | no | no | The native `sleep` action suits minimal and distroless images that have no shell. An `exec` hook on an image without a `sleep` binary fails, the kubelet records a `FailedPreStopHook` event, and SIGTERM follows immediately - silently removing the protection. ## Sizing the numbers The grace period is a **single budget** shared by the hook and the application: - `terminationGracePeriodSeconds` (default **30**) starts counting when termination begins. - The preStop sleep consumes part of it. - Only the remainder is left for the process to drain after SIGTERM before SIGKILL. Pick the sleep from measured propagation delay (a few seconds is typical), then set the grace period to **sleep + longest in-flight work + a small margin**. Raising the sleep without raising the grace period just moves the SIGKILL earlier relative to SIGTERM. ## What the sleep does not fix - It does not remove **existing keep-alive connections**; those stay open until the application closes them, which is the application runtime's job. - It does not help if the process ignores SIGTERM, for example when a shell wrapper runs as PID 1 - the container then simply waits out the grace period and is SIGKILLed. - It is a heuristic: if routing takes longer than the sleep, the race reappears, only less often. - It does not shorten a rollout. Every old pod now lives at least as long as the sleep, so a Deployment with many replicas, or a node drain touching many pods, takes correspondingly longer to finish its terminations. The practical test is empirical: roll the Deployment while a load generator runs, and compare the error count with and without the hook. If errors remain, measure how long routing actually takes to converge on the cluster and adjust the sleep, rather than guessing a larger number.

  • Should the application start failing its readiness probe as soon as it receives SIGTERM to get removed from the Service faster?
    It adds little for removal: the EndpointSlice controller already marks a terminating pod's endpoint not ready the moment `deletionTimestamp` is set, whatever the probe says. The delay you are fighting is propagation to kube-proxy and ingress controllers, which a probe does not shorten. The preStop sleep addresses that delay directly.
  • What happens if the preStop hook is still running when the grace period ends?
    The kubelet stops waiting for the hook, then still sends the stop signal with a minimum two-second window before SIGKILL, so the pod lives roughly two seconds past its grace period and the application gets almost no time to drain. A native sleep action longer than the grace period is rejected at admission; an exec sleep is not.
  • Does the preStop sleep also help when a node is drained with `kubectl drain`?
    Yes. A drain evicts pods through the Eviction API, and an eviction ends in the same graceful pod deletion, so the same deletionTimestamp, preStop hook and grace period apply. `kubectl drain --grace-period` can override the pod's own grace period, which also shrinks the time left after the sleep.

It is like a shop that locks its door the instant the head office removes it from the online map, while customers already walking over with the old map still arrive; waiting a few minutes before locking lets the map catch up.

saying these in an interview costs you the question

  • Kubernetes removes the pod from the Service before sending SIGTERM
  • Handling SIGTERM correctly in the app is enough for zero errors
  • The preStop sleep is added on top of the grace period
  • The readiness probe must fail before the endpoint is removed
  • A preStop exec sleep works in any container image
open as a page

What does `kubectl delete pod --grace-period=0 --force` actually do in Kubernetes, and when is it safe to use?

level: juniorimportance: should knowfreq 46%

basics

~20 s

Force deletion removes the Pod object from the API server at once, without waiting for the kubelet to confirm the containers stopped. The processes may keep running on the node, so use it only when the pod is known to be dead.

open as a page

On a 3-node kubeadm Kubernetes cluster, why would each old payments-authorization API pod sit in Terminating for 47 seconds during a rollout, then exit with code 137 mid-request, and how do you fix it?

level: seniorimportance: should knowfreq 38%

basics

~20 s

An exec preStop sleep longer than terminationGracePeriodSeconds consumes the whole budget; the kubelet then allows only a 2-second stop window, so SIGKILL follows SIGTERM. Shorten the sleep and size the grace period to sleep plus drain.

open as a page