Describe what happens between 'kubectl delete pod' and the container disappearing in Kubernetes: which signal is sent, what terminationGracePeriodSeconds controls, and why in-flight requests can still fail during that window.
answer
- deletionTimestamp -> Terminating
- endpoint removal and SIGTERM are concurrent
- one budget: preStop + SIGTERM handling
- preStop sleep outlasts dataplane convergence
- SIGKILL at the end; signal reaches PID 1 only
basics
~20 sThe API server sets a deletionTimestamp and a grace period (default 30s). In parallel the Pod is removed from Service endpoints, and the kubelet runs any preStop hook then sends SIGTERM to PID 1 of each container; anything alive when the grace period ends gets SIGKILL. Requests still fail because endpoint removal propagates asynchronously.
solid answer
~60 sDeletion is a graceful, *concurrent* dance: 1. The API server marks the Pod with `deletionTimestamp` and the grace period (default 30s). The Pod shows as Terminating. 2. **In parallel**, two independent chains start. The endpoint controllers remove the Pod from its EndpointSlices, which then has to propagate to every kube-proxy, ingress and mesh sidecar on every node. And the kubelet begins shutdown: run the `preStop` hook, then send **SIGTERM** to PID 1 of each container. 3. When the grace period expires - one clock covering preStop *plus* SIGTERM handling - anything still alive gets **SIGKILL**. The failure mode is the race in step 2: dataplane updates are eventually consistent, so a proxy can still open new connections to a Pod that already received SIGTERM. The standard fix is a `preStop` sleep of a few seconds so the app keeps serving while routes drain, plus a real SIGTERM handler that stops accepting new work and finishes in-flight requests. Make sure the process actually receives the signal as PID 1, and that the grace period exceeds your longest request.
code
yaml · 9 linesspec:
terminationGracePeriodSeconds: 60
containers:
- name: api
image: registry.example.com/api:2.3.1
lifecycle:
preStop:
sleep:
seconds: 10go deeper
State the order: SIGTERM first, a default 30-second grace period, then SIGKILL if the process has not exited.
Add that preStop runs before SIGTERM and shares the same grace budget, and that endpoint removal happens in parallel rather than beforehand.
Explain the race that causes deploy-time 502s, prescribe the preStop sleep plus a real drain handler, and size the grace period against the longest request and against node-drain budgets.
Treat it as a platform default: standard preStop and grace settings baked into the chart, an SLO for zero-error deploys, and awareness of how node shutdown, preemption and sidecar ordering interact with it.
## The sequence When a Pod is deleted - explicitly, by a rolling update, by an eviction or by a scale-down - the API server does **not** remove the object immediately. It sets `metadata.deletionTimestamp` and records the grace period, and the object stays visible as `Terminating`. Two independent chains then run **concurrently**: **Chain A - routing removal.** The Pod's `Ready` condition is flipped and the endpoint/EndpointSlice controllers remove its address. Each node's kube-proxy (or ingress controller, or mesh sidecar) watches those objects and rewrites its iptables/IPVS/eBPF rules. This is eventually consistent and takes anywhere from milliseconds to seconds across a large cluster. **Chain B - container shutdown.** The kubelet runs the container's `preStop` hook if one is defined, waits for it, then sends **SIGTERM** to PID 1 in each container. It waits out the remaining grace period, then sends **SIGKILL** to whatever is still alive and tears down the sandbox. Only once the kubelet confirms termination is the API object actually removed. Crucially, **the grace period is one budget for preStop plus SIGTERM handling**, not one for each. A 20-second preStop under the default 30-second grace leaves the app about 10 seconds to drain before SIGKILL. ## Why requests still fail Nothing sequences chain A before chain B. A Pod can receive SIGTERM while proxies on other nodes still consider it a valid backend. If the app's SIGTERM handler immediately closes the listening socket, those in-flight and just-arrived connections get connection-refused or reset - visible to users as a burst of 502/504 on every deploy. The accepted remedy is a small `preStop` hook that simply sleeps (5-15 seconds, tuned to how fast your dataplane converges). During the sleep the container keeps serving normally while the endpoint update propagates. Only afterwards does SIGTERM arrive and the app drains. Since Kubernetes 1.30 you can express this with the native `sleep` lifecycle action instead of shelling out, which matters for images without a shell. ## What the application must do - **Receive the signal.** SIGTERM goes to PID 1 of the container only. If the image starts the app through a shell wrapper that does not forward signals, PID 1 swallows it and the app is eventually SIGKILLed - a container-image concern, fixed with exec-form entrypoints or a tiny init that forwards signals. - **Handle it properly.** Stop accepting new connections or messages, finish in-flight work, close database connections and consumer-group memberships, flush buffers, exit 0. Frameworks usually provide a graceful-shutdown hook; make sure it is wired. - **Fit in the budget.** `terminationGracePeriodSeconds` must exceed preStop plus the longest legitimate request or task. Long-running work may need hundreds of seconds; conversely a very long grace period slows node drains and rollouts, so choose deliberately. Setting it to 0 forces immediate SIGKILL and risks data loss. ## Related mechanics worth mentioning - Rolling updates delete Pods the same way, so this behaviour - not the Deployment strategy - decides whether a release drops requests. - Node drains and graceful node shutdown honour the same grace period, but node shutdown has its own, often shorter, kubelet-level budget: a Pod needing 120 seconds may still be cut off. - Preemption and eviction respect it too. Hard node failure does not: nothing graceful happens, the Pod is simply lost once the node stops reporting. - `kubectl delete pod --force --grace-period=0` removes the API object without waiting for the kubelet's confirmation. On a live node the container can briefly keep running; for StatefulSet Pods it risks two instances sharing one identity. ## What a good answer sounds like "deletionTimestamp, then concurrently endpoint removal and preStop-then-SIGTERM, SIGKILL at the end of one shared grace budget; endpoint updates are async, so add a preStop sleep and a real SIGTERM handler, and size the grace period above the longest request."
- Your app has a proper SIGTERM handler but you still see 502s on every deploy. What do you check first?The race between endpoint removal and SIGTERM: the app closes its listener before every proxy has stopped routing to it. Add a preStop sleep of several seconds so the Pod keeps serving while EndpointSlice updates propagate, and confirm the ingress or load balancer actually watches endpoints rather than caching backends for longer.
- What does terminationGracePeriodSeconds: 0 do, and when is it appropriate?It effectively means immediate SIGKILL with no chance to drain or flush, so in-flight requests and buffered writes are lost. It is only appropriate for stateless throwaway workloads or for forcing cleanup of a Pod stuck on an unreachable node, and even then it is risky for StatefulSet identities.
It is like a shop closing: the sign in the window (endpoint removal) and the staff packing up (SIGTERM) happen at the same time. If nobody stalls a few minutes, customers still walk in after the tills are shut.
saying these in an interview costs you the question
- Believing Kubernetes removes the Pod from endpoints before sending SIGTERM
- Thinking preStop gets its own time budget separate from the grace period
- Saying SIGKILL is sent first, or that SIGTERM goes to every process in the container
- Assuming a framework's graceful shutdown is enough without checking the signal reaches PID 1
- Using --force --grace-period=0 as a routine deletion habit