On a 3-node kubeadm Kubernetes cluster, why would each old payments-authorization API pod sit in Terminating for 47 seconds during a rollout, then exit with code 137 mid-request, and how do you fix it?
answer
- hook outlived its budget
- two-second kubelet floor
- 137 means SIGKILL
- grace plus two versus exactly grace
- admission checks the native sleep
basics
~20 sAn exec preStop sleep longer than terminationGracePeriodSeconds consumes the whole budget; the kubelet then allows only a 2-second stop window, so SIGKILL follows SIGTERM. Shorten the sleep and size the grace period to sleep plus drain.
solid answer
~50 sThe pod spec has `terminationGracePeriodSeconds: 45` and an exec `preStop` hook running `sleep 60`. The kubelet waits for the hook at most the full 45 seconds, then finds no time left and floors the remainder at its 2-second minimum, so the runtime sends SIGTERM and SIGKILLs 2 seconds later: 45 + 2 = 47, exit code 137, in-flight authorizations cut. I would confirm it from the spec, the `Killing` event timing and the container's `lastState.terminated.exitCode`. The contrasting signature is a pod that lasts exactly 45 seconds, which points to a process ignoring SIGTERM, typically a shell wrapper as PID 1. The fix: a native `preStop.sleep` of about 8 seconds, measured against routing propagation, and a grace period of sleep plus the longest authorization plus margin, for example 37. The native action is also validated against the grace period, so this misconfiguration cannot be admitted again.
code
yaml · 13 linesapiVersion: v1
kind: Pod
metadata:
name: payments-authz-broken
spec:
terminationGracePeriodSeconds: 45
containers:
- name: api
image: registry.example.com/payments-authz:1.8.2
lifecycle:
preStop:
exec:
command: ["sleep", "60"]go deeper
Recall that 137 means the container was killed with SIGKILL and that preStop time comes out of the grace period.
Explain the kubelet's order - preStop, remainder, 2-second floor, SIGTERM, SIGKILL - and reproduce the 45 plus 2 arithmetic.
Show how you separate the three signatures from timings and events, then derive sleep and grace from measured propagation and the longest request.
Discuss enforcing sleep-and-grace conventions across teams, for example preferring the validated native sleep, and the rollout and drain time that long budgets cost.
## The symptom, stated precisely A payments-authorization API runs as a Deployment on a three-node kubeadm cluster. On each rollout the old pods show `Terminating` for **47 seconds**, their containers end with **exit code 137** (128 + 9, meaning SIGKILL), and clients see authorizations fail mid-flight. The pod spec contains: - `terminationGracePeriodSeconds: 45` - a `preStop` hook of type `exec` running `sleep 60`, copied from an older service ## How the kubelet spends the budget The kubelet stops a container in a fixed order: 1. It takes the grace period - the pod's `deletionGracePeriodSeconds`, which comes from `terminationGracePeriodSeconds` unless the delete request overrode it. 2. If a `preStop` hook exists, it runs it and waits **at most the whole grace period** for it. Time the hook uses is subtracted from the budget. 3. Whatever remains is passed to the container runtime as the stop timeout. If the remainder is below **2 seconds**, the kubelet raises it to 2 - a floor that exists to avoid needless SIGKILLs. 4. The runtime sends the stop signal (SIGTERM by default), waits that timeout, then sends SIGKILL. Applied to this spec: | Step | Seconds elapsed | |---|---| | Deletion; hook `sleep 60` starts | 0 | | Kubelet stops waiting for the hook | 45 | | Remaining budget 0, floored to 2; SIGTERM sent | 45 | | SIGKILL; exit code 137 | 47 | The application received SIGTERM with two seconds left, so any authorization longer than that was killed. ## Telling it apart from the other full-grace failure A pod that always consumes its grace period has two common causes, and the timing separates them: - **Total = grace + 2 seconds** - the preStop hook outlived the budget, as here. - **Total = exactly the grace period** - the hook finished, SIGTERM was sent, and the process never exited. The usual reason is a process that does not act on SIGTERM, most often a shell-form command where `/bin/sh` is PID 1 and does not pass the signal on. Fix that in the image's entrypoint, not in Kubernetes. - **Total = a few seconds with errors** - the hook failed (for example no `sleep` binary), producing a `FailedPreStopHook` event, and SIGTERM came immediately. Evidence to collect: ```bash kubectl get deploy payments-authz -o jsonpath='{.spec.template.spec.terminationGracePeriodSeconds}{"\n"}{.spec.template.spec.containers[0].lifecycle}{"\n"}' kubectl get events --field-selector involvedObject.name=payments-authz-7d9c5b6f4-q2x8m ``` The `Killing` event marks when stopping began; the kubelet journal on the node, at verbosity 2 or higher, logs "PreStop hook not completed in grace period". A replacement pod has no memory of the old one, so capture the timings during the rollout. ## Fixing the budget Work out the numbers from what the service actually needs: - **Sleep** - long enough for endpoint removal to reach every kube-proxy and the ingress controller. Measure it; here 8 seconds. - **Drain** - the longest request that must finish. The authorization call has a 25-second upstream timeout. - **Margin** - a few seconds for the process to close connections and exit; here 4. That gives `terminationGracePeriodSeconds: 37` and `preStop.sleep.seconds: 8`. Two more rules keep it fixed: - Use the **native `sleep` action**. API validation rejects a sleep longer than `terminationGracePeriodSeconds`, so the original misconfiguration would have been refused at admission; an `exec` sleep is never checked. - Make sure the application exits on SIGTERM once its in-flight work is done, rather than waiting to be killed. ## Other deleters use the same budget - `kubectl drain` evicts pods, and an eviction becomes the same graceful delete; its `--grace-period` flag can shorten the budget. - A graceful node shutdown gives pods at most the kubelet's `shutdownGracePeriod` (default `0s`, which disables the feature), so a long pod grace period can still be cut short when a machine powers off. - A liveness or startup probe can carry its own `terminationGracePeriodSeconds`, used instead of the pod's value when that probe triggers the restart. Once the new values ship, verify them the same way the fault was found: roll the Deployment under realistic authorization traffic, time how long old pods stay in `Terminating` (it should now be the sleep plus the real drain, well under 37 seconds), and confirm the containers exit with code 0 rather than 137. A payments path deserves that check in every release that touches the pod spec.
- Why would changing that hook to `preStop.sleep.seconds: 60` with the same 45-second grace period not even be accepted?API validation checks the native sleep action against the pod's `terminationGracePeriodSeconds` and rejects a value above it. The exec form runs an arbitrary command, so the API server cannot know how long it takes and does not check it; that is how the broken spec was admitted.
- The same service runs on interruptible nodes that give a short notice before shutdown. What limits its termination time there?When the kubelet's graceful node shutdown is enabled, pods are stopped within the kubelet's `shutdownGracePeriod` budget, minus the part reserved for critical pods, whatever the pod asks for. If that budget or the interruption notice is shorter than the pod's grace period, SIGKILL arrives earlier, so the preStop sleep and drain must fit the shorter window.
- Does a PodDisruptionBudget change how long an evicted pod gets to shut down?No. A PodDisruptionBudget only decides whether an eviction is allowed at that moment. Once allowed, the pod goes through the normal graceful deletion with its own grace period, unless the caller, such as `kubectl drain --grace-period`, overrides it.
saying these in an interview costs you the question
- Exit code 137 always means the container ran out of memory
- The grace period restarts after the preStop hook finishes
- A longer preStop sleep always makes shutdown safer
- Kubernetes rejects any preStop hook longer than the grace period
- A pod lasting exactly its grace period proves the hook is too long
- Raising terminationGracePeriodSeconds alone fixes SIGTERM being ignored