skip to content

Why would a 7-replica Kubernetes Deployment deleted with kubectl delete --cascade=foreground still exist with a deletionTimestamp twenty minutes later, and how do you unblock it safely?

level: seniorimportance: should knowfreq 42%

answer

  1. a marked object cannot be un-marked
  2. finalizers can only be removed
  3. walk the chain top-down
  4. unreachable nodes never confirm termination
  5. force skips grace, not finalizers

basics

~20 s

Its foregroundDeletion finalizer stays until every blocking dependent is gone, and a Pod is usually stuck on an unreachable node or on its own finalizer. Fix that lowest blocker; stripping finalizers by hand skips the cleanup they guard.

solid answer

~40 s

Foreground deletion sets the Deployment's `deletionTimestamp` and adds the `foregroundDeletion` finalizer. The Deployment then waits for its ReplicaSet, which waits the same way for its 7 Pods, because each Pod's reference sets `blockOwnerDeletion: true`. Twenty minutes means something at the bottom has not finished. I read the finalizers top-down with `kubectl get ... -o jsonpath`, then list the ReplicaSet and Pods with their `deletionTimestamp`, `finalizers` and `nodeName`. The usual culprits are a Pod on a `NotReady` node, which the kubelet never confirms as stopped, and a Pod carrying a custom finalizer whose operator is down. I fix that blocker: recover or remove the node, or restore the operator. `--force --grace-period=0` skips graceful termination but not finalizers. Patching a finalizer away skips its cleanup, and removing `foregroundDeletion` from the Deployment only hides the stuck Pod.

code

bash · 3 lines
bash
kubectl get deployment flag-eval -o jsonpath='{.metadata.deletionTimestamp}{" "}{.metadata.finalizers}{"\n"}'
kubectl get rs,pods -l app=flag-eval -o custom-columns=NAME:.metadata.name,DELETING:.metadata.deletionTimestamp,FINALIZERS:.metadata.finalizers,NODE:.spec.nodeName
kubectl get nodes

go deeper

for a junior

Recall that an object with a deletionTimestamp is being deleted and waits for its metadata.finalizers list to empty before it disappears.

for a middle

Explain how foregroundDeletion chains from the Deployment to its ReplicaSet to its Pods, and why blockOwnerDeletion makes each level wait.

for a senior

Diagnose the chain top-down, fix the lowest blocker, and explain what force deletion and manual finalizer removal each skip, and what that costs.

for a principal

Treat finalizer-adding operators as dependencies of deletion: set ownership, alerting on stale deletionTimestamp values, and a runbook for when manual removal is allowed.

## What `deletionTimestamp` really means In Kubernetes, a delete request does not always remove an object at once. If the object has entries in `metadata.finalizers`, the API server only **marks** it: it sets `metadata.deletionTimestamp` and leaves it in storage. A **finalizer** is a string key that some component promised to remove once its cleanup is done. The object leaves storage only when the list is empty. Three rules follow, and they shape every fix: - `deletionTimestamp` **cannot be unset**. A marked object is going away. There is no undo. - Once it is set, finalizers can **only be removed**. The API server rejects an update that adds a new one. - Anything that keeps a finalizer in place keeps the object alive. `kubectl delete` waits by default (`--wait` is true), so the command itself appears to hang. ## Why a foreground delete gets stuck Take the 7-replica feature-flag evaluation Deployment, `flag-eval`, on a 12-node GPU model-serving cluster, deleted with `kubectl delete deployment flag-eval --cascade=foreground`. That request sends `propagationPolicy: Foreground`: 1. The Deployment gets a `deletionTimestamp` and the **`foregroundDeletion`** finalizer. 2. The garbage collector in `kube-controller-manager` deletes the ReplicaSet with Foreground, so the ReplicaSet also gets `foregroundDeletion`. 3. The garbage collector deletes the 7 Pods. Each Pod's ownerReference has `blockOwnerDeletion: true`, so the ReplicaSet waits for every one of them. 4. Only when the last Pod is gone do the finalizers come off, first on the ReplicaSet and then on the Deployment. Twenty minutes later the Deployment still exists, so something **below** it has not finished. The usual causes: | Blocker | What you see | Why it blocks | |---|---|---| | Pod on an unreachable node | Pod `Terminating`, its node `NotReady` | the kubelet never confirms the containers stopped, so the Pod is not removed | | Pod with a custom finalizer | `metadata.finalizers` lists a key some operator owns | that operator is down or failing, so it never removes its key | | Garbage collector not making progress | nothing has a new `deletionTimestamp` | check the garbage collector's errors in the `kube-controller-manager` logs | ## Diagnosing top-down Follow the chain rather than guessing: 1. Read the owner: `kubectl get deployment flag-eval -o jsonpath='{.metadata.deletionTimestamp} {.metadata.finalizers}'`. Seeing `foregroundDeletion` confirms it is waiting on dependents. 2. List what still exists under it, with each object's deletion state, finalizers and node: ```bash kubectl get rs,pods -l app=flag-eval -o custom-columns=NAME:.metadata.name,DELETING:.metadata.deletionTimestamp,FINALIZERS:.metadata.finalizers,NODE:.spec.nodeName kubectl get nodes ``` 3. For a dependent that is still there, compare its `metadata.ownerReferences[].uid` with the owner's UID, so you know it really belongs to this chain. 4. For a Pod with a foreign finalizer, find the component that owns that key and check why it is not reconciling. ## Unblocking safely Fix the **lowest** blocker, not the top of the chain. - **Unreachable node.** Bring the node back, or remove its Node object if the machine is really gone. Once the Node is gone, the control plane cleans up Pods still bound to it. As a last resort, `kubectl delete pod <name> --grace-period=0 --force` removes the Pod from the API without waiting for the kubelet. The container **may still be running** on the partitioned machine, still holding its GPU and still answering requests if it can reach its clients. - **Custom finalizer.** Restore the owning operator so it finishes its cleanup and removes its key. Stripping the key by hand, for example with `kubectl patch pod <name> --type=merge -p '{"metadata":{"finalizers":null}}'`, **skips that cleanup**. Whatever the finalizer protected, such as an external registration or a lease, is leaked and must be cleaned up by hand. - **Force deletion is not finalizer removal.** `--force --grace-period=0` skips graceful termination only. A Pod whose finalizer list is not empty stays in storage anyway. Removing `foregroundDeletion` from the **Deployment** makes it vanish at once, but it only hides the problem. Its remaining dependents now point at a missing owner and are still garbage collected in the background, and the stuck Pod stays stuck for the same reason as before. ## Preventing the next one - Alert on objects whose `deletionTimestamp` is older than a few minutes. - Treat every operator that adds finalizers as a dependency of deletion: if it is down, deletion stops. - Prefer Background for routine cleanup. Use Foreground only when a caller must wait for the children to be gone.

  • A teammate proposes clearing the finalizers on the Deployment to finish the job. What actually happens?
    Removing `foregroundDeletion` lets the Deployment leave storage at once. Its ReplicaSet and Pods now point at a missing owner, so the garbage collector still deletes them in the background. The Pod that was stuck stays stuck for the same reason, whether that is an unreachable node or a finalizer nobody removes. The command returns, but the real problem and the GPU it holds are still there.
  • Can you cancel a deletion by removing deletionTimestamp or adding a finalizer to protect the object?
    No. `deletionTimestamp` cannot be unset once it is set, and the API server rejects any update that adds a new finalizer to an object that is being deleted. Finalizers can only be removed. The only way back is to let the object go and create it again.
  • When is force-deleting a Pod on a NotReady node risky?
    Force deletion removes the Pod from the API without waiting for the kubelet, so its container may still be running on the partitioned machine. For stateless Pods in a Deployment this usually means a short overlap. For workloads that need a single writer or a stable identity, two copies can run at once. Confirm the machine is really down before forcing.

saying these in an interview costs you the question

  • kubectl delete --force --grace-period=0 removes an object regardless of its finalizers
  • Clearing finalizers by hand is always safe because Kubernetes cleans up anyway
  • You can cancel a pending deletion by removing its deletionTimestamp
  • Removing foregroundDeletion from the Deployment also frees the stuck Pod
  • Restarting kube-controller-manager is the first fix for any stuck deletion