skip to content

Explain what a Kubernetes finalizer is, how a controller should use one to release external resources before an object goes away, and how you would debug a namespace or custom resource stuck in a Terminating state.

level: seniorimportance: must knowfreq 50%

answer

  1. Delete sets deletionTimestamp, not removal
  2. Add finalizer before creating external state
  3. Cleanup idempotent; 404 = success; remove own finalizer
  4. Stuck = controller gone or cleanup failing
  5. Namespace: read status.conditions for the failing API group

basics

~20 s

A finalizer is a string in metadata.finalizers that blocks final deletion. A delete request only sets deletionTimestamp; the object persists until every finalizer is removed. The controller does its cleanup, then removes its own finalizer. Stuck Terminating usually means the responsible controller is gone or failing.

solid answer

~50 s

When an object has entries in `metadata.finalizers`, a DELETE does not remove it: the API server sets `metadata.deletionTimestamp` and the object stays, now read-mostly. Each controller named by a finalizer performs its cleanup and then removes **its own** finalizer string; when the list is empty the API server deletes the object for real. The controller pattern in reconcile: 1. If `deletionTimestamp` is nil and my finalizer is absent, add it (before creating anything external). 2. If `deletionTimestamp` is set and my finalizer is present, run cleanup — idempotently, since it may re-run — then remove the finalizer and return. 3. If it is set and my finalizer is gone, do nothing. To debug stuck Terminating: read `metadata.finalizers` to see who is blocking; check whether that controller is running and what its logs say; for namespaces check `status.conditions`, which name the API groups whose resources cannot be enumerated (usually a broken aggregated API service). Force-removing a finalizer is a last resort and leaks whatever it protected.

code

go · 23 lines
go
const finalizer = "db.example.com/cleanup"

if app.DeletionTimestamp.IsZero() {
    if !controllerutil.ContainsFinalizer(&app, finalizer) {
        controllerutil.AddFinalizer(&app, finalizer)
        if err := r.Update(ctx, &app); err != nil {
            return ctrl.Result{}, err
        }
    }
    // ...normal reconcile, may now create external resources...
    return ctrl.Result{}, nil
}

if controllerutil.ContainsFinalizer(&app, finalizer) {
    if err := r.deleteExternalBucket(ctx, &app); err != nil { // 404 treated as success
        return ctrl.Result{}, err // requeued with backoff
    }
    controllerutil.RemoveFinalizer(&app, finalizer)
    if err := r.Update(ctx, &app); err != nil {
        return ctrl.Result{}, err
    }
}
return ctrl.Result{}, nil

go deeper

for a junior

Say that a finalizer blocks deletion until the controller finishes cleanup, and that a delete only sets a deletionTimestamp while the object still exists.

for a middle

Show the two-branch reconcile, the ordering rule that the finalizer is added before external state, and that each controller removes only its own entry.

for a senior

Add the debugging path — read finalizers, find the owning controller, read logs and conditions, check namespace status conditions and aggregated APIServices — and explain why force-removal leaks resources.

for a principal

Discuss deletion as a distributed transaction with no rollback: guarantee versus liveness, bounded escape hatches, uninstall ordering for operators, and how to keep stuck deletions observable across a fleet.

## What a finalizer is `metadata.finalizers` is a list of arbitrary strings, conventionally domain-qualified like `db.example.com/cleanup`. Their only meaning is: *this object may not be removed from etcd while this list is non-empty*. They are a pre-delete hook implemented as data. On DELETE the API server checks the list. If non-empty, it does not remove the object; it stamps `metadata.deletionTimestamp` and increments generation. The object remains readable and is now conventionally treated as "being deleted" — most controllers stop making changes to it apart from removing their own finalizer. Once every finalizer is removed the API server completes the deletion. ## Why controllers need them OwnerReferences and cascading garbage collection cover Kubernetes objects in the same namespace. They cannot cover: - **External systems**: a cloud bucket, DNS record, load balancer, database user, license seat, or an entry in a third-party API. - **Cluster-scoped objects created by a namespaced custom resource**, which ownership rules forbid. - **Ordered shutdown**: quiescing an application, taking a final backup, deregistering from a service mesh, draining traffic before storage disappears. Without a finalizer, the custom resource can vanish while the controller is offline, and the knowledge of what to clean up disappears with it. The finalizer keeps the object — and its spec and status — alive until the cleanup has actually happened. ## The correct reconcile shape ``` if obj.DeletionTimestamp.IsZero() { ensure finalizer present // add before creating external state ...normal reconcile... } else { if finalizer present { run cleanup (idempotent) on success: remove finalizer, update object on failure: return error -> requeued with backoff } return } ``` Key disciplines: - **Add the finalizer before creating external state**, otherwise a delete arriving in the gap leaks the resource. - **Cleanup must be idempotent and tolerate "already gone"**: treat a 404 from the external API as success, because the pass may re-run after a crash. - **Remove only your own finalizer.** Others belong to other controllers. - **Never block forever.** If cleanup cannot succeed, surface it: set a condition, emit an Event, and let backoff retry. A silent permanent block is how objects get stuck. - **Consider a bounded escape hatch** for cleanups that can legitimately fail forever — an annotation that means "give up and release", used deliberately by an operator, is better than users editing finalizers by hand. - **Expect conflicts**: the object is being written by others too; use patches and retry on conflict rather than whole-object updates. ## Finalizers you did not write The platform uses them too: `foregroundDeletion` implements foreground cascading deletion; `kubernetes.io/pvc-protection` and `pv-protection` keep volumes from being deleted while in use; namespace deletion is itself finalizer-driven. Seeing these is normal, not a bug. ## Debugging stuck Terminating **A custom resource stuck deleting.** Read the object: `kubectl get <kind> <name> -o jsonpath='{.metadata.finalizers}'`. Each string names the responsible party. Then check whether that controller is running at all — an operator uninstalled while its custom resources still exist is the single most common cause, because nobody is left to remove the finalizer. If it is running, read its logs and the object's conditions: cleanup is probably failing on a real error such as expired cloud credentials or a dependency that refuses deletion. **A namespace stuck Terminating.** Namespace deletion requires deleting every resource inside it, so it is blocked either by an object with an unresolved finalizer or by an API group the namespace controller cannot enumerate. `kubectl get namespace X -o yaml` and read `status.conditions`: fields like `NamespaceDeletionDiscoveryFailure` name the failing API group, which is usually an **aggregated APIService** whose backing service is gone. Fix or delete the broken APIService and deletion resumes. Otherwise enumerate leftover objects across all API resources in that namespace and inspect their finalizers. **Force removal.** Editing an object to empty its `finalizers` list makes deletion complete immediately — and skips the cleanup the finalizer existed to guarantee. That means orphaned cloud resources, still-billing load balancers, undeleted data. Do it only after establishing what will leak and recording it for manual cleanup. For namespaces there is a `/finalize` subresource used for the same purpose, with the same caveat. Uninstalling an operator before deleting its custom resources is the mistake that leads people here; the correct order is delete the custom resources first, watch them go, then remove the operator and its CRDs. ## Interview framing Define the mechanism precisely (delete becomes a timestamp; removal happens when the list empties), show the two-branch reconcile with the finalizer added *before* external state, and then demonstrate the debugging path from `metadata.finalizers` to the owning controller to its logs. Finish with the honest statement that force-removal is data loss by another name.

  • Why is it wrong to force-remove a finalizer as a first response to a stuck object?
    The finalizer exists to guarantee cleanup, so removing it completes the deletion while skipping that cleanup: cloud load balancers, buckets, DNS records or database users keep existing and keep costing money, with no Kubernetes object left to point at them. The right sequence is to find the owning controller, read its logs, fix the real failure, and only force removal once you know exactly what will leak and have written it down for manual cleanup.
  • Someone uninstalled an operator while its custom resources still existed, and now those resources will not delete. What happened and what is the fix?
    The objects still carry the operator's finalizer, but no controller is running to remove it, so deletion blocks indefinitely. The clean fix is to reinstall the operator, let it process the deletions, then delete the custom resources first and remove the operator and its CRDs afterwards. If reinstalling is impossible, you must strip the finalizers manually and clean up the external resources by hand.
  • Why must the finalizer be added before the controller creates external resources?
    If a delete arrives between creating the external resource and adding the finalizer, the object is removed immediately and the controller loses the record of what it provisioned, leaking it. Adding the finalizer first guarantees that any delete after that point leaves the object in place with its spec and status until cleanup has run.
  • What should the controller do if cleanup keeps failing?
    Return the error so the item is requeued with exponential backoff, and make the failure visible: set a Degraded or DeletionBlocked condition, emit an Event, and expose it in metrics. It should not remove the finalizer to escape, and it should not block silently. For failures that can never succeed, provide an explicit, auditable escape such as an annotation an operator sets deliberately.

A finalizer is a checkout hold on a hotel room: you can announce departure at any time, but the room is not released until housekeeping and the minibar audit each sign off.

saying these in an interview costs you the question

  • Thinking a DELETE request removes the object immediately even when finalizers are present
  • Adding the finalizer after creating external resources, leaving a leak window
  • Removing another controller's finalizer, or clearing the whole list from your own reconcile
  • Force-clearing finalizers as the first troubleshooting step instead of finding the blocked controller
  • Deleting an operator or its CRDs before deleting the custom resources that carry its finalizers
  • Writing cleanup that fails when the external resource is already gone instead of treating 404 as success

context