skip to content

You delete a running Pod with kubectl delete pod and an equivalent Pod appears seconds later. Explain the mechanism that recreated it, and why Kubernetes controllers are described as level-triggered rather than edge-triggered.

level: middleimportance: must knowfreq 58%

answer

  1. ReplicaSet counts pods matching a selector
  2. edge = react to transition, level = react to current value
  3. events only enqueue a key; reconcile re-reads the world
  4. missed event delays convergence, never breaks it
  5. bare pods have no controller, stay deleted

basics

~20 s

A controller (the ReplicaSet behind the Deployment) constantly compares desired replica count with the pods it actually sees, and creates one when the count is short. Level-triggered means it acts on current state, not on the delete event, so it recovers even if it missed the event.

solid answer

~60 s

The Pod was recreated by the ReplicaSet controller. The Deployment owns a ReplicaSet whose spec says "3 pods matching this selector"; the controller lists pods matching that selector, counts 2, and creates one more. Nothing about your delete command was special — the controller reacts to the *count being wrong*, not to your action. That is the difference between level-triggered and edge-triggered. **Edge-triggered** logic reacts to transitions: "on pod-deleted, create a pod." If the process is restarting or the event is dropped, that edge is lost forever and the system stays broken. **Level-triggered** logic reads the current level of the world on every pass: desired 3, observed 2, therefore act. Missing an event only delays convergence; the next pass fixes it. In practice Kubernetes is edge-*driven* but level-*based* in its logic: watch events wake a controller quickly, but the reconcile function then re-reads full current state rather than trusting the event payload. Periodic resyncs and requeues cover anything the event stream missed. That is exactly why the system self-heals after network blips, controller restarts, or a node dying.

code

bash · 6 lines
bash
kubectl delete pod web-7d4f9c8b6-x2k9p
kubectl get pods -l app=web -w
kubectl describe rs -l app=web | grep -A5 Events
# a bare pod has no owner and is not recreated
kubectl run scratch --image=busybox --restart=Never -- sleep 3600
kubectl delete pod scratch

go deeper

for a junior

Name the ReplicaSet as the recreator and state the loop: desired count versus observed count. Note the new pod has a new name and IP.

for a middle

Define edge versus level triggering precisely and explain that events only enqueue a key while reconcile re-reads state; mention idempotency and resync.

for a senior

Emphasize the failure modes level-triggering buys off — dropped watches, controller restarts, node loss — and where self-healing does not apply.

for a principal

Argue the design tradeoff: eventual convergence with bounded staleness beats exactly-once event processing in a distributed system, and shape controller APIs so every reconcile is recomputable from cluster state alone.

## What actually recreated the pod A Deployment does not manage pods directly. It manages ReplicaSets; a ReplicaSet declares "there should be N pods matching this label selector". The ReplicaSet controller runs a loop: list the pods matching the selector, compare the count with spec.replicas, and create or delete the difference. Your `kubectl delete pod` simply removed one object; the very next pass observed 2 where 3 were wanted and created a replacement. The new pod is not the old one restored — it is a fresh pod with a new name, new IP, and empty ephemeral storage. The same mechanism explains node failure recovery: when a node goes away, its pods eventually leave the observed set, the count drops, and replacements are created elsewhere. Nothing special-cased "node failure"; the count was simply wrong. ## Edge-triggered versus level-triggered The terms come from hardware interrupts. **Edge-triggered** means acting on transitions. The logic is "when X happens, do Y." It is efficient, but it is only correct if you observe every transition. Miss one — a dropped message, a process restart, a watch that expired — and the system is permanently wrong, because nothing will tell you again. It also forces you to keep internal state to interpret each edge. **Level-triggered** means acting on the current value. The logic is "whatever happened, here is what I want and here is what exists; close the gap." Duplicate signals are harmless (the second pass finds no gap), missed signals only delay convergence, and a controller that restarts with zero memory behaves correctly immediately because it re-reads the world. ## How Kubernetes combines both A well-written controller is triggered by edges but reasons on levels. Watch events (and periodic resyncs) merely enqueue a key — usually just "namespace/name", with the event payload discarded. The reconcile function then fetches the current object and the current related objects, computes the delta, and acts. This gives the responsiveness of events with the robustness of polling. Several important properties follow: - **Idempotency is required.** Reconcile may run many times for one change — duplicate events, resyncs, retries after error. Running it twice must be the same as running it once. - **Reconcile must not assume it knows why it was called.** It gets a key, not a story. Any logic of the form "this must be a delete because…" is a bug. - **Convergence is eventual, not instantaneous.** Between the delete and the replacement there is a real gap where capacity is reduced. Level-triggering guarantees you get there, not that you were never wrong. - **Errors are handled by requeueing**, typically with rate-limited exponential backoff, rather than by aborting. ## Where self-healing stops Self-healing only covers objects a controller owns. A bare Pod created directly, with no ReplicaSet, Job, or StatefulSet above it, has no controller watching its count — delete it and it stays gone. Likewise, if you edit a managed pod's image in place, the ReplicaSet does not correct it, because its selector-and-count contract is satisfied; only pod-template changes at the Deployment level roll pods. Understanding what each controller actually compares tells you exactly what it will and will not fix.

  • If controllers re-read current state anyway, why bother with watches instead of just polling every few seconds?
    Watches give low latency and low cost at the same time: the API server pushes changes so controllers react in milliseconds without every controller re-listing every resource on a timer, which would be crushing load at cluster scale. The watch is an optimization on when to run reconcile; the reconcile logic itself remains level-based, so correctness never depends on the watch being perfect.
  • You edit the image field on a pod that a Deployment manages. Does the controller revert it?
    No. The ReplicaSet controller only ensures the right number of pods match its selector; it does not diff each pod against the template. The edited pod keeps running with your image until it is deleted or the Deployment rolls, at which point the replacement comes from the template. Drift correction at the pod-content level is not part of that contract.

Edge-triggered is a doorbell — miss the ring and you never know someone came. Level-triggered is looking out the window every minute: you always see who is standing there, no matter how many times they rang.

saying these in an interview costs you the question

  • Saying the deleted pod was restarted or restored — it is a brand-new pod with a new identity
  • Claiming the controller reacted to the delete event specifically, rather than to the wrong count
  • Believing bare pods are recreated after deletion
  • Saying level-triggered means polling only, and that Kubernetes does not use watches
  • Assuming a controller will revert any manual edit to any managed object

context