skip to content

Kubernetes controllers are described as level-triggered rather than edge-triggered, and their reconcile function must be idempotent. Explain what that means and why a controller written as a set of event handlers for create, update and delete tends to break.

level: middleimportance: must knowfreq 56%

answer

  1. Edge = react to change; level = read current state
  2. Reconcile gets only namespace/name, no diff
  3. Deterministic child names, create-or-update
  4. Guard external side effects; record in status
  5. Resync re-checks everything and is harmless

basics

~20 s

Level-triggered means the controller reads current desired and actual state on every pass and closes the gap, instead of reacting to individual change events. Reconcile may run many times for one change, in any order, after restarts, so it must be safe to repeat and must not depend on having seen previous events.

solid answer

~50 s

**Edge-triggered** logic reacts to transitions: "on update, do X". If an event is missed — controller restart, watch disconnect, coalesced updates — the transition is lost forever and the system stays wrong. **Level-triggered** logic reads the *current level*: fetch the object, observe what actually exists, compute the difference, act. An event is only a hint that it is worth looking again; the truth is always re-read. That is why `Reconcile(ctx, req)` receives only a namespace/name, not a diff or an event type. It must be **idempotent**: running it twice with the same state produces the same result and no duplicate side effects. Practically that means create-or-update rather than blind create, guarding external actions with a check or a recorded marker in `status`, tolerating partial progress from a crashed previous run, and returning an error or a requeue instead of trying to finish everything in one pass. The payoff is robustness: crashes, missed events and periodic resyncs all converge to the same correct state.

code

go · 29 lines
go
func (r *AppReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) {
    var app v1.App
    if err := r.Get(ctx, req.NamespacedName, &app); err != nil {
        return ctrl.Result{}, client.IgnoreNotFound(err)
    }

    desired := buildDeployment(&app) // deterministic name: app.Name + "-web"
    var found appsv1.Deployment
    err := r.Get(ctx, client.ObjectKeyFromObject(desired), &found)
    switch {
    case apierrors.IsNotFound(err):
        if err := r.Create(ctx, desired); err != nil {
            return ctrl.Result{}, client.IgnoreAlreadyExists(err)
        }
    case err != nil:
        return ctrl.Result{}, err
    default:
        if !equalSpec(&found, desired) {
            found.Spec = desired.Spec
            if err := r.Update(ctx, &found); err != nil {
                return ctrl.Result{}, err // conflict: requeued, re-read next pass
            }
        }
    }

    app.Status.ObservedGeneration = app.Generation
    app.Status.ReadyReplicas = found.Status.ReadyReplicas
    return ctrl.Result{}, r.Status().Update(ctx, &app)
}

go deeper

for a junior

Define edge versus level triggering and say that reconcile can run many times for one change, so it must be repeatable.

for a middle

Show the reconcile body — get, observe, diff, act, update status — and name the idempotency techniques: deterministic names, create-or-update, tolerate AlreadyExists and Conflict.

for a senior

Discuss failure modes across restarts and watch gaps, guarding external side effects with existence checks and status records, and why resync makes the system eventually correct.

for a principal

Position level-triggered reconciliation as the reason Kubernetes converges despite unreliable delivery, and connect it to API design choices such as the status subresource, observedGeneration and optimistic concurrency.

## Two ways to write a controller **Edge-triggered** design responds to *changes*: "when a resource is created, provision it; when the replica count changes from 3 to 5, add two." The handler encodes the transition, so correctness depends on receiving every event exactly once, in order. **Level-triggered** design responds to *state*: "whatever happened, here is the object as it exists now, and here is what exists in the world; make the world match." Kubernetes chose this, and it shapes every controller you will write or debug. ## Why edge-triggered breaks in a distributed system The delivery guarantees simply are not there: - **Watches disconnect.** Network blips, API server restarts and etcd compaction all interrupt the stream. On reconnect a client may resync from a fresh list, so intermediate transitions are never delivered. - **Events are coalesced.** If an object changes from 3 to 4 to 5 replicas quickly, the controller may only ever observe 5. A handler that reasoned "increment by one" is now wrong. - **Controllers restart.** Rollouts, OOM kills, node drains and leader-election changes all mean a fresh process with no memory of past events. - **Ordering across resources is not guaranteed.** The Pod your controller created may be observed before the parent update that caused it. - **Delivery is at-least-once.** The same change can be processed more than once. Each of these breaks transition-based logic and none of them break state-based logic. ## What Reconcile actually receives A controller-runtime reconciler is called with just a request key: namespace and name. Deliberately. There is no event type and no old/new pair, precisely so you cannot write edge-triggered code by accident. The canonical body: 1. **Get** the object by key. If it is not found, it was deleted — usually nothing to do because dependents are garbage collected, so return cleanly. 2. **Observe** the actual world: list or get the owned resources, query the external system if there is one. 3. **Compute the difference** between spec and observation. 4. **Act** to close the gap, using create-or-update semantics. 5. **Update status** to reflect what was observed, typically including `observedGeneration` and conditions. 6. **Return** — done, or an error, or a request to be called again later. ## What idempotent means concretely Running the same reconcile twice on the same state must not produce two of anything or corrupt progress: - **Deterministic names.** Derive child object names from the parent (`<cr-name>-config`) so a re-run addresses the same object instead of creating a second one. Never use random suffixes for objects you must find again. - **Create-or-update, not create.** Try to get; if absent create; if present compute the desired form and patch only if it differs. Treat AlreadyExists as success and re-read. - **Guard non-Kubernetes side effects.** Creating a cloud bucket or a database user is not naturally idempotent. Use an idempotency key or a deterministic name, check for existence first, and record the outcome (an ID or a condition) in `status` so a re-run recognises completed work. Do the external action first, then record it — and be able to recover if the process dies between the two, because that gap always exists. - **Tolerate partial progress.** A previous run may have created two of three resources. The next pass must simply continue, which falls out naturally from observe-then-diff. - **Do not accumulate in-memory state across invocations.** Anything you need must be re-derivable from the cluster, because the process can restart between calls. - **Handle conflicts.** Updates use optimistic concurrency via `resourceVersion`; a conflict error means someone else wrote first. Re-read and retry rather than force-overwriting, and prefer patches over whole-object updates for fields you do not own. - **Separate spec and status writes.** With the status subresource, writing status does not bump generation, which prevents your own status write from triggering an endless reconcile loop. ## Resync: the safety net Informers periodically re-deliver everything in their cache (a resync). Under edge-triggered logic that would be a storm of spurious work; under level-triggered logic it is harmless and useful — every object gets re-checked, so any drift, missed event or bug-induced divergence is eventually corrected. "Eventually correct" is the property level-triggering buys. ## Reasoning about it in an interview A strong answer states the definition, then names the failure modes (restart, missed watch event, coalesced updates, duplicate delivery) and shows what the code does about them: re-read state, deterministic child names, create-or-update, guarded external calls, status as the record of completed work. A weak answer talks only about "handling create and delete events", which reveals an edge-triggered mental model.

  • Your reconcile creates a bucket in an external cloud API. How do you keep that idempotent?
    Give the bucket a deterministic name derived from the custom resource (including a stable UID if names can be reused), check for existence before creating, and treat an already-exists response as success. Record the resulting identifier in `status` so subsequent passes recognise the work as done. Accept that a crash between creating the bucket and writing status is possible, so the existence check — not the status field alone — must be authoritative.
  • Why does the reconcile signature give you only a namespace and name rather than the changed object?
    Because handing you the object or a diff would invite edge-triggered logic based on a possibly stale or missed transition. Forcing a fresh read from the informer cache keeps the controller level-based: it always works from current state, so restarts, dropped watch events and coalesced updates all converge correctly.
  • What is observedGeneration and why do controllers write it?
    `metadata.generation` increments when the spec changes, but not when status changes. A controller copies it into `status.observedGeneration` once it has acted on that spec version. Comparing the two tells users and automation whether the reported status reflects the current desired state or a previous one, which is essential for readiness gates and for tools waiting on rollouts.

Edge-triggered is a doorbell you might not hear; level-triggered is glancing at the door every so often to see whether anyone is standing there.

saying these in an interview costs you the question

  • Writing separate onCreate/onUpdate/onDelete handlers and assuming every event arrives exactly once
  • Assuming reconcile runs once per change, so duplicate side effects 'cannot' happen
  • Keeping in-memory state between reconciles instead of re-deriving from the cluster
  • Creating child objects with random names, so a re-run cannot find what the previous run made
  • Treating a Conflict error as a bug and forcing an overwrite instead of re-reading and retrying
  • Believing periodic resync is wasted work rather than the drift-correction safety net

context