What is kube-controller-manager, and what does a single controller inside it actually do on each iteration?
answer
- one binary, many goroutine loops
- informer cache → work queue → sync
- compare spec vs observed, minimal write
- level-triggered: state, not events
- ownerReferences chain Deployment→RS→Pod
basics
~20 sIt is one control-plane process hosting many independent controllers (Deployment, ReplicaSet, Node, Job, endpoints, service accounts…). Each watches its objects via the API server, compares desired spec to observed state, and makes one corrective API call — repeatedly, until they match.
solid answer
~60 s**kube-controller-manager** is a single binary that runs dozens of controllers as goroutines sharing one set of informers and one leader election. Each controller owns a resource type: the ReplicaSet controller keeps N Pods alive, the Deployment controller manages ReplicaSets for rollouts, the Job controller drives Pods to completion, the Node lifecycle controller marks unreachable nodes NotReady and taints/evicts, and there are many smaller ones (endpoints/EndpointSlice, service account token, namespace deletion, PV binding). An iteration is always the same shape: an event or resync puts an object **key** on a work queue; a worker pops the key, reads the current object from a local cache, computes the difference between `spec` (desired) and observed reality, and issues the minimal API write — create a Pod, delete a Pod, patch status. Then it drops the key. If the write fails or state is still wrong, the key is requeued with backoff. Controllers are **level-triggered**: they act on current state, not on the event that woke them, so a missed or duplicated event is harmless.
go deeper
Name it as the component that keeps reality matching your YAML, and give one concrete loop such as ReplicaSet keeping replica count.
Describe informer cache, work queue with a namespace/name key, spec-versus-observed comparison, and the Deployment→ReplicaSet→Pod ownership chain.
Emphasise level-triggering and idempotency, periodic resync as drift repair, requeue with backoff, and diagnosing a wedged controller from queue-depth and retry metrics.
Discuss why bundling loops in one process with shared informers bounds API-server load, the split-out precedent of cloud-controller-manager, and the rules a custom controller or operator must follow to be a good citizen.
## One process, many loops kube-controller-manager is a single process on the control plane that hosts a large collection of independent controllers. They are bundled for operational simplicity — one deployment unit, one shared informer cache, one leader election, one set of flags — but conceptually each controller is standalone and could run separately (in fact cloud-specific controllers were split out into cloud-controller-manager for exactly that reason). Representative members: - **ReplicaSet controller** — for each ReplicaSet, count Pods matching its selector and owned by it; create or delete Pods until the count equals `spec.replicas`. - **Deployment controller** — never touches Pods. It creates and resizes ReplicaSets to execute a rolling update, respecting `maxSurge`/`maxUnavailable`, and records rollout status. - **Job / CronJob controllers** — run Pods to completion, track successes/failures, honour backoff limits; CronJob creates Jobs on schedule. - **Node lifecycle controller** — watches node heartbeats (Lease objects), marks a silent node `NotReady`, applies `node.kubernetes.io/unreachable` taints and drives eviction of its Pods. - **EndpointSlice controller** — keeps the list of ready Pod IPs behind each Service in sync. - **Namespace controller** — on namespace deletion, deletes contained resources before removing the namespace. - **PersistentVolume binder / attach-detach**, **service account token**, **garbage collector** (deletes objects whose owner disappeared), **TTL**, **resource quota**, and others. ## The anatomy of one iteration Every controller follows the same skeleton: 1. **Informers** maintain a local, watch-backed cache of the object types it cares about, so reads are in-memory and never hammer the API server. 2. Event handlers do almost nothing: they compute a **key** (`namespace/name`) and push it onto a **rate-limited work queue**. Related objects map back to their owner — a Pod event enqueues the *ReplicaSet* key via the ownerReference. 3. Workers pop keys. The queue **deduplicates**: ten rapid events for one object collapse into one processing pass, and a key being processed is not handed to a second worker, which removes most concurrency hazards. 4. `syncHandler(key)` reads the object from cache, lists its children from cache, computes the delta, and performs the smallest possible write through the API server. 5. On error the key is requeued with exponential backoff; on success it is forgotten. A periodic **resync** re-enqueues everything (typically every few minutes) so drift caused by a lost event or an external mutation is repaired anyway. ## Level-triggered, not edge-triggered This is the design point interviewers probe. A controller does not react to *what changed*; it reacts to *what is*. It never asks "was a Pod deleted?" — it asks "how many Pods exist right now versus how many should?" Consequences: duplicate events are harmless, missed events are repaired by the next event or resync, and a controller that crashes mid-work resumes correctly from current state with no journal or checkpoint. Idempotency comes for free because each sync recomputes from scratch. ## Ownership and chaining Controllers compose by layering, connected by `metadata.ownerReferences`. Deployment → ReplicaSet → Pod: the Deployment controller only manipulates ReplicaSets; the ReplicaSet controller only manipulates Pods; the scheduler binds Pods; kubelets run them. No component calls another; they communicate solely by reading and writing shared API objects. Owner references also drive the garbage collector — delete a Deployment and cascading deletion removes its ReplicaSets and then their Pods. ## Practical implications - **Spec is yours, status is theirs.** You write `spec`; controllers write `status`. Editing `status` by hand is pointless — the next sync overwrites it. - **Fighting controllers.** Manually deleting a Pod owned by a ReplicaSet recreates it; manually editing a ReplicaSet's replicas under a Deployment gets reverted on the next Deployment sync. Change the top of the chain. - **Load on the API server.** Because everything flows through watches and writes, a hot loop (a bad custom controller writing status every second) is a control-plane load problem, which is why work queues are rate-limited by default. - **Observability.** Controller-manager exposes queue depth, work duration and retry metrics; a growing queue depth for one controller is the signal that reconciliation is falling behind or wedged on errors.
- Why is level-triggered reconciliation preferred over acting on individual change events?Because it makes the controller tolerant of lost, duplicated or reordered events and of its own crashes. Each sync recomputes the answer from current state, so replaying a sync is harmless and no durable event log or cursor is needed. An edge-triggered design would have to guarantee exactly-once delivery of every change, which is unattainable across a network.
- You delete a Pod that belongs to a Deployment and it comes back. Which controller recreated it, and what would actually reduce the Pod count?The ReplicaSet controller recreated it: its Pod count fell below spec.replicas so it created a replacement. The Deployment controller was not involved directly. To reduce the count you must change the desired state at the top of the chain — scale the Deployment — since editing the ReplicaSet is reverted by the Deployment controller on its next sync.
saying these in an interview costs you the question
- Saying the Deployment controller creates Pods — it creates and resizes ReplicaSets; the ReplicaSet controller creates Pods.
- Describing controllers as reacting to events (edge-triggered) rather than reconciling current state.
- Claiming controllers talk directly to kubelets or nodes; all coordination is through the API server.
- Thinking each controller is a separate process or Pod by default.
- Believing hand-editing an object's status changes behaviour.