Walk through the machinery that sits between a change in the Kubernetes API and your controller's reconcile function being called: watches, informers, caches and work queues. Why is a work queue used rather than calling reconcile directly from the event handler?
answer
- LIST then WATCH; informer keeps a cache
- Handlers enqueue keys, never reconcile inline
- Queue: dedupe, one worker per key, backoff retry
- RequeueAfter for polling; never sleep in reconcile
- Cache is stale-ish; watch scope drives memory
basics
~20 sA shared informer watches the API server, keeps an in-memory cache of objects, and on each event pushes only the object's key onto a rate-limited work queue. Workers pop keys and reconcile, reading from the cache. The queue deduplicates keys, bounds concurrency, and provides retry with backoff.
solid answer
~60 sThe chain is: **watch → informer + cache → event handler → work queue → worker → reconcile**. An **informer** does an initial LIST, then WATCHes from that resourceVersion, storing objects in a local **cache (store/indexer)**. Controllers read from that cache instead of hitting the API server, which is why reads are cheap but may be slightly stale. Informers are **shared**: many controllers watching Pods use one watch stream. The informer's handlers do not reconcile. They enqueue the object's **key** (namespace/name) onto a rate-limited work queue, and mapping functions turn events on owned or related objects into the owner's key. The queue matters because it: **deduplicates** — many rapid changes to one object collapse into one pending key, and one item is never processed by two workers at once; **decouples** rate — a burst of events never blocks the watch stream; **bounds concurrency** via a fixed worker count; and **owns retries** — a failed reconcile is requeued with exponential backoff and jitter instead of hot-looping. Requeue-after supports polling external state.
code
go · 9 linesfunc (r *AppReconciler) SetupWithManager(mgr ctrl.Manager) error {
return ctrl.NewControllerManagedBy(mgr).
For(&v1.App{}). // primary resource: enqueue its own key
Owns(&appsv1.Deployment{}). // owned children map back to the owner key
Watches(&corev1.ConfigMap{}, // arbitrary relation: custom mapping
handler.EnqueueRequestsFromMapFunc(r.appsUsingConfigMap)).
WithOptions(controller.Options{MaxConcurrentReconciles: 4}).
Complete(r)
}go deeper
Name the chain — watch, informer with cache, event handler enqueues a key, worker calls reconcile — and say the queue exists to retry and avoid duplicate work.
Explain deduplication, per-key serialization, and the difference between returning an error, Requeue and RequeueAfter, plus that reads come from a possibly stale cache.
Add operational depth: backoff rate limiters, MaxConcurrentReconciles tuning, watch scoping for memory, leader election, and the metrics you use to spot a controller falling behind.
Discuss controller throughput as a capacity problem across the fleet — API-server and etcd pressure, cache memory versus cluster size, resync period cost, and sharding strategies when one controller cannot keep up.
## The full path 1. **Watch.** The client issues a LIST to get a consistent snapshot plus a `resourceVersion`, then a WATCH from that point, receiving Added/Modified/Deleted events over a long-lived connection. 2. **Informer.** The informer maintains that stream, applies events to a local **store** (a thread-safe cache, often with indexes such as by-namespace or by-owner), and invokes registered event handlers. If the connection drops or the resourceVersion becomes too old, it re-lists and resumes — which is exactly why controllers cannot rely on seeing every intermediate transition. 3. **Shared informer factory.** Multiple controllers in one process that care about the same resource share a single informer and therefore a single watch and a single cache copy, which is what keeps memory and API-server load manageable. 4. **Event handlers.** Instead of doing work, handlers compute one or more **keys** to reconcile and add them to the queue. For an owned object, the handler maps the child to its owner (controller-runtime's `Owns()` does this by reading the ownerReference); for arbitrary relationships you supply a mapping function (`Watches(..., handler.EnqueueRequestsFromMapFunc(...))`). 5. **Work queue.** A deduplicating, rate-limited queue. It guarantees an item currently being processed is not handed to another worker, and that additions while an item is in flight cause it to be re-queued once afterwards. 6. **Workers.** A fixed pool (`MaxConcurrentReconciles`) pops keys and calls `Reconcile`. The reconciler reads from the informer cache, so a `Get` is a local map lookup, not an API call. 7. **Result handling.** Returning an error requeues with **exponential backoff**; `Result{Requeue: true}` requeues with backoff; `Result{RequeueAfter: d}` schedules after a fixed delay; a clean empty result drops the item and resets its backoff. ## Why the queue exists **Deduplication.** A single spec change can generate several events (spec, status, owned-object churn). Level-based reconciliation means one pass covers them all, so collapsing many keys into one pending item is not just an optimisation — it prevents redundant work proportional to event volume. **Serialization per object.** The queue ensures one key is processed by at most one worker at a time, so you do not need locks against yourself. Different objects still proceed in parallel. **Backpressure.** Reconciles can be slow (external API calls). If handlers reconciled inline, a slow reconcile would stall the watch stream, cause the informer to fall behind, and eventually force expensive re-lists. The queue decouples arrival rate from processing rate. **Retry policy in one place.** Transient failures — conflicts, throttling, a temporarily unreachable dependency — are normal. The rate limiter (typically an item-level exponential backoff from milliseconds to minutes, combined with an overall bucket limiter) turns them into polite retries instead of a hot loop hammering the API server. Without it, a permanently failing object can consume an entire cluster's API budget. **Delayed work.** `RequeueAfter` lets a controller poll something that has no watch — an external cloud resource, a certificate expiry, a timeout — without sleeping inside reconcile and occupying a worker. ## Consequences you must design around - **The cache is eventually consistent.** Immediately after you create an object, a cached read may not see it. Use deterministic names and treat AlreadyExists as success; if you must read your own write, do a direct API read or simply let the next reconcile see it. - **Never block in reconcile.** Sleeping or waiting on a long external call ties up a worker. Return `RequeueAfter` and check again. - **Concurrency is per controller.** Raising `MaxConcurrentReconciles` increases throughput but also API write pressure and the chance of conflicts; tune with queue depth and reconcile latency metrics. - **Watch scope drives memory.** Caching all Pods cluster-wide can be hundreds of megabytes; narrowing by namespace, label selector, or trimming cached fields is the usual fix on large clusters. - **Resync period re-enqueues everything** on a timer. Useful for drift correction, but a short period on a large object count is a self-inflicted load problem. - **Multiple replicas need leader election** so only one instance's queue is active; otherwise two controllers fight over the same objects. ## Observability The standard signals to watch are work-queue depth, queue latency (how long items wait), reconcile duration and error rate, and the number of requeues. Rising depth with flat throughput means you are worker-bound or a dependency is slow; a high error rate with growing backoff means one or more objects are permanently failing and dragging the loop. ## Interview framing Describe the chain in order, then justify the queue with the four properties — dedupe, serialize per key, backpressure, retry with backoff — and mention the cache's staleness as a real design constraint. That combination shows you have operated a controller, not only read the tutorial.
- Your controller reads an object it just created and gets NotFound. What is happening and how do you handle it?Reads go through the informer cache, which is updated asynchronously from the watch stream, so your own write may not be visible yet. The level-based fix is to not depend on it: use deterministic names, treat AlreadyExists as success, and let the next reconcile observe the object. If you truly must read-after-write, bypass the cache with a direct API client read, accepting the extra API-server load.
- How would you diagnose a controller that is falling behind?Look at work-queue depth and queue latency alongside reconcile duration and error rate. Growing depth with normal per-item latency means too few concurrent workers or too many objects; growing depth with long latency points at a slow dependency inside reconcile, often an external API call that should be made asynchronous or given a RequeueAfter. A high error rate with rising backoff means specific objects are permanently failing.
- What does raising MaxConcurrentReconciles risk?More parallel workers mean more simultaneous writes, so more optimistic-concurrency conflicts on shared objects, more API-server and etcd load, and more pressure on any external system the reconcile calls. It also increases memory and can trip client-side rate limits. Raise it deliberately while watching conflict rate, API latency and queue depth rather than as a default.
- Why do controller deployments use leader election?Running several replicas gives failover, but every replica has its own informers and queue, so without coordination they would reconcile the same objects concurrently and fight over writes. Leader election lets only the elected instance run its controllers; the others stand by and take over if the lease is not renewed.
The informer is a receptionist taking calls and writing names on a single to-do list; workers take names off the list. If the receptionist did each job personally, the phone line would back up.
saying these in an interview costs you the question
- Saying the controller queries the API server on every reconcile, missing the informer cache entirely
- Doing the reconcile work directly inside the informer event handler, blocking the watch stream
- Assuming every watch event results in exactly one reconcile, ignoring deduplication and coalescing
- Sleeping inside reconcile to wait for an external resource instead of returning RequeueAfter
- Treating the informer cache as strongly consistent and depending on read-after-write
- Running multiple controller replicas without leader election and expecting them to cooperate