A Kubernetes controller's reconcile function fails for one object because a downstream API returns errors. Why does the controller use a rate-limited work queue instead of simply retrying in a loop?
answer
- queue holds keys, not snapshots
- dedupe + single-flight per key
- AddRateLimited on error, Forget on success
- exponential per-item × global token bucket
- RequeueAfter ≠ error; resync is the floor
basics
~20 sThe work queue deduplicates keys, guarantees one worker per key at a time, and requeues failures with exponential backoff. A tight retry loop would block other objects, hammer the API server, and turn one broken object into a control-plane load problem.
solid answer
~50 sControllers enqueue **keys** (`namespace/name`), not objects. The client-go work queue gives three properties: 1. **Deduplication** — a key already queued is not added twice, so a burst of events for one object collapses into a single sync that reads the latest state from cache. 2. **Single-flight per key** — a key being processed is held aside until `Done()`, so two workers never reconcile the same object concurrently. That is why worker parallelism is safe. 3. **Rate-limited requeue** — on error the worker calls `AddRateLimited`, which schedules a retry with per-item exponential backoff (commonly 5ms doubling to ~1000s) plus a global token-bucket cap. On success `Forget()` resets the item's backoff. So a permanently broken object retries ever more slowly instead of spinning, other keys keep flowing through the remaining workers, and the API server is protected. Periodic resync still re-enqueues everything, so nothing is abandoned forever, and metrics like queue depth, work duration and retries expose a wedged controller.
go deeper
Know that failures are retried with a growing delay rather than in a tight loop, so one bad object cannot block everything.
Explain keys versus objects, deduplication, one-worker-per-key, and AddRateLimited/Forget/Done.
Discuss the exponential-plus-token-bucket limiter, requeue-after versus error, and diagnosing wedged or hot-looping controllers from workqueue metrics.
Frame the queue as the control plane's admission control — bounding self-inflicted API load — and set expectations for operators the platform will host, including idempotency and no blocking calls in reconcile.
## Keys, not objects The first design decision is that the queue holds **keys** — the string `namespace/name` — rather than object snapshots. When a worker pops a key it re-reads the object from the informer cache, so it always acts on the newest known state. A stale snapshot from three events ago is never processed. This is the queue-level expression of level-triggered reconciliation: the event tells you *who* to look at, never *what* to do. ## The three guarantees **Deduplication.** Adding a key that is already pending is a no-op. During a rollout, a single ReplicaSet may receive dozens of Pod events per second; all of them collapse into one pending key and therefore one sync that sees the final state. Without this, event volume would translate directly into API write volume. **Single-flight per key.** The queue tracks items currently being processed. If a key is re-added while in flight, it is remembered as "dirty" and re-queued only after the worker signals `Done()`. Consequently no two workers ever reconcile the same object simultaneously — the reason a controller can run, say, five workers without locking. It also means a change arriving mid-sync is not lost: it triggers exactly one more sync afterwards. **Rate-limited requeue.** Errors are expected: a conflicting write, a not-yet-created dependency, a throttled cloud API. The standard handler is: - error → `AddRateLimited(key)`: retry after a delay drawn from the rate limiter; - success → `Forget(key)`: clear that key's failure count so its backoff resets; - always → `Done(key)`. The default limiter is a *max-of* two limiters: a **per-item exponential** one (base 5ms, doubling, capped near 1000s) and a **global token bucket** (e.g. 10 qps, burst 100). The per-item part stops one sick object from spinning; the bucket stops a thousand simultaneously-sick objects from spiking the API server. ## Why a naive retry loop is wrong Imagine `for { if err := reconcile(obj); err == nil { break } }`: - **Head-of-line blocking.** The worker is stuck on one object; every other object waits behind it. One broken ConfigMap reference stalls the whole controller. - **Amplification.** A failure that returns instantly (a 409 conflict, a validation error) retries thousands of times per second. The API server sees a synthetic DDoS from its own control plane. This is the classic "hot loop" incident, and it is why default limiters exist at all. - **Lost concurrency safety.** Without a queue tracking in-flight keys you need your own locking to prevent two workers touching one object. - **No observability.** Queue depth, latency (`workqueue_queue_duration_seconds`), processing time (`workqueue_work_duration_seconds`), retry counts and `unfinished_work_seconds` are all queue-level metrics; a bare loop exposes none of them. ## Backoff versus requeue-after Two different situations deserve different handling, and conflating them is a common mistake: - **Error** — something went wrong; use rate-limited requeue so repeated failure decays. - **Not an error, just not done yet** — for example waiting for a cloud volume to finish attaching. Return success but requeue after a fixed interval (`AddAfter` / `RequeueAfter`). Returning a fake error to force a retry pollutes error metrics and drives exponential backoff on a perfectly healthy object, so the poll interval grows to many minutes and the resource looks hung. ## Resync as the floor Even a key whose backoff has grown to fifteen minutes is re-enqueued by the informer's periodic resync (typically every 10–30 minutes, or by any new event on the object or its children). Nothing is permanently dropped. Combined with level-triggering, this means the worst consequence of misconfigured backoff is slow convergence, not lost convergence. ## Operating and debugging Symptoms map cleanly to metrics: *depth grows without bound* → reconciles are slower than event arrival, so add workers or make the sync cheaper; *depth is small but adds/retries are huge* → a hot loop, usually a controller writing status on every sync and thereby waking itself (a self-triggering feedback cycle — the fix is to write only when the status actually changed); *work duration spikes* → a blocking external call inside the sync, which should be moved out or bounded by a timeout so one slow dependency does not consume all workers. ## Applying this to your own controllers Everything above is the client-go/controller-runtime machinery that Kubernetes' built-in controllers use and that operators inherit for free. The obligations on your reconcile function are therefore: make it idempotent, keep it short and non-blocking, return errors only for genuine failures, call `Forget` on success (controller-runtime does this for you), and never mutate objects taken from the informer cache — copy first, because that cache is shared across all controllers in the process.
- A controller's reconcile is waiting for an external volume to become ready. Should it return an error so the key is retried?No. Returning an error inflates error metrics and applies exponential backoff to a healthy object, so the poll interval balloons and the resource appears stuck. Return success with an explicit requeue-after of a sensible interval, which keeps the retry cadence constant and truthful. Reserve errors for actual failures you want to decay and alert on.
- Controller metrics show a tiny queue depth but an enormous add rate. What is the likely cause?A self-triggering hot loop: the controller writes to an object it also watches — most often patching status unconditionally on every sync — so its own write generates an event that re-enqueues the key. The fix is to compare the computed status against the observed status and skip the write when nothing changed, which breaks the feedback cycle.
A hospital triage desk: each patient appears once on the list, one doctor sees them at a time, and someone who needs to come back is given a later appointment rather than being seen again immediately while everyone else waits.
saying these in an interview costs you the question
- Saying the queue stores object copies, so a worker acts on stale state.
- Claiming two workers may reconcile the same key concurrently, requiring your own mutex.
- Returning an error purely to schedule a retry for a normal 'not ready yet' condition.
- Believing an item that exhausts backoff is dropped forever — periodic resync re-enqueues it.
- Mutating an object obtained from the informer cache instead of deep-copying it.