skip to content

A controller's reconcile fails for one object because a dependency is temporarily unavailable, and the controller then hammers the Kubernetes API with retries. Explain how requeue and rate-limited backoff are supposed to work in a controller, and what the different reconcile return values mean.

level: middleimportance: should knowfreq 36%

answer

  1. err -> requeue with exponential backoff
  2. RequeueAfter -> fixed delay, no error metric
  3. empty result -> forget item, reset backoff
  4. Limiter = per-item exponential + shared token bucket
  5. Permanent errors go to conditions, not retries

basics

~20 s

Returning an error requeues the key through a rate limiter with exponential backoff, so retries slow from milliseconds to minutes. Returning RequeueAfter schedules a re-check at a fixed delay for polling. Returning an empty result with no error drops the item and resets its backoff. Never retry in a loop inside reconcile.

solid answer

~50 s

Controller-runtime gives three outcomes from `Reconcile`: - **`return ctrl.Result{}, err`** — failure. The key goes back on the work queue through the rate limiter, whose per-item exponential backoff typically starts around 5 ms and doubles up to about 1000 s, plus an overall bucket limiter across items. This is the right response to transient errors: conflicts, throttling, an unreachable dependency. - **`return ctrl.Result{RequeueAfter: d}, nil`** — not failed, just not finished. Re-check after a fixed delay; used to poll external state that has no watch, or to re-check a certificate expiry. - **`return ctrl.Result{}, nil`** — converged. The item leaves the queue and its backoff counter is **forgotten**, so the next failure starts from a short delay again. The anti-patterns are retrying in a `for` loop or sleeping inside reconcile, both of which hold a worker and bypass backoff. Also avoid returning an error for an expected non-error state — it inflates error metrics and produces unpredictable delays where a deliberate `RequeueAfter` is clearer.

code

go · 21 lines
go
// transient: dependency not reachable right now
if err := r.callExternal(ctx); err != nil {
    return ctrl.Result{}, err // exponential backoff via the rate limiter
}

// still provisioning: poll again without signalling failure
if !ready {
    return ctrl.Result{RequeueAfter: 30 * time.Second}, nil
}

// permanent user error: surface it, do not retry-storm
if app.Spec.StorageClass == "" {
    meta.SetStatusCondition(&app.Status.Conditions, metav1.Condition{
        Type: "Ready", Status: metav1.ConditionFalse,
        Reason: "InvalidSpec", Message: "spec.storageClass must be set",
    })
    return ctrl.Result{}, r.Status().Update(ctx, &app)
}

// converged
return ctrl.Result{}, nil

go deeper

for a junior

Know the three return values and that failures are retried automatically with increasing delay rather than immediately.

for a middle

Explain the per-item exponential backoff plus shared bucket limiter, when to use RequeueAfter instead of an error, and that success forgets the item.

for a senior

Add the judgement: separate transient from permanent failures, surface user errors in conditions and Events, and tune concurrency and requeue intervals against queue depth, latency and API-server load.

for a principal

Treat retry policy as a fleet-level capacity concern — how thousands of failing objects interact with API-server and dependency capacity, and what defaults keep controllers safe neighbours during a widespread outage.

## Where retries live A controller does not retry inside reconcile. Retry is a property of the **work queue** that feeds it. The queue is rate-limited: when an item is added after a failure, the limiter decides how long to hold it before making it available to a worker again. Keeping retry outside the reconcile body means the worker is free in the meantime, backoff state survives across invocations, and every controller in the process shares one consistent policy. ## The default rate limiter The standard controller rate limiter is the maximum of two components: 1. **Per-item exponential backoff.** Each key has a failure counter. Delay roughly doubles per consecutive failure — commonly from about 5 milliseconds up to a cap around 1000 seconds (roughly 16 minutes). The counter is cleared when the item is *forgotten*, which happens on a successful reconcile. 2. **An overall bucket limiter.** A token bucket (commonly on the order of tens of items per second with a burst) across all items in that queue, so a mass failure — say 5,000 objects all failing because one dependency is down — cannot produce a thundering herd against the API server. The combination is what makes a controller a good API-server citizen: a single broken object retries with growing patience, and a fleet-wide failure is smoothed. ## The three return values in detail **Error.** `return ctrl.Result{}, err` logs the error, increments the controller's error metric, and requeues with backoff. Use it for anything genuinely wrong or transient. Note that a returned error together with a non-zero `Result` is contradictory — the error path wins, so pick one. **RequeueAfter.** `return ctrl.Result{RequeueAfter: 30*time.Second}, nil` is a *scheduled* re-check with no error semantics: no error metric, no exponential growth, exactly the delay you asked for. Correct uses: polling a cloud API that has no watch; re-checking after a TTL or expiry; implementing a timeout for a phase that is still progressing. Because you control the delay, you also own the cost — a 1 s requeue across 10,000 objects is a self-inflicted load test. **Empty result, no error.** `return ctrl.Result{}, nil` means "nothing more to do right now". The item is forgotten and its backoff resets. The controller will hear about the object again through its watches, or at the next resync. This should be the common path for a converged object. (`Result{Requeue: true}` requeues through the rate limiter without an error; it is largely equivalent to an immediate backoff-governed retry and is being phased out in favour of `RequeueAfter`.) ## Distinguishing transient from permanent Backoff handles transient failures well and permanent ones badly: a spec that references a nonexistent storage class will fail forever, retrying every 16 minutes and lighting up error metrics with no chance of success. Good controllers separate the two: - **Transient** (network, conflict, throttling, dependency starting up): return the error and let backoff work. - **Permanent / user error** (invalid spec, missing referenced object that only the user can create): do not treat it as a controller error. Record it in `status.conditions` with a clear reason, emit an Event so `kubectl describe` shows it, and return without an error — possibly with a modest `RequeueAfter` if the situation could resolve on its own, or simply wait for the user's edit, which arrives as a watch event. That distinction is the main thing separating a controller that is pleasant to operate from one whose dashboards are permanently red. ## Interaction with conflicts Optimistic-concurrency conflicts on update are normal and expected, especially with concurrency above one. The idiomatic handling is to return the error and let the next pass re-read fresh state — not to retry the write in a tight loop with stale data. If conflicts dominate your error rate, the usual causes are too many concurrent reconciles on shared objects, writing the whole object instead of patching, or two controllers writing the same fields. ## Observing the behaviour The signals to watch: reconcile error rate by controller, reconcile duration, work-queue depth, queue latency, and retry/requeue counts. A queue whose depth is flat but whose latency grows suggests items sitting in long backoff. Steady low-rate errors on the same few keys usually means permanent failures that should have been surfaced as conditions instead. ## Interview framing Name the three return values and their exact semantics, describe the two-part rate limiter, and then make the judgement point: retries belong in the queue rather than inside reconcile, and permanent user errors belong in status and Events rather than in the retry loop.

  • When should a controller NOT return an error even though it cannot finish its work?
    When the cause is a user or configuration problem it cannot fix — an invalid field, a referenced Secret that does not exist yet, a storage class that is not installed. Retrying cannot succeed, so returning an error only inflates error metrics and burns retries. The right move is to write a clear condition and Event describing what the user must change, then return cleanly; the user's edit arrives as a watch event and triggers a fresh reconcile.
  • Why is retrying inside the reconcile function with a loop and sleep a bad idea?
    It occupies a worker for the whole duration, reducing throughput for every other object, and it bypasses the queue's rate limiter, so it can hammer a dependency that is already struggling. It also loses the retry state on restart. Returning an error or a RequeueAfter delegates timing to the queue, which keeps the worker free and applies consistent backoff.
  • What resets an item's exponential backoff?
    A successful reconcile: when the reconciler returns without an error and without requesting a requeue, the queue forgets the item and clears its failure counter. The next failure therefore starts again at the minimum delay rather than continuing from the previous long interval.

saying these in an interview costs you the question

  • Implementing retry loops with sleeps inside reconcile instead of returning an error or RequeueAfter
  • Returning an error for expected states like 'dependency not created yet', producing permanent error noise
  • Assuming a failed reconcile retries immediately and forever at the same rate, with no backoff
  • Setting a very short RequeueAfter across thousands of objects and then blaming the API server for load
  • Treating optimistic-concurrency conflicts as bugs and force-writing instead of re-reading on the next pass
  • Thinking backoff state lives in the reconciler rather than in the work queue's rate limiter

context