skip to content

Watch and Informer Mechanics

Controllers stay in sync through list-watch streams rather than polling, using resourceVersion for ordering, bookmarks for resumption and shared informer caches. Senior rounds probe it for optimistic concurrency and how the cluster scales event delivery.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

5

Your code updates a Kubernetes object and the API server responds with HTTP 409 Conflict saying the object has been modified. What causes that, and what are the correct ways to handle it?

level: middleimportance: must knowfreq 45%

answer

  1. resourceVersion = optimistic-concurrency token
  2. 409 = someone wrote first; nothing was applied
  3. retry must re-GET, not reuse the stale object
  4. patch or server-side apply avoids whole-object conflicts
  5. never blank resourceVersion to force a write

basics

~20 s

Kubernetes updates use optimistic concurrency: your object carries the resourceVersion you read, and the server rejects the write if the live object has changed since. Handle it by re-reading the object fresh, reapplying your change, and retrying — or by sending a patch instead of a full update.

solid answer

~60 s

Every object carries a `resourceVersion`. A full update (PUT) includes it, and the API server accepts the write only if it still matches the live object; otherwise it returns **409 Conflict** with "the object has been modified; please apply your changes to the latest version and try again". That is optimistic concurrency: no locks, no blocking, just a rejected write when someone else got there first — often another controller, or your own stale informer cache. Two correct responses: 1. **Retry on conflict.** Re-GET the object *from the API server, not the cache*, reapply your mutation to the fresh copy, and update again. `retry.RetryOnConflict` in client-go does exactly this with backoff. Never loop retrying the same stale object — it will fail forever. 2. **Patch instead of update.** A strategic-merge or JSON patch expresses only the change, so unrelated concurrent edits do not conflict at all. Server-side apply is the declarative form and conflicts only on fields another manager owns. What is *not* acceptable: stripping or blanking resourceVersion to force the write. That disables the safety check and silently discards the other writer's changes.

code

go · 14 lines
go
err := retry.RetryOnConflict(retry.DefaultRetry, func() error {
    live, err := clientset.AppsV1().Deployments(ns).Get(ctx, name, metav1.GetOptions{})
    if err != nil {
        return err
    }
    live.Spec.Template.Labels["rollout"] = revision
    _, err = clientset.AppsV1().Deployments(ns).Update(ctx, live, metav1.UpdateOptions{})
    return err
})

// narrow, independent change: patch instead, no resourceVersion involved
patch := []byte(`{"metadata":{"labels":{"team":"payments"}}}`)
_, err = clientset.AppsV1().Deployments(ns).Patch(
    ctx, name, types.StrategicMergePatchType, patch, metav1.PatchOptions{})

go deeper

for a junior

State the rule: the object changed since you read it, so re-read and try again; never delete resourceVersion.

for a middle

Explain optimistic concurrency, the retry-with-fresh-read loop, and when a patch removes the conflict entirely.

for a senior

Discuss why controllers see conflicts constantly given eventually-consistent caches, choosing patch type versus update, bounding retries, and status-subresource separation.

for a principal

Frame it as concurrency policy across the platform: which fields are independent enough to patch, where read-modify-write must be preserved for correctness, and how contention on hot objects shapes controller design.

## The mechanism Kubernetes never takes locks on objects. Instead every object has a `metadata.resourceVersion` that changes on every write. When you send a full update, you send the object you read, resourceVersion included. The API server compares it with the stored version: - match: the write is applied and the version advances; - mismatch: **409 Conflict**, nothing is written. This is optimistic concurrency control, and it converts a lost-update bug into a visible error. Without it, a read-modify-write pair from two controllers would silently clobber one another: both read version 5, both write their own full object, and the second erases the first's field. ## Why it fires so often in controllers Controllers usually read from an informer cache, which is eventually consistent. If the object changed a moment ago, your cached copy already holds an outdated resourceVersion, so the very first update attempt conflicts. Popular objects also attract many writers — status updates, finalizer additions, label edits from several operators. Conflicts on hot objects are normal, not exceptional; the design assumption is that you retry. ## Handling it well **Retry with a fresh read.** The essential detail is that each retry must re-fetch the object from the API server (a live GET, not the informer's Lister) and reapply the mutation to that fresh copy, because only that gives you a current resourceVersion. client-go's `retry.RetryOnConflict(retry.DefaultRetry, fn)` implements the loop; you supply a function that gets, mutates and updates. Bound the retries — infinite loops on a permanently contended object are a real outage mode — and let the controller's workqueue requeue with backoff if retries are exhausted. **Prefer a patch.** A full update carries the entire object and therefore conflicts on any concurrent change anywhere in it. A patch carries only your intent: - *strategic merge patch* — the Kubernetes-aware default for built-in types, merging lists by key rather than replacing them; - *JSON merge patch* — simple object merge, replaces lists wholesale; - *JSON patch* — an operation list, and the only one that can express a test-then-set precondition; - *server-side apply* — declarative, with per-field ownership; it conflicts only when you contest a field another field manager owns. Patches usually need no resourceVersion at all, so they are the pragmatic choice for narrow mutations like adding a label, a finalizer or an annotation. ## Know when you still want the conflict If your change depends on the value you read — incrementing a counter, appending to a list, or making a decision from the current status — a blind patch can corrupt data. There you *want* optimistic concurrency: read, compute, update with resourceVersion, and let 409 force a recomputation on the new value. ## Anti-patterns - **Clearing resourceVersion to force the write.** It turns the update into an unconditional overwrite and reintroduces lost updates. Reviewers should treat it as a defect. - **Retrying without re-reading.** Guaranteed to fail repeatedly and to hammer the API server. - **Retrying the wrong errors.** 409 Conflict means retry; 422 Invalid, 403 Forbidden and 404 Not Found will not improve with repetition. - **Updating spec and status in one call for resources with a status subresource.** They are separate endpoints; mixing them either drops changes or produces avoidable conflicts. - **Unbounded retry loops on a contended object**, which starve the workqueue. ## The mindset Treat 409 as normal control flow in a system where many autonomous loops write concurrently. The question a strong candidate answers is not "how do I stop it happening" but "is my change independent enough to patch, or dependent enough that I must re-read and recompute".

  • When should you deliberately keep using a full update with resourceVersion instead of switching to a patch?
    Whenever the new value depends on the value you read — incrementing a counter, appending to a list, or deciding from current status. There the conflict is the feature: it forces you to recompute against the latest state instead of overwriting a change you never saw. Blind patches in that situation produce silent data corruption.
  • Why does clearing metadata.resourceVersion before an update make the error disappear, and why is that wrong?
    An empty resourceVersion tells the API server to skip the precondition, so the write is applied unconditionally. The 409 stops appearing because the safety check is gone, not because the race is gone — any concurrent change is now silently overwritten, which is exactly the lost-update bug the mechanism exists to prevent.

Editing a wiki page: you saved based on revision 5, but revision 6 already exists, so the site refuses and asks you to merge against the newest text rather than silently overwriting someone's paragraph.

saying these in an interview costs you the question

  • Blanking or removing resourceVersion to make the update succeed
  • Retrying the update with the same stale object instead of re-reading
  • Reading the retry copy from the informer cache rather than the API server
  • Calling 409 a server bug or an API server overload symptom
  • Retrying non-retryable errors such as 422 Invalid or 403 Forbidden the same way
  • Unbounded retry loops on a hotly contended object

context

open as a page

Describe how a Kubernetes client keeps an up-to-date view of a set of objects using the list-then-watch pattern, and what the resourceVersion value returned by the API server means in that flow.

level: middleimportance: must knowfreq 46%

basics

~20 s

The client LISTs the objects once, notes the collection's resourceVersion, then opens a WATCH starting after that version and receives ADDED/MODIFIED/DELETED events as a stream. resourceVersion is an opaque cursor into the API server's change history, not a number to compare or interpret.

open as a page

What is a shared informer in the Kubernetes client libraries, and what are the consequences of reading objects from its local cache instead of from the API server?

level: seniorimportance: should knowfreq 38%

basics

~20 s

An informer runs one list-watch per resource type, keeps the objects in an in-memory indexed store, and fires handlers on changes; shared means many controllers in the process reuse that one stream and cache. Cache reads are fast and free but eventually consistent, so they can be stale and must never be mutated.

open as a page

What is a BOOKMARK event in the Kubernetes watch protocol, and which problem does it solve for a long-lived watch client?

level: seniorimportance: nice to knowfreq 20%

basics

~20 s

A BOOKMARK is a periodic watch event carrying only an up-to-date resourceVersion and no object change. It lets an idle client advance its resume cursor, so after a disconnect it can resume instead of getting 410 Gone and re-listing everything.

open as a page

As a cluster grows to thousands of nodes and dozens of controllers, watch traffic becomes a significant load on the control plane. How do you reason about that cost and reduce it on the client side?

level: principalimportance: nice to knowfreq 22%

basics

~20 s

Cost scales with objects times watchers times change rate, plus the memory each client caches. Reduce it by sharing informers per process, scoping watches with label and field selectors, trimming cached objects, avoiding short resyncs and relist storms, and keeping fast-churning data out of watched objects.

open as a page