A client calls the Kubernetes API with `?watch=true&resourceVersion=<rv>` and after some idle time receives HTTP 410 Gone with "too old resource version". Explain the contract the API server's watch endpoint offers and what the client is required to do.
answer
- watch = long-lived stream of ADDED/MODIFIED/DELETED/BOOKMARK
- resourceVersion = opaque cursor from etcd revision
- bounded watch cache + etcd compaction → 410 Gone
- recovery = relist, take list's rv, rewatch
- allowWatchBookmarks keeps the cursor fresh; rv=0 list may be stale
basics
~20 sA watch streams change events occurring after a given resourceVersion, which the server can only replay from a bounded in-memory window. Once that version ages out, the server returns 410 and the client must re-list to get a fresh resourceVersion and re-watch.
solid answer
~50 s`?watch=true` opens a long-lived streamed response of typed events — `ADDED`, `MODIFIED`, `DELETED`, `BOOKMARK` — each carrying an object whose `metadata.resourceVersion` is an opaque cursor derived from etcd's revision. The server sends events that occurred *after* the `resourceVersion` you supplied. It cannot do that indefinitely: history lives in a bounded per-resource watch cache, and etcd compacts old revisions. When your cursor falls out of that window, the server returns **410 Gone**. That is a normal, expected condition after a disconnect, a slow consumer, or a burst of churn — not a failure. The required recovery is **re-LIST, then re-watch**: take the `resourceVersion` from the list's `metadata`, and restart the watch from it. Never resume from a stale cursor, and never diff by polling. Two things reduce 410s: `allowWatchBookmarks=true`, which periodically ships a BOOKMARK event carrying a fresh resourceVersion with no object payload, and simply consuming events promptly. Informers implement this loop for you.
code
bash · 6 lineskubectl get --raw \
'/api/v1/namespaces/default/pods?watch=true&allowWatchBookmarks=true&resourceVersion=0'
# LIST first to obtain a valid cursor
kubectl get --raw '/api/v1/namespaces/default/pods?limit=500' \
| jq -r '.metadata.resourceVersion'go deeper
Know that a watch is a streaming endpoint delivering ADDED/MODIFIED/DELETED events and that 410 means "start over with a list".
Explain resourceVersion as an opaque cursor and the relist-then-rewatch recovery, and know that history is bounded rather than infinite.
Cover the watch cache and etcd compaction as the two bounds, bookmarks, selectors and pagination as cost controls, and why level-triggered reconciliation follows from the delivery guarantees.
Reason about watch load across a fleet — relist storms after a control-plane restart, the memory cost of unpaginated lists, and standardising every team on informer libraries rather than hand-rolled watch loops.
## What a watch is on the wire `GET /api/v1/namespaces/default/pods?watch=true&resourceVersion=12345` does not return a collection. It returns a long-lived HTTP response — chunked JSON objects, or a framed protobuf stream, and over HTTP/2 a single stream on the shared connection — where each frame is a `WatchEvent`: ``` {"type":"ADDED","object":{...}} {"type":"MODIFIED","object":{...}} {"type":"DELETED","object":{...}} {"type":"BOOKMARK","object":{"metadata":{"resourceVersion":"12420"}}} ``` The stream stays open until the client closes it, the server times it out (there is a randomised request timeout so thousands of watches do not all reconnect at once), or an error terminates it. ## resourceVersion is a cursor, not a number you own Every object carries `metadata.resourceVersion`, and every list carries one in its `metadata`. It is derived from etcd's global revision, but the API contract declares it **opaque**: do not parse it, compare it numerically across resource types, or persist it as if it were meaningful. Its only supported use is handing it back to the server. The semantics of the `resourceVersion` parameter differ by request: - **Watch with an rv** — send events strictly *after* that revision. - **Watch with no rv** — the server sends the current state as a series of synthetic `ADDED` events, then continues live. Convenient, but it hides whether you are seeing history or the present. - **`resourceVersion=0` on a LIST** — "any reasonably recent cached version", served from the watch cache without a quorum read. Cheap and the informer default, but it may be **stale**, so it is wrong for read-then-write correctness. - **LIST with no rv** — a quorum read through etcd: authoritative, expensive. ## Why 410 Gone happens The API server keeps a per-resource ring buffer of recent events in its **watch cache** (on the order of a thousand entries, sized adaptively), and etcd itself **compacts** revisions older than a few minutes. Both bound how far back a watch can start. If your cursor is older than the oldest retained event, the server has no way to tell you what you missed — and silently skipping changes would break the correctness of every controller — so it must fail loudly: ``` 410 Gone: too old resource version: 12345 (18631) ``` Common causes: the client was disconnected long enough for the window to roll; the consumer was too slow and the server dropped it rather than buffering without limit; a burst of churn (a big rollout, a namespace deletion) blew through the buffer quickly; or an etcd compaction ran. ## The required client behaviour **Relist and rewatch.** 1. LIST the resource; take `metadata.resourceVersion` from the list response. 2. Reconcile local state against that list — objects deleted while you were disconnected appear only as absences, which is precisely why you cannot skip the list step. 3. Start a new watch from that resourceVersion. Doing anything else is a bug: resuming from the stale cursor loops on 410, and polling instead of watching turns O(changes) into O(objects × frequency) load on the API server. Use `?allowWatchBookmarks=true`. The server then periodically emits a `BOOKMARK` event whose object carries only a current `resourceVersion` and no payload. Even when nothing your watch cares about is changing — say you filtered to one namespace while the rest of the cluster churns — bookmarks keep your cursor near the head of the stream, so a brief disconnect resumes cleanly instead of landing outside the window. Bookmarks are cheap and dramatically reduce relist storms after a control-plane restart. ## Filtering, chunking and cost Watches accept `labelSelector` and `fieldSelector`, evaluated server-side, which is the correct way to reduce both network and client CPU. Large initial LISTs should use `limit` and `continue` for pagination, and clients that only need names and labels can request metadata-only responses (`Accept: application/json;as=PartialObjectMetadataList;g=meta.k8s.io;v=v1`) so the server never serialises full objects. An unpaginated LIST of every Pod in a large cluster is the single most common way to spike API-server memory. ## What a watch does not give you - **Not a durable queue.** No acknowledgement, no replay beyond the window, no delivery guarantee across a disconnect. - **Not necessarily every intermediate state.** Coalescing means you may not observe each transition; you are guaranteed to converge on the latest state, which is why controllers are written as level-triggered reconcilers over the current object rather than as edge-triggered handlers that assume they saw each step. - **Not ordered across resources.** Ordering holds per watched resource, not globally. Client-go's informer machinery implements this contract — list, watch, bookmark handling, 410 recovery, resync — which is why almost nothing should hand-roll a watch loop in production.
- What does allowWatchBookmarks=true buy you?The server periodically sends a BOOKMARK event carrying only an up-to-date resourceVersion and no object payload. That keeps the client's cursor near the head of the stream even when nothing it watches is changing, so a short disconnect can resume instead of falling outside the retained window and forcing a full relist. On a busy cluster with many narrowly-filtered watchers it removes most 410-driven relist storms.
- Why is a LIST with resourceVersion=0 cheaper, and when is it wrong?It is served from the API server's in-memory watch cache rather than a quorum read from etcd, so it avoids consensus latency and etcd load — which is why informers use it on startup. It may return slightly stale data, so it is wrong whenever you are about to make a decision or a write that depends on having the newest state; use a plain LIST (quorum read) for read-then-write correctness.
- Why are Kubernetes controllers written as level-triggered reconcilers rather than event handlers?A watch does not guarantee you observe every intermediate state: events can coalesce, and after a 410 the client relists and sees only the current state. A controller that reacts to transitions would silently diverge. Reconciling the observed current object against desired state converges correctly no matter which events were missed, and makes the relist path just another input.
It is a live broadcast with a short rewind buffer, not a recorded archive. If you were away longer than the buffer, the station cannot replay the gap — you must ask for the current state and start following again.
saying these in an interview costs you the question
- Treating 410 Gone as a server bug rather than the defined contract
- Resuming a watch from the stale resourceVersion after a 410, looping forever
- Parsing or arithmetic on resourceVersion, or comparing it across resource types
- Assuming resourceVersion=0 gives a strongly consistent read
- Replacing watches with polling loops, or assuming a watch is a durable message queue