What is a BOOKMARK event in the Kubernetes watch protocol, and which problem does it solve for a long-lived watch client?
answer
- BOOKMARK = resourceVersion only, no object change
- keeps an idle watcher's resume cursor fresh
- avoids 410 Gone and full relist on reconnect
- opt in with allowWatchBookmarks=true
- reduces relist stampede after API server restart
basics
~20 sA BOOKMARK is a periodic watch event carrying only an up-to-date resourceVersion and no object change. It lets an idle client advance its resume cursor, so after a disconnect it can resume instead of getting 410 Gone and re-listing everything.
solid answer
~60 sA watch client's resume cursor only advances when it receives an event. If it watches a rarely-changing resource — or a narrow selector that matches nothing for hours — its remembered resourceVersion gets older and older while the cluster moves on. When the connection eventually drops and the client tries to resume, that version has fallen out of the API server's history window, so it gets **410 Gone** and must do a full LIST. **BOOKMARK** events fix that. The API server periodically sends `{"type":"BOOKMARK", "object": {...}}` where the object carries only an updated `metadata.resourceVersion` and no meaningful content. The client stores that version and does nothing else. Now an idle watcher's cursor stays fresh, reconnects resume cleanly, and expensive relists — which are the costly operation for the API server, especially when many clients relist at once after a restart — are avoided. Clients opt in with `allowWatchBookmarks=true`; client-go's reflectors set it by default. A client that does not understand BOOKMARK must simply ignore unknown event types, which is why it is opt-in rather than always on.
code
bash · 2 lineskubectl get --raw '/api/v1/namespaces/default/pods?watch=true&allowWatchBookmarks=true&resourceVersion=182773'
# {"type":"BOOKMARK","object":{"kind":"Pod","apiVersion":"v1","metadata":{"resourceVersion":"191204"}}}go deeper
Know the one-line purpose: a heartbeat event carrying only an updated resourceVersion.
Explain the stale-cursor problem for quiet watches and how bookmarks prevent 410 Gone on reconnect.
Tie it to cluster-wide behaviour: relist stampedes after API server rollouts, and why aggressive selectors are safe with bookmarks.
Discuss it as backpressure design — keeping recovery cheap for thousands of clients so control-plane restarts degrade gracefully.
## The stale-cursor problem A watch resumes from a resourceVersion, and that value only moves forward when the client is delivered an event. Consider a controller watching a custom resource with maybe ten objects that change once a week, or a kubelet watching pods with a field selector for its own node on a quiet node. It lists, gets version 1000, and then hears nothing for hours while cluster-wide versions climb into the millions. The API server retains only a bounded window of recent history in its watch cache, and etcd compacts old revisions underneath. Version 1000 is long gone. When the client's connection is closed — proxies time watches out, API servers roll during upgrades, load balancers reset idle connections — the reconnect asks to resume from 1000 and gets `410 Gone: too old resource version`. The only recovery is a full LIST of the collection. Individually that is merely wasteful. In aggregate it is a real hazard: after an API server restart, thousands of clients reconnect, many of them stale, and a stampede of full LISTs hits the server precisely when it is least able to absorb it. ## What a bookmark is A BOOKMARK is a synthetic event in the same stream, with `type: "BOOKMARK"`. Its object is of the watched kind but is deliberately empty apart from `metadata.resourceVersion` — no spec, no status, nothing to apply. The contract is: the server is telling you "everything up to this version has been delivered to you; nothing you care about changed". The client's handling is trivial: update the stored resume version, do not touch the cache, do not invoke any handlers. Because it carries no object payload, the bandwidth cost is negligible even when sent to many watchers. ## Opting in and defaults Clients request them with the query parameter `allowWatchBookmarks=true` on the watch request. It is opt-in because older or hand-written clients might mishandle an unexpected event type — the general rule that clients should ignore unknown types is not universally obeyed. Bookmarks reached GA in Kubernetes v1.17 and are enabled by default in the standard client-go Reflector, so anything built on informers already benefits without configuration. The server sends them on its own schedule (roughly once a minute per watch, with jitter, and it may skip them when the stream is already busy since real events serve the same purpose). ## Why it matters in practice - **Fewer 410s and fewer relists.** Reconnects after transient network failures resume in place. - **Cheaper API server restarts and rollouts.** Reconnecting clients mostly resume rather than relist, flattening the thundering herd. - **Narrow selectors become safe.** Watching a very specific label or field selector no longer implies a stale cursor, so clients can filter aggressively to reduce traffic without paying for it on reconnect. ## What it does not do Bookmarks do not guarantee a watch can always be resumed — a long enough outage still ages out. They deliver no data and never trigger reconcile. And they are not a substitute for correct 410 handling: a client must still be able to relist and rebuild its cache, because bookmarks only make that path rarer.
- Should a controller's reconcile logic run when a BOOKMARK event arrives?No. A bookmark carries no object content and signals no change, so handlers must ignore it apart from recording the new resume resourceVersion. Triggering reconcile on bookmarks would generate pointless work across every watched resource roughly every minute.
- Do bookmarks make handling 410 Gone unnecessary?No. They make aged-out cursors much less likely, but a client that is disconnected or paused long enough will still fall outside the server's retained history window. Correct clients must always be able to re-LIST and rebuild their cache; bookmarks only reduce how often that expensive path is taken.
A tour guide who calls out the current page number every few minutes even when nothing new is read aloud, so anyone who steps out can rejoin at the right place instead of starting the book over.
saying these in an interview costs you the question
- Thinking a bookmark contains a changed object that should be applied to the cache
- Believing bookmarks make relisting unnecessary
- Assuming they are always on for every client rather than requested with allowWatchBookmarks
- Confusing them with periodic informer resyncs, which replay cached objects to handlers