As a cluster grows to thousands of nodes and dozens of controllers, watch traffic becomes a significant load on the control plane. How do you reason about that cost and reduce it on the client side?
answer
- cost = change rate x watchers, plus cache memory, plus relist
- share informers: one per type per process
- selectors: namespace, label, field (nodeName pattern)
- trim managedFields out of the cache
- LIST storms after apiserver restart are the real risk
basics
~20 sCost scales with objects times watchers times change rate, plus the memory each client caches. Reduce it by sharing informers per process, scoping watches with label and field selectors, trimming cached objects, avoiding short resyncs and relist storms, and keeping fast-churning data out of watched objects.
solid answer
~1 minModel the cost in three parts. **Fan-out**: every change to a watched object is serialized and pushed to each interested watcher, so cost is roughly object change rate times number of watchers. **Client memory**: each informer caches every object in its scope, so watching all Pods or Secrets cluster-wide can be hundreds of megabytes per process. **Recovery**: a full LIST is far more expensive than a watch, and a control-plane restart can trigger thousands of simultaneous relists. Client-side levers, in rough order of payoff: - **Share informers** — one per resource type per process; duplicated informers are the most common self-inflicted multiplier. - **Scope the watch** — namespace scope, plus label or field selectors (the kubelet pattern: `spec.nodeName=<me>`), so both fan-out and cache size shrink. - **Trim what you cache** — a transform function stripping managedFields and bulky annotations can cut informer memory substantially. - **Do not resync aggressively** — a one-minute resync re-reconciles everything for no consistency benefit; minutes-to-hours, or zero, is right. - **Reduce write churn** — controllers that rewrite status every loop generate the events everyone else pays for; only write when something changed. - **Allow bookmarks** so reconnects resume instead of relisting. Server-side capacity work (watch cache sizing, request prioritisation, API server scaling) belongs with the API server itself; the client-side discipline above is what application and platform teams control.
code
go · 15 linesfactory := informers.NewSharedInformerFactoryWithOptions(
clientset,
30*time.Minute, // long resync: safety net, not a consistency mechanism
informers.WithNamespace("payments"),
informers.WithTweakListOptions(func(o *metav1.ListOptions) {
o.FieldSelector = "spec.nodeName=" + nodeName
o.LabelSelector = "app.kubernetes.io/part-of=checkout"
}),
informers.WithTransform(func(obj interface{}) (interface{}, error) {
if m, err := meta.Accessor(obj); err == nil {
m.SetManagedFields(nil)
}
return obj, nil
}),
)go deeper
Know that watches are not free and that informers cache whole objects in memory; do not open a watch per goroutine.
Apply the concrete levers: shared informers, namespace and selector scoping, sensible resync intervals.
Quantify the cost model, diagnose a controller amplifying writes, and tune caching and reconnect behaviour under load.
Set platform-wide policy — which components may watch cluster-wide, what churn a new custom resource may generate, and when data does not belong in the Kubernetes API at all.
## Where the cost actually is A watch is cheap while idle and expensive under churn. Three distinct costs matter. **Fan-out on change.** When an object changes, the API server must deliver it to every watcher whose scope matches. Serialization and encoding dominate, and it is per watcher — a hundred controllers watching Pods means a hundred deliveries per pod change. High-churn objects therefore have super-linear blast radius: a controller that writes status on every reconcile multiplies its own load onto every other watcher of that type. **Client memory.** An informer holds the full object for everything in its scope. On a large cluster, all Pods can be gigabytes across processes; Secrets are worse per byte and more sensitive. Memory pressure in controllers then causes slow event draining, which increases lag and makes the cache staler — a feedback loop. **Recovery.** LIST is the expensive operation: it materialises and serializes a whole collection. It is fine occasionally, catastrophic in a herd. Anything that causes many clients to relist at once — an API server rollout, a network partition healing, aged-out resume cursors — is the real scaling event to design against. ## Client-side levers 1. **Share, don't duplicate.** A SharedInformerFactory gives one watch and one cache per resource type per process. Multiple independent informers for the same type inside one binary is a common and avoidable multiplier, especially when several controllers are bundled into one manager. 2. **Narrow the scope.** Namespace-scoped informers, label selectors, and field selectors cut both delivery and memory. The canonical example is the kubelet watching only pods with `spec.nodeName` equal to its own node — without that, every node would receive every pod change in the cluster. Bookmarks make narrow selectors safe by keeping the resume cursor fresh despite low event volume. 3. **Trim objects before caching.** Transform functions applied on the way into the store can drop `managedFields`, last-applied annotations and other bulk that controllers never read, cutting memory materially with no behaviour change. 4. **Right-size resync.** Resync replays the cache through handlers; it adds reconcile work without refreshing anything from the server. Short intervals multiply CPU across every object for no consistency benefit. Prefer long intervals or none, and rely on correct event handling plus explicit requeues. 5. **Reduce write amplification.** Controllers should write only when the computed state differs from what is stored — no unconditional status updates each pass, no timestamp fields that change every reconcile. Every avoided write is a delivery avoided at every watcher. 6. **Keep high-frequency data out of the API.** Metrics, heartbeats and per-request state do not belong in objects everyone watches. Lease objects exist precisely to keep node heartbeats small and separate from the Node object; the general principle is to isolate churn into small, narrowly watched resources. 7. **Handle failures gracefully.** Exponential backoff with jitter on reconnect, bounded retry, and correct 410 handling turn a potential stampede into a spread-out recovery. ## The judgement part The interesting decisions are architectural. Do you build one controller with a broad watch or several with narrow ones? Should a platform component watch all namespaces, or run per-tenant with scoped caches? Is a new custom resource's status going to be written every few seconds by design, and who will pay for that? Should this data live in the Kubernetes API at all, or in a purpose-built store, given that the API is a coordination substrate rather than a high-throughput datastore? A strong answer names the cost model explicitly (objects times watchers times churn, plus memory, plus recovery), picks the levers with the highest ratio of saving to complexity, and treats control-plane capacity as a shared resource that individual controllers can and do degrade for everyone.
- A team proposes a controller that updates a lastSeen timestamp in the status of every Pod once per second. What is your review response?Reject it on blast radius: it converts every pod into a high-churn object, and each write is delivered to every watcher of Pods across the cluster, so the cost is multiplied by the number of controllers and kubelets watching. If a heartbeat is genuinely needed, put it in a small dedicated object with few watchers — the pattern Lease objects exist for — or keep it outside the API entirely in metrics.
- Why is a full LIST the operation to design against rather than the watch itself?A steady watch delivers only deltas and is cheap per client, while a LIST materialises and serializes an entire collection at once. The danger is correlated relists — an API server rollout or a healed partition causing thousands of clients to list simultaneously — which is why bookmarks, jittered backoff and correct resume handling matter more than shaving individual watch cost.
saying these in an interview costs you the question
- Assuming watches are free because they are push-based
- Setting a short resync interval believing it improves cache freshness
- Creating a separate informer per controller for the same resource in one process
- Watching all Secrets or all Pods cluster-wide when a selector would do
- Ignoring the relist stampede risk after control-plane restarts
- Treating the Kubernetes API as a general-purpose high-write datastore