A large cluster's control plane is degrading: API latency spikes, controllers time out, and kube-apiserver memory keeps climbing. How do you reason about its capacity, and how do you protect it from expensive clients?
answer
- LIST = memory; WATCH = per-replica cache; writes = etcd bound
- replicas scale throughput, not etcd writes or cache memory
- clients: informers, limit/continue, selectors, PartialObjectMetadata, bookmarks
- APF FlowSchema + PriorityLevelConfiguration → 429 instead of starvation
- attribute by user_agent; watch apiserver_flowcontrol_* metrics
basics
~20 sClassify the load: expensive unpaginated LISTs dominate memory, watches dominate steady state. Add replicas for throughput, then constrain clients with pagination, selectors and informers, and enforce fairness with API Priority and Fairness so one client cannot starve the rest.
solid answer
~60 sSeparate the request classes, because they cost differently. - **LIST** is the memory killer: an unpaginated list of every Pod deserialises the whole collection into the API server's heap and serialises it again per client. A handful of misbehaving clients doing this on a timer will drive OOM. - **WATCH** is cheap per event but each replica holds its own watch cache and connections, so memory scales with object count *per replica*. - **GET/write** load ultimately lands on etcd, which does not scale horizontally. Then act on three levers. **Capacity**: the API server is stateless, so add replicas behind the load balancer — but that multiplies watch-cache memory and does nothing for etcd write limits. **Client behaviour**, the real fix: informers instead of polling, `limit`/`continue` pagination, field and label selectors, metadata-only responses, and `allowWatchBookmarks`. **Isolation**: API Priority and Fairness classifies requests into flow schemas and priority levels, so a runaway operator gets 429s while leader election and node heartbeats keep flowing. Measure with `apiserver_request_duration_seconds`, `apiserver_longrunning_requests`, and the `apiserver_flowcontrol_*` series before changing anything.
code
bash · 8 lineskubectl get --raw /metrics \
| grep 'apiserver_request_total' | grep 'verb="LIST"' | sort -t' ' -k2 -rn | head
kubectl get --raw /metrics | grep apiserver_longrunning_requests
kubectl get --raw /metrics | grep apiserver_flowcontrol_rejected_requests_total
kubectl get flowschemas
kubectl get prioritylevelconfigurationsgo deeper
Know that listing everything repeatedly is expensive and that informers exist so clients watch instead of poll.
Distinguish LIST from WATCH cost, apply pagination and selectors, and recognise 429 as backpressure to be retried rather than an error to escalate.
Attribute load from metrics and audit, apply APF FlowSchemas to contain a heavy client, and explain why replicas do not fix etcd or cache memory.
Decide what the control plane is allowed to be used for: enforce client conventions organisation-wide, encode starvation policy in FlowSchemas, and know when splitting clusters or moving state out of the API beats further tuning.
## First, classify the traffic kube-apiserver capacity is not one number. The request classes have qualitatively different costs, and conflating them leads to the wrong fix. **LIST** is the expensive one. Serving a collection means fetching objects (from the watch cache, or from etcd for a quorum read), decoding, converting to the requested version, and serialising per client. A `LIST /api/v1/pods` on a cluster with 100k Pods materialises hundreds of megabytes transiently, and several concurrent copies is an OOM. Worse, this is exactly what naive clients do on a loop. **WATCH** is the opposite profile: expensive to establish, then cheap per event. Steady state cost is memory for the watch cache plus per-connection overhead. Crucially the watch cache is **per replica**, so three API servers hold three copies. This is why scaling out helps latency but not memory pressure. **GET and single-object writes** are individually cheap, but writes must reach etcd and every write is a Raft round trip across the quorum. Beyond a point, write throughput is an etcd property, not an API-server one. **Long-running requests** — watches, `exec`, `port-forward`, `log -f` — occupy handlers indefinitely and are accounted separately from short requests. ## Understand what "add replicas" buys The API server is stateless, so replicas scale request throughput and give HA. What replicas do **not** fix: etcd write capacity (shared), watch-cache memory (duplicated), or a client that lists all Pods every ten seconds (now it just picks a victim). Scaling out is necessary but rarely sufficient, and it is the fix people reach for first because it requires no negotiation with application teams. ## The durable fix is client behaviour Most control-plane pathology traces to a handful of clients: - **Polling instead of watching.** A CI script or a homegrown operator listing all Pods every few seconds. Replace with informers, which list once and then watch. - **Unpaginated LISTs.** Use `limit` and `continue` so the server streams pages instead of building the whole collection. - **No selectors.** Server-side `fieldSelector` and `labelSelector` shrink both the work and the payload; filtering client-side means the server already paid the full cost. - **Full objects when metadata suffices.** Requesting `PartialObjectMetadata` (`Accept: application/json;as=PartialObjectMetadataList;g=meta.k8s.io;v=v1`) skips serialising specs and statuses entirely — a large win for garbage collectors and inventory tools. - **No bookmarks.** `allowWatchBookmarks=true` keeps cursors fresh so a control-plane restart does not trigger a synchronised relist storm from every client at once — one of the classic ways a recovering control plane immediately falls over again. - **Storing large objects.** Huge ConfigMaps and Secrets, or CRDs with unbounded fields, inflate every list of those resources. Cap them. Finding the offenders: the `apiserver_request_total` series broken down by `verb`, `resource` and `user_agent` usually names them within minutes, and audit logs at `Metadata` level confirm it. Insist that clients set a distinguishing user-agent — an organisation where everything reports `Go-http-client/2.0` cannot diagnose this at all. ## Isolation: API Priority and Fairness Older clusters had only global `--max-requests-inflight` / `--max-mutating-requests-inflight` counters: a blunt cap under which one aggressive client could consume the whole budget and starve node heartbeats and leader election, cascading into node NotReady and controller failovers. APF replaces this with two object types. A **FlowSchema** matches requests by subject (user, group, service account) and resource, assigns them to a priority level, and defines a *flow distinguisher* (typically by user or namespace). A **PriorityLevelConfiguration** owns a share of the concurrency budget and queues within it, fair-queued across flows. Defaults ship with sensible levels: `system` (node heartbeats), `leader-election`, `workload-high`, `workload-low`, `global-default`, and an exempt level for the most critical system traffic. The effect is that a runaway operator saturates *its own* priority level and receives **429 Too Many Requests** with a `Retry-After`, while leader election and kubelet heartbeats continue. Client-go honours the retry, so well-behaved clients back off automatically. Observe with `apiserver_flowcontrol_request_wait_duration_seconds`, `apiserver_flowcontrol_current_inqueue_requests`, and `apiserver_flowcontrol_rejected_requests_total`; a custom FlowSchema pinning a known-heavy batch client to a low level is the standard remediation. ## Metrics to reason with - `apiserver_request_duration_seconds` by verb/resource/scope — latency, and which class is slow. Cluster-scoped LISTs standing out is the giveaway. - `apiserver_request_total` by code/user_agent — offenders and 429 rates. - `apiserver_longrunning_requests` — watch and exec counts. - `apiserver_current_inflight_requests` and the flowcontrol series — saturation. - `etcd_request_duration_seconds` and etcd's own disk-commit latency — to decide whether the bottleneck is really downstream, since etcd is disproportionately sensitive to disk fsync latency. - Process memory versus object counts per resource, to confirm the LIST hypothesis. ## Making the judgement The principal-level content is deciding what you are optimising and what you will refuse. Adding replicas is cheap and buys latency headroom immediately; it is the right first move under fire but must not become the answer. Capping client behaviour requires organisational work — a platform library that everyone uses, a rule that operators must use informers, a review gate on new controllers, and user-agent conventions so attribution is possible. APF is the safety net that keeps a single bad actor from taking the cluster down while that work happens; encoding "which traffic is allowed to starve" in FlowSchemas is a genuine architectural statement about the cluster. And sometimes the honest answer is a split: separate clusters for high-churn workloads, or moving a chatty system off custom resources entirely, because a control plane is a shared consistency service, not a general-purpose database.
- Why does adding a third API-server replica sometimes make memory pressure worse rather than better?Each replica maintains its own watch cache and its own watch connections to etcd, so the cached object set is duplicated per replica rather than shared. Replicas add request-serving capacity and HA, but total control-plane memory grows roughly linearly with replica count, and etcd now fields more watchers. If the actual problem is expensive LISTs or oversized objects, replicas spread the damage instead of reducing it.
- A client suddenly receives 429 Too Many Requests from the API server. What should it do, and what does the response tell you?It should honour the `Retry-After` header and back off — client-go does this automatically. The 429 means API Priority and Fairness queued or rejected the request at its assigned priority level, which is the system deliberately protecting higher-priority traffic such as leader election and node heartbeats. Diagnose by finding which FlowSchema matched the client and whether that priority level's concurrency share is genuinely too small or the client is simply too chatty.
- When is the right answer to move workloads to a separate cluster rather than tune the API server?When the load is inherent rather than accidental — very high object churn, an operator legitimately reconciling hundreds of thousands of custom resources, or per-tenant CRDs growing without bound. The control plane is a strongly consistent coordination service backed by etcd, and etcd's write throughput and database size do not scale horizontally. At that point splitting the workload across clusters, or moving high-frequency state out of the API entirely, is cheaper and safer than continuing to tune.
It is a single checkout counter for a whole warehouse. Adding tills helps until the stockroom behind them is the limit — and without lane assignment, one customer emptying every shelf blocks the people buying a single item.
saying these in an interview costs you the question
- Treating replica count as the primary capacity lever, ignoring etcd and per-replica cache memory
- Recommending raising max-requests-inflight on a modern cluster instead of using APF
- Letting clients poll collections on a timer instead of using informers
- Blaming etcd without checking whether unpaginated LISTs are driving API-server heap
- Deploying clients with a default user-agent, making attribution impossible during an incident