skip to content

Cluster Architecture

The pieces that make a cluster run and the machinery joining them: kube-apiserver as the only writer to etcd, the controller and scheduler loops, the kubelet on each node, and the list-watch model between them. Interviewers use it to tell operators from users.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

explore

questions

page 2 of 2

Kubernetes lets you add API endpoints either with a CustomResourceDefinition or through the API aggregation layer and an APIService object. Explain how aggregation works inside kube-apiserver and when you would choose it over a CRD.

level: seniorimportance: should knowfreq 30%

basics

~20 s

An APIService registers a group/version and points at a Service; kube-apiserver then proxies every request for that group/version to that extension server, which implements storage and semantics itself. Choose it when you need non-etcd storage or custom verbs.

open as a page

A client calls the Kubernetes API with `?watch=true&resourceVersion=<rv>` and after some idle time receives HTTP 410 Gone with "too old resource version". Explain the contract the API server's watch endpoint offers and what the client is required to do.

level: seniorimportance: should knowfreq 40%

basics

~20 s

A watch streams change events occurring after a given resourceVersion, which the server can only replay from a bounded in-memory window. Once that version ages out, the server returns 410 and the client must re-list to get a fresh resourceVersion and re-watch.

open as a page

kube-controller-manager runs on all three control-plane nodes of a highly available cluster. Why don't three copies of the ReplicaSet controller each create their own set of Pods?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Because of leader election: all three instances start, but each competes to hold a Lease object in the API server. Only the lease holder runs its control loops; the others idle as hot standbys and take over if the lease expires.

open as a page

A Kubernetes controller's reconcile function fails for one object because a downstream API returns errors. Why does the controller use a rate-limited work queue instead of simply retrying in a loop?

level: seniorimportance: should knowfreq 42%

basics

~20 s

The work queue deduplicates keys, guarantees one worker per key at a time, and requeues failures with exponential backoff. A tight retry loop would block other objects, hammer the API server, and turn one broken object into a control-plane load problem.

open as a page

An etcd member backing a Kubernetes cluster starts rejecting writes with "mvcc: database space exceeded". What causes this, and what is the difference between compaction and defragmentation in fixing it?

level: seniorimportance: should knowfreq 42%

basics

~20 s

etcd keeps every historical revision and enforces a space quota (2 GiB by default). Compaction discards revisions older than a chosen point, freeing space inside the file; defragmentation rewrites the file so freed pages return to the filesystem. Afterwards you must disarm the NOSPACE alarm before writes resume.

open as a page

In a production cluster a Deployment's replica count flips between two values every few seconds and no human is running kubectl. How do you diagnose and fix that kind of fight over a single field?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Two writers each treat that field as theirs and keep correcting each other's drift. Find them via metadata.managedFields, the audit log and controller logs, then give the field exactly one owner — usually by removing it from the declarative manifest and letting the autoscaler or operator own it.

open as a page

What is a shared informer in the Kubernetes client libraries, and what are the consequences of reading objects from its local cache instead of from the API server?

level: seniorimportance: should knowfreq 38%

basics

~20 s

An informer runs one list-watch per resource type, keeps the objects in an in-memory indexed store, and fires handlers on changes; shared means many controllers in the process reuse that one stream and cache. Cache reads are fast and free but eventually consistent, so they can be stale and must never be mutated.

open as a page

You are designing the control plane for a self-managed, bare-metal Kubernetes cluster spread across racks. How do you choose between three and five control-plane nodes, and where do you place them?

level: principalimportance: should knowfreq 40%

basics

~20 s

Pick the member count from the failures you must survive: three tolerate one loss, and five tolerate two, including one planned. Spread members so no single rack or power domain holds a majority, which needs at least three failure domains with low-latency links.

open as a page

Why would a Kubernetes client-go program log "client-side throttling, not priority and fairness", and how does that differ from an HTTP 429?

level: juniorimportance: nice to knowfreq 32%

basics

~20 s

That delay comes from client-go's own token-bucket rate limiter (5 QPS, burst 10 unless you change them), which holds requests before they leave the process. An HTTP 429 means the API server's Priority and Fairness filter rejected a request that did arrive.

open as a page

The control-plane node running the active kube-controller-manager loses power, and a shipment-tracking rollout stalls for 47 seconds. What governs Kubernetes leader-election failover, and would you tune it?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

A standby takes over only after the leader's Lease looks expired: 15 seconds by default, plus a jittered 2-second retry. The rest of a 47-second stall usually comes from endpoint failover, stuck connections and cache sync.

open as a page

What is a BOOKMARK event in the Kubernetes watch protocol, and which problem does it solve for a long-lived watch client?

level: seniorimportance: nice to knowfreq 20%

basics

~20 s

A BOOKMARK is a periodic watch event carrying only an up-to-date resourceVersion and no object change. It lets an idle client advance its resume cursor, so after a disconnect it can resume instead of getting 410 Gone and re-listing everything.

open as a page

A large cluster's control plane is degrading: API latency spikes, controllers time out, and kube-apiserver memory keeps climbing. How do you reason about its capacity, and how do you protect it from expensive clients?

level: principalimportance: nice to knowfreq 25%

basics

~20 s

Classify the load: expensive unpaginated LISTs dominate memory, watches dominate steady state. Add replicas for throughput, then constrain clients with pagination, selectors and informers, and enforce fairness with API Priority and Fairness so one client cannot starve the rest.

open as a page

As a cluster grows to thousands of nodes and dozens of controllers, watch traffic becomes a significant load on the control plane. How do you reason about that cost and reduce it on the client side?

level: principalimportance: nice to knowfreq 22%

basics

~20 s

Cost scales with objects times watchers times change rate, plus the memory each client caches. Reduce it by sharing informers per process, scoping watches with label and field selectors, trimming cached objects, avoiding short resyncs and relist storms, and keeping fast-churning data out of watched objects.

open as a page

showing 31–43 of 43