Cluster Architecture
The pieces that make a cluster run and the machinery joining them: kube-apiserver as the only writer to etcd, the controller and scheduler loops, the kubelet on each node, and the list-watch model between them. Interviewers use it to tell operators from users.
part ofKubernetesoverview, primer and where to startread it →on this pageshowhide
explore
- kube-apiserver6 questions
- etcd4 questions
- Controller Manager and Scheduler Roles4 questions
- Node Components5 questions
- Desired-State Reconciliation4 questions
- Watch and Informer Mechanics5 questions
- API Groups & Versioning3 questions
- HA Control Planes5 questions
- API Priority & Fairness4 questions
- Aggregated APIs & Metrics3 questions
questions
page 2 of 2Kubernetes lets you add API endpoints either with a CustomResourceDefinition or through the API aggregation layer and an APIService object. Explain how aggregation works inside kube-apiserver and when you would choose it over a CRD.
basics
~20 sAn APIService registers a group/version and points at a Service; kube-apiserver then proxies every request for that group/version to that extension server, which implements storage and semantics itself. Choose it when you need non-etcd storage or custom verbs.
A client calls the Kubernetes API with `?watch=true&resourceVersion=<rv>` and after some idle time receives HTTP 410 Gone with "too old resource version". Explain the contract the API server's watch endpoint offers and what the client is required to do.
basics
~20 sA watch streams change events occurring after a given resourceVersion, which the server can only replay from a bounded in-memory window. Once that version ages out, the server returns 410 and the client must re-list to get a fresh resourceVersion and re-watch.
kube-controller-manager runs on all three control-plane nodes of a highly available cluster. Why don't three copies of the ReplicaSet controller each create their own set of Pods?
basics
~20 sBecause of leader election: all three instances start, but each competes to hold a Lease object in the API server. Only the lease holder runs its control loops; the others idle as hot standbys and take over if the lease expires.
A Kubernetes controller's reconcile function fails for one object because a downstream API returns errors. Why does the controller use a rate-limited work queue instead of simply retrying in a loop?
basics
~20 sThe work queue deduplicates keys, guarantees one worker per key at a time, and requeues failures with exponential backoff. A tight retry loop would block other objects, hammer the API server, and turn one broken object into a control-plane load problem.
An etcd member backing a Kubernetes cluster starts rejecting writes with "mvcc: database space exceeded". What causes this, and what is the difference between compaction and defragmentation in fixing it?
basics
~20 setcd keeps every historical revision and enforces a space quota (2 GiB by default). Compaction discards revisions older than a chosen point, freeing space inside the file; defragmentation rewrites the file so freed pages return to the filesystem. Afterwards you must disarm the NOSPACE alarm before writes resume.
In a production cluster a Deployment's replica count flips between two values every few seconds and no human is running kubectl. How do you diagnose and fix that kind of fight over a single field?
basics
~20 sTwo writers each treat that field as theirs and keep correcting each other's drift. Find them via metadata.managedFields, the audit log and controller logs, then give the field exactly one owner — usually by removing it from the declarative manifest and letting the autoscaler or operator own it.
You are designing the control plane for a self-managed, bare-metal Kubernetes cluster spread across racks. How do you choose between three and five control-plane nodes, and where do you place them?
basics
~20 sPick the member count from the failures you must survive: three tolerate one loss, and five tolerate two, including one planned. Spread members so no single rack or power domain holds a majority, which needs at least three failure domains with low-latency links.
Why would a Kubernetes client-go program log "client-side throttling, not priority and fairness", and how does that differ from an HTTP 429?
basics
~20 sThat delay comes from client-go's own token-bucket rate limiter (5 QPS, burst 10 unless you change them), which holds requests before they leave the process. An HTTP 429 means the API server's Priority and Fairness filter rejected a request that did arrive.
The control-plane node running the active kube-controller-manager loses power, and a shipment-tracking rollout stalls for 47 seconds. What governs Kubernetes leader-election failover, and would you tune it?
basics
~20 sA standby takes over only after the leader's Lease looks expired: 15 seconds by default, plus a jittered 2-second retry. The rest of a 47-second stall usually comes from endpoint failover, stuck connections and cache sync.
What is a BOOKMARK event in the Kubernetes watch protocol, and which problem does it solve for a long-lived watch client?
basics
~20 sA BOOKMARK is a periodic watch event carrying only an up-to-date resourceVersion and no object change. It lets an idle client advance its resume cursor, so after a disconnect it can resume instead of getting 410 Gone and re-listing everything.
A large cluster's control plane is degrading: API latency spikes, controllers time out, and kube-apiserver memory keeps climbing. How do you reason about its capacity, and how do you protect it from expensive clients?
basics
~20 sClassify the load: expensive unpaginated LISTs dominate memory, watches dominate steady state. Add replicas for throughput, then constrain clients with pagination, selectors and informers, and enforce fairness with API Priority and Fairness so one client cannot starve the rest.
As a cluster grows to thousands of nodes and dozens of controllers, watch traffic becomes a significant load on the control plane. How do you reason about that cost and reduce it on the client side?
basics
~20 sCost scales with objects times watchers times change rate, plus the memory each client caches. Reduce it by sharing informers per process, scoping watches with label and field selectors, trimming cached objects, avoiding short resyncs and relist storms, and keeping fast-churning data out of watched objects.
showing 31–43 of 43