skip to content

Cluster Architecture

The pieces that make a cluster run and the machinery joining them: kube-apiserver as the only writer to etcd, the controller and scheduler loops, the kubelet on each node, and the list-watch model between them. Interviewers use it to tell operators from users.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

explore

questions

page 1 of 2

In Kubernetes, what does metrics-server provide for `kubectl top` and the HorizontalPodAutoscaler, and what is it deliberately not designed to be?

level: juniorimportance: must knowfreq 72%

answer

  1. narrow, current, in-memory
  2. kubelet /metrics/resource scrape
  3. APIService v1beta1.metrics.k8s.io
  4. two points make a CPU rate
  5. no history, no custom metrics

basics

~20 s

metrics-server scrapes current CPU and memory usage from every kubelet, keeps only the latest points in memory, and serves them as the metrics.k8s.io API. It is not a monitoring system: no history, no custom metrics, no alerting.

solid answer

~40 s

metrics-server is a small cluster add-on that periodically scrapes each kubelet's `/metrics/resource` endpoint for CPU and memory usage of nodes, pods and containers. It keeps only the most recent scrapes in memory and serves them through the `metrics.k8s.io` API, which it plugs into kube-apiserver through an `APIService` object named `v1beta1.metrics.k8s.io`. That API is what `kubectl top` and a CPU- or memory-based HorizontalPodAutoscaler read. It is deliberately *not* a monitoring system: it stores no history, loses everything on restart, serves no application or custom metrics, sends nothing to a backend and does not claim to be an accurate source for billing or capacity reports. For trends, dashboards and alerts you run a real metrics backend alongside it.

code

bash · 3 lines
bash
kubectl top node
kubectl top pod -n checkout --sort-by=memory --containers
kubectl get --raw /apis/metrics.k8s.io/v1beta1/namespaces/checkout/pods | head -c 400

go deeper

for a junior

Remember three facts: it serves current CPU and memory, it is what kubectl top and a CPU-based HPA read, and it keeps no history.

for a middle

Explain the path: kubelet /metrics/resource, periodic scrape, two points in memory, served as metrics.k8s.io through an APIService and the aggregation layer.

for a senior

Show you know when it is enough and when it is not: it covers autoscaling and quick triage, while trends, alerting and chargeback need a separate metrics backend.

for a principal

Frame the cost trade-off: on many small edge clusters, metrics-server alone may be the right footprint, provided the organisation accepts it cannot answer historical questions.

## What metrics-server is **metrics-server** is a Kubernetes SIG-maintained add-on, usually running as a Deployment in `kube-system`, that implements the **resource metrics API**. That API lives in the API group `metrics.k8s.io` and exposes two read-only resources: - `nodes` — current CPU and memory usage per Node - `pods` — current CPU and memory usage per Pod, broken down by container It exists so that core Kubernetes features have a *standard, always-present* answer to one narrow question: "how much CPU and memory is this thing using right now?" The two main consumers are `kubectl top` and the **HorizontalPodAutoscaler** when it targets CPU or memory utilisation. ## How the data flows 1. Every kubelet already measures container usage through the container runtime and publishes it on its HTTPS port (10250 by default) at `/metrics/resource`. 2. metrics-server scrapes that endpoint on every node at a fixed interval, set by `--metric-resolution`. The upstream manifest sets `15s`; the flag rejects anything below `10s`. 3. It keeps the **last two scrapes** per container and node in memory. CPU usage is a rate, so it is computed from the difference between those two points; memory is the latest working-set value. 4. It serves the result as an ordinary-looking Kubernetes API. An `APIService` object named `v1beta1.metrics.k8s.io` tells kube-apiserver to forward every request under `/apis/metrics.k8s.io/v1beta1/` to the `metrics-server` Service. 5. `kubectl top pods` and the HPA controller call kube-apiserver as usual; the **aggregation layer** proxies the call to metrics-server, which answers from memory. Because the answer comes from memory, a request for a 1,180-pod namespace in a ticket-booking checkout is cheap — no query engine, no disk. ## What it deliberately is not The project README states its non-goals plainly: it is not for non-Kubernetes clusters, not "an accurate source of resource usage metrics", and not for autoscaling on anything other than CPU and memory. | Expectation | Reality in metrics-server | |---|---| | History / trends | None — only the latest window; a restart starts from empty | | Application metrics (requests per second, queue depth) | Not served; those come from custom or external metrics adapters | | Dashboards and alerting | Not provided; there is no query language and no rules | | Export to a backend | Nothing is pushed anywhere | | Billing-grade accuracy | Explicitly disclaimed; it is sized for autoscaling decisions | | Object state (desired vs ready replicas) | Not its job; that is kube-state-metrics | A metrics backend such as Prometheus is a **separate, parallel pipeline**. It scrapes its own targets and stores time series; it does not use metrics-server, and metrics-server does not feed it. ## Reading `kubectl top` correctly - The numbers are **recent usage**, not requests or limits. A pod requesting `500m` but idle shows a few millicores. - Right after a pod starts, or right after metrics-server restarts, a pod may be missing: two scrapes are needed before a CPU rate exists. - `kubectl top pod --containers` shows the per-container split; `--sort-by=cpu` or `--sort-by=memory` orders the list. - `kubectl top node` values come from the kubelet's node-level measurement, so they include system daemons and do not equal the sum of the pods. ## Operating it on a small cluster On a 5-node edge cluster in a retail store, metrics-server is often the *only* metrics component, because shipping a full monitoring stack to every store is expensive. That is fine for `kubectl top` and CPU-based scaling of the checkout, but the team must accept that nobody can answer "what was checkout's memory at 14:05 yesterday?" from it. Its footprint grows with the number of pods and nodes it tracks, so its memory request is sized to cluster scale rather than traffic. It also depends on the aggregation layer being configured on kube-apiserver and on kubelets presenting serving certificates it can verify; when either is broken, `kubectl top` fails even though every workload is healthy. ## Common interview traps - **"A metrics backend makes metrics-server redundant"** — they coexist; removing metrics-server breaks `kubectl top` and resource-based HPAs even when a metrics backend is healthy. - **"It reads cAdvisor directly"** — current releases scrape the kubelet's `/metrics/resource` endpoint; the kubelet is the single source. - **"More replicas give more history"** — running several metrics-server replicas is for availability of the API, not for retention; each replica still holds only the latest window. - **"`kubectl top` is authoritative for sizing requests"** — it is a point-in-time glance; sizing needs percentiles over days, which only a stored time series can give.

  • Why does a freshly started pod sometimes not appear in `kubectl top pods` for a short while?
    CPU usage is a rate, so metrics-server needs two scrapes of the same container to compute it. Until the second scrape after the pod starts has landed, which can take one or two `--metric-resolution` intervals, there is no usable point and the pod is left out of the response. The same gap appears for every pod right after metrics-server itself restarts, because its store starts empty.
  • What must be true of a Kubernetes cluster before metrics-server can serve anything?
    kube-apiserver must have the aggregation layer configured, with front-proxy certificates. The control plane must be able to reach the metrics-server pod, and metrics-server must reach every kubelet on its published address and port. Kubelets need webhook authentication and authorization enabled, and their serving certificates must be signed by the cluster CA, unless verification is switched off, which is acceptable only for testing.
  • Can you point a metrics backend at metrics-server to get history instead of running node-level exporters?
    It is the wrong tool. metrics-server serves only CPU and memory, in the Kubernetes API format rather than a scrape format, and is sized for the latest window. A metrics backend should scrape kubelets and node exporters directly, which yields far richer series. The two pipelines are meant to run side by side, each for its own purpose.

metrics-server is a car's speedometer, not its trip log: it tells you how fast you are going right now and forgets the moment you look away.

saying these in an interview costs you the question

  • metrics-server stores a week of usage history for capacity planning
  • kubectl top shows the pod's CPU requests and limits
  • metrics-server is required for a metrics backend such as Prometheus to work
  • metrics-server can serve requests-per-second for autoscaling
  • metrics-server reads usage from etcd
  • kubectl top node equals the sum of all pods on the node
open as a page

In a Kubernetes control plane, what is kube-apiserver responsible for, and why is it the only component that talks to etcd directly?

level: juniorimportance: must knowfreq 75%

basics

~20 s

kube-apiserver is the cluster's front door: a REST API that authenticates, authorises, validates and persists every object, and streams changes to watchers. Routing all writes through it gives one place for auth, validation, admission and audit.

open as a page

In a Kubernetes cluster, what does the kube-scheduler actually do when a new Pod is created, and what does it not do?

level: juniorimportance: must knowfreq 78%

basics

~20 s

kube-scheduler watches for Pods with an empty spec.nodeName, picks a suitable node, and writes that choice back to the API server (a binding). It does not start the container — the kubelet on the chosen node does that.

open as a page

In a Kubernetes control plane, what is etcd, what exactly is stored in it, and which components are allowed to talk to it?

level: juniorimportance: must knowfreq 70%

basics

~20 s

etcd is the cluster's only persistent database: a distributed, strongly consistent key-value store holding every API object (Pods, Deployments, Services, ConfigMaps, Secrets, RBAC, node state). Only kube-apiserver connects to it; everything else reads and writes through the API server.

open as a page

If every kube-apiserver in a Kubernetes cluster becomes unreachable, what keeps running on the nodes and what stops working?

level: juniorimportance: must knowfreq 74%

basics

~20 s

Running Pods keep running: the kubelet keeps their containers alive and kube-proxy's installed Service rules keep routing. Anything that needs a write or fresh state stops working: scheduling, scaling, rollouts, rescheduling off failed nodes and kubectl.

open as a page

What is the kubelet responsible for on a Kubernetes worker node, and what does its pod sync loop actually do?

level: juniorimportance: must knowfreq 70%

basics

~20 s

The kubelet is the node agent. It watches the API server for Pods assigned to its node, tells the container runtime to create the containers described in each PodSpec, mounts volumes, runs probes, and reports Pod and node status back. It manages Pods only, not Deployments or ReplicaSets.

open as a page

In a Kubernetes object such as a Deployment, what is the difference between the spec section and the status section, and who writes each one?

level: juniorimportance: must knowfreq 72%

basics

~10 s

spec is the desired state you declare; status is the observed state the system reports. You and your tools write spec, controllers write status. Kubernetes keeps acting until status matches spec.

open as a page

In Kubernetes API versioning, what do alpha, beta and GA versions promise, and how long must a deprecated version stay served?

level: middleimportance: must knowfreq 64%

basics

~20 s

Alpha versions are off by default and may vanish in any release; beta versions stay served at least 9 months or 3 releases after deprecation; GA versions are never removed within Kubernetes v1. New betas are off by default since v1.24.

open as a page

How does Kubernetes API Priority and Fairness decide which priority level an incoming API request lands in, and what happens when that level is full?

level: middleimportance: must knowfreq 55%

basics

~20 s

kube-apiserver checks each request against the FlowSchema objects, lowest matchingPrecedence first, and sends it to the PriorityLevelConfiguration the matching schema names. If that level has no free seats, the request waits in a fair queue by flow or gets HTTP 429.

open as a page

Trace what kube-apiserver does with an HTTPS POST of a Deployment manifest, from the moment the request arrives until the object is durably stored. Name the stages in order and say what each can reject.

level: middleimportance: must knowfreq 60%

basics

~10 s

TLS, then authentication (who), authorisation (may they), decode plus defaulting, mutating admission, schema validation, validating admission, then a write to etcd guarded by optimistic concurrency, followed by audit and watch fan-out.

open as a page

What is kube-controller-manager, and what does a single controller inside it actually do on each iteration?

level: middleimportance: must knowfreq 72%

basics

~20 s

It is one control-plane process hosting many independent controllers (Deployment, ReplicaSet, Node, Job, endpoints, service accounts…). Each watches its objects via the API server, compares desired spec to observed state, and makes one corrective API call — repeatedly, until they match.

open as a page

Why is the Kubernetes cluster store etcd almost always deployed with 3 or 5 members rather than 2 or 4, and what happens to the cluster when a majority of those members is unavailable?

level: middleimportance: must knowfreq 60%

basics

~20 s

etcd uses Raft: a write must be replicated to a majority (quorum) before it commits, so an N-member cluster tolerates (N-1)/2 failures. Even sizes add a member without adding fault tolerance. Without quorum etcd rejects writes, so the control plane freezes — but running Pods keep running.

open as a page

In a highly available Kubernetes control plane, how do stacked and external etcd topologies differ, and what failures can each tolerate?

level: middleimportance: must knowfreq 63%

basics

~20 s

Stacked etcd runs a member on each control-plane node beside kube-apiserver; external etcd uses dedicated hosts. Stacked needs fewer machines but couples failures: three stacked nodes tolerate one loss, and each loss removes an API server and an etcd member.

open as a page

You delete a running Pod with kubectl delete pod and an equivalent Pod appears seconds later. Explain the mechanism that recreated it, and why Kubernetes controllers are described as level-triggered rather than edge-triggered.

level: middleimportance: must knowfreq 58%

basics

~20 s

A controller (the ReplicaSet behind the Deployment) constantly compares desired replica count with the pods it actually sees, and creates one when the count is short. Level-triggered means it acts on current state, not on the delete event, so it recovers even if it missed the event.

open as a page

Your code updates a Kubernetes object and the API server responds with HTTP 409 Conflict saying the object has been modified. What causes that, and what are the correct ways to handle it?

level: middleimportance: must knowfreq 45%

basics

~20 s

Kubernetes updates use optimistic concurrency: your object carries the resourceVersion you read, and the server rejects the write if the live object has changed since. Handle it by re-reading the object fresh, reapplying your change, and retrying — or by sending a patch instead of a full update.

open as a page

Describe how a Kubernetes client keeps an up-to-date view of a set of objects using the list-then-watch pattern, and what the resourceVersion value returned by the API server means in that flow.

level: middleimportance: must knowfreq 46%

basics

~20 s

The client LISTs the objects once, notes the collection's resourceVersion, then opens a WATCH starting after that version and receives ADDED/MODIFIED/DELETED events as a stream. resourceVersion is an opaque cursor into the API server's change history, not a number to compare or interpret.

open as a page

How do you take a backup of the Kubernetes cluster store with etcdctl, and what is the actual procedure to restore a cluster from that backup?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Back up with etcdctl snapshot save (v3 API, with client TLS certs), on a schedule, stored off-box. To restore: stop the API servers, stop etcd and move the old data directory aside, run etcdctl snapshot restore into a fresh directory, start etcd as a new single-member cluster, restart the control plane, then re-add members.

open as a page

How does a Kubernetes node register itself with the cluster and prove it is still alive, and what would you check when a node shows status NotReady while its workloads are still serving traffic?

level: seniorimportance: must knowfreq 48%

basics

~20 s

On startup the kubelet authenticates (usually a bootstrap token plus a CSR) and creates the Node object with its capacity and labels. It then renews a Lease in the kube-node-lease namespace every ~10s as a heartbeat and updates node conditions less often. NotReady means the kubelet stopped reporting or reported a problem — not that containers died.

open as a page

Using kubectl, how do you find which API group and version a Kubernetes resource kind is served at, and which fields it accepts?

level: juniorimportance: should knowfreq 58%

basics

~10 s

Run kubectl api-resources to list each resource's group/version, kind, short name and scope, kubectl api-versions to list every served group/version, and kubectl explain to read the fields the cluster's published schema accepts.

open as a page

Kubernetes defines the metrics.k8s.io, custom.metrics.k8s.io and external.metrics.k8s.io APIs; what serves each one, and where does kube-state-metrics fit?

level: middleimportance: should knowfreq 46%

basics

~20 s

Each group is served by a separate server registered through an APIService: metrics-server for metrics.k8s.io, and adapters for custom and external metrics. kube-state-metrics is not an API at all but a plain exporter of object state.

open as a page

Why is an unpaginated LIST expensive for the Kubernetes API server, and how do the limit and continue parameters let a client page through results?

level: middleimportance: should knowfreq 44%

basics

~20 s

A full LIST makes kube-apiserver load, convert and serialize every matching object in one response, and Priority and Fairness charges it several seats. With limit, the server returns pages plus a continue token for the next request until the token comes back empty.

open as a page

Kubernetes manifests declare fields such as `apiVersion: apps/v1` and `kind: Deployment`. Explain how group, version and kind map onto the REST URLs the API server exposes, and how the same stored object can be served at more than one version.

level: middleimportance: should knowfreq 45%

basics

~10 s

apiVersion is group/version and kind is the type; together they select a REST path like /apis/apps/v1/namespaces/ns/deployments. The server stores one storage version and converts to whichever served version a client requests.

open as a page

In an HA Kubernetes control plane, why do kubelets and clients reach kube-apiserver through a load balancer, and how should it check backends?

level: middleimportance: should knowfreq 52%

basics

~20 s

kube-apiserver replicas are stateless and all active, so clients need one stable address that survives any single server's loss. Put a TCP pass-through load balancer or virtual IP in front of them, and have it health-check each server's /readyz endpoint.

open as a page

What is the Container Runtime Interface (CRI) in Kubernetes, and how does the kubelet use it to run containers with containerd?

level: middleimportance: should knowfreq 50%

basics

~20 s

CRI is the gRPC API the kubelet uses to talk to a container runtime over a local socket. The kubelet calls it to pull images and to create pod sandboxes and containers; containerd implements CRI natively and delegates actual container creation to a low-level OCI runtime such as runc.

open as a page

What is kube-proxy's job on a Kubernetes node, and what still works if kube-proxy stops running there?

level: middleimportance: should knowfreq 50%

basics

~20 s

kube-proxy runs on every node (usually as a DaemonSet) and programs the node's kernel so traffic to a Service's virtual IP is redirected to a healthy backend Pod. It is a control agent writing rules, not a data-path proxy: if it stops, existing rules keep working but new Services and endpoint changes stop being applied.

open as a page

What is a static pod in Kubernetes, how does it differ from a Pod created through the API server, and where are static pods actually used?

level: middleimportance: should knowfreq 45%

basics

~20 s

A static pod is defined by a manifest file in a directory the kubelet watches (typically /etc/kubernetes/manifests). The kubelet runs it directly, with no scheduler or controller involved. It creates a read-only mirror Pod in the API so you can see it, but deleting that mirror does not stop it — you must remove the file.

open as a page

Compare kubectl create, kubectl replace and kubectl apply for updating a live Kubernetes object, and explain what server-side apply changed about tracking field ownership and conflicts.

level: middleimportance: should knowfreq 48%

basics

~20 s

create fails if the object exists; replace overwrites the whole object, dropping fields you omitted; apply merges your manifest with the live object. Server-side apply moves that merge into the API server and records per-field owners in metadata.managedFields, so a conflicting write is rejected until you force it.

open as a page

On a 5-node Kubernetes edge cluster, `kubectl top pods` fails with `error: Metrics API not available` while Prometheus dashboards look fine. How do you diagnose it?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Prometheus scrapes targets directly, so it proves nothing about the aggregation path. Check the v1beta1.metrics.k8s.io APIService's Available condition and reason, then fix the failing hop: Service and endpoints, apiserver-to-pod reachability and TLS, or metrics-server-to-kubelet scraping.

open as a page

Before upgrading a 64-node Kubernetes cluster, how do you find every client and manifest still using an API version the target release stops serving, and migrate them safely?

level: seniorimportance: should knowfreq 47%

basics

~10 s

Read the removal list for each release you cross, find live callers with kube-apiserver's apiserver_requested_deprecated_apis metric and k8s.io/deprecated audit annotations, scan Git for dormant manifests, then rewrite and verify with fresh audit events.

open as a page

On a 140-node Kubernetes cluster shared by 22 teams, other teams' controllers start getting HTTP 429s after one team deploys a new operator. How do you confirm API Priority and Fairness is rejecting them and contain the noisy client?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Confirm that the 429s come from Priority and Fairness by checking the X-Kubernetes-PF response headers and apiserver_flowcontrol_rejected_requests_total. Find the flow in the APF debug dumps, then move the operator into its own small priority level while its team fixes the call pattern.

open as a page

showing 1–30 of 43