skip to content

kube-apiserver

Every component, kubectl included, talks only to the API server, and it is the sole writer to etcd. Tracing one request through authentication, admission, validation and persistence is the classic 'what happens when you kubectl apply' question.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

6

In a Kubernetes control plane, what is kube-apiserver responsible for, and why is it the only component that talks to etcd directly?

level: juniorimportance: must knowfreq 75%

answer

  1. only etcd client; everyone else is an HTTPS client
  2. authn → authz → admission → validation → persist → watch fan-out
  3. chokepoint = RBAC, audit, encryption-at-rest, versioning in one place
  4. stateless → HA by replicas; scheduler/CM use leader election
  5. apiserver down: no changes, but running Pods keep running

basics

~20 s

kube-apiserver is the cluster's front door: a REST API that authenticates, authorises, validates and persists every object, and streams changes to watchers. Routing all writes through it gives one place for auth, validation, admission and audit.

solid answer

~50 s

kube-apiserver is the only component with a public interface into the cluster and the only one that reads or writes etcd. Everything else — kubectl, the scheduler, controller-manager, kubelets, controllers and operators — talks to it over HTTPS and never touches storage. Its jobs: terminate TLS; authenticate the caller; authorise the request (usually RBAC); run admission control to mutate and then validate the object; apply schema validation and defaulting; persist to etcd with optimistic concurrency; serve reads; and stream changes to watchers. Making it the sole etcd client is a deliberate chokepoint. Authentication, authorisation, admission policy, validation, audit logging, encryption at rest, and API versioning all get exactly one implementation and cannot be bypassed. It also lets etcd's client surface stay small and lets the API evolve independently of how objects are stored. It is stateless, so HA means running several replicas behind a load balancer; the shared state lives entirely in etcd.

code

bash · 6 lines
bash
kubectl get --raw /healthz
kubectl get --raw /apis | head
kubectl get --raw '/api/v1/namespaces/default/pods?limit=1'

# see the request path kubectl takes
kubectl get pods -v=8 2>&1 | grep -m1 'GET https'

go deeper

for a junior

State that it is the cluster's REST front door, the only etcd client, and that all other components are its clients. Naming authn/authz/validation/persist in order is enough.

for a middle

Add why the chokepoint matters — one place for RBAC, admission, audit, encryption at rest — and explain that it is stateless so HA is just replicas.

for a senior

Discuss the watch cache absorbing read load, storage version versus served versions, and precisely what degrades when the API is unavailable versus what keeps running.

for a principal

Frame it as the consistency and policy boundary of the whole system, and reason about it as the scaling chokepoint that etcd sits behind.

## The one door in Kubernetes is built as a set of loosely coupled components that share no state directly. The shared state lives in etcd, and exactly one process is allowed to touch it: kube-apiserver. Every other participant — the `kubectl` on your laptop, kube-scheduler, kube-controller-manager, cloud-controller-manager, every kubelet, every custom operator, the dashboard, CI pipelines — is an ordinary HTTPS client of the same REST API. There is no back channel and no privileged internal protocol; the scheduler uses the same endpoints available to you, differing only in credentials and permissions. ## What it actually does for a request For a write it terminates TLS, authenticates the caller (client certificate, service-account bearer token, OIDC token, or an authentication webhook), authorises the verb-on-resource (RBAC in almost all clusters, plus the Node and webhook authorisers), decodes the body and applies **defaulting**, runs **mutating admission** (built-in plugins and any registered mutating webhooks), runs **schema validation** on the resulting object, runs **validating admission** (plugins, validating webhooks, and CEL-based admission policies), and finally persists the object to etcd with optimistic concurrency. Then it records an audit entry and dispatches the change to everyone watching that resource. For a read it serves GET and LIST, frequently from an in-memory watch cache rather than hitting etcd, and it serves long-lived `watch` streams that push subsequent changes. It also hosts **discovery** (`/api`, `/apis`, the OpenAPI schema) so clients can learn which groups, versions, kinds, verbs and subresources exist, which is how `kubectl` knows what `kubectl get widgets` should do. ## Why the sole-etcd-client rule matters Every cross-cutting concern gets a single implementation at that chokepoint: - **Security.** Authentication, authorisation and admission are unbypassable. If controllers could write to etcd directly, RBAC would be advisory and any compromised controller could forge any object. - **Correctness.** Validation and defaulting run exactly once, in one place, so no client can persist an object that violates the schema — including objects written by components you did not author. - **Auditability.** One audit log captures who did what to which object, because there is no other path. - **Data-at-rest encryption.** Secret encryption is implemented in the apiserver's storage layer; direct etcd writers would bypass it. - **Decoupling.** The serialized form in etcd is a *storage version*, and the apiserver converts between it and whichever API version the client asked for. That indirection is what lets Kubernetes promote APIs from beta to GA without rewriting stored data or breaking old clients. - **Load control.** Because every request funnels through it, the apiserver can prioritise and throttle traffic (API Priority and Fairness) and absorb read load in its watch cache instead of hammering etcd. ## Stateless by design kube-apiserver holds no durable state of its own; its caches are derived. That makes horizontal scaling and HA trivial in principle: run three replicas behind a load balancer (or, for in-cluster traffic, the `kubernetes` Service in the `default` namespace, whose endpoints point at them). Any replica can serve any request. Contrast this with kube-scheduler and kube-controller-manager, which run active/passive with leader election because two actors reconciling the same objects would conflict. What is *not* free about scaling: each replica maintains its own watch cache, so memory grows with object count per replica, and every replica keeps its own watch connections to etcd. The apiserver scales out for request throughput; etcd remains the shared limit for write throughput. ## What happens when it is down A useful test of understanding. If the apiserver is unavailable: no `kubectl`, no new deployments, no scheduling, no controller reconciliation, no scaling, and no Service/Endpoint updates. But **already-running Pods keep running** — the kubelet has its local view and container runtimes need no control plane to keep processes alive, and kube-proxy keeps serving traffic with the dataplane rules it already programmed. The cluster stops changing rather than stops working. That distinction is the classic follow-up. ## Where it sits physically In kubeadm-style clusters it runs as a static Pod on each control-plane node, defined by a manifest in `/etc/kubernetes/manifests/` that the kubelet starts without any API involvement — solving the bootstrap chicken-and-egg problem. In managed offerings (EKS, GKE, AKS) the provider runs it and you only ever see the endpoint.

  • If every kube-apiserver replica is down, what stops working and what keeps running?
    All changes stop: no kubectl, no scheduling, no controller reconciliation, no scaling, no Endpoint updates, no new Pods. Already-running Pods keep running because the kubelet works from its existing view and the container runtime needs no control plane, and kube-proxy keeps forwarding traffic using the dataplane rules it already programmed. Failures are not repaired though — a crashed Pod will not be rescheduled until the API is back.
  • Why do the scheduler and controller-manager use leader election while the apiserver does not?
    The apiserver is stateless and idempotent per request, so any replica can serve any call and all replicas can be active behind a load balancer. The scheduler and controller-manager are active reconcilers; two instances acting on the same objects would issue conflicting writes and duplicate work, so they elect a single leader through a Lease object and the others stand by.

It is the only teller window at the bank. Nobody reaches into the vault themselves, so identity checks, paperwork rules, and the transaction log all exist in exactly one place.

saying these in an interview costs you the question

  • Saying controllers or kubelets read etcd directly
  • Describing the apiserver as merely a proxy that forwards requests to etcd without validation
  • Claiming the apiserver needs leader election for HA
  • Believing running Pods die the instant the control plane goes down
  • Treating admission control as something clients can opt out of

context

open as a page

Trace what kube-apiserver does with an HTTPS POST of a Deployment manifest, from the moment the request arrives until the object is durably stored. Name the stages in order and say what each can reject.

level: middleimportance: must knowfreq 60%

basics

~10 s

TLS, then authentication (who), authorisation (may they), decode plus defaulting, mutating admission, schema validation, validating admission, then a write to etcd guarded by optimistic concurrency, followed by audit and watch fan-out.

open as a page

Kubernetes manifests declare fields such as `apiVersion: apps/v1` and `kind: Deployment`. Explain how group, version and kind map onto the REST URLs the API server exposes, and how the same stored object can be served at more than one version.

level: middleimportance: should knowfreq 45%

basics

~10 s

apiVersion is group/version and kind is the type; together they select a REST path like /apis/apps/v1/namespaces/ns/deployments. The server stores one storage version and converts to whichever served version a client requests.

open as a page

Kubernetes lets you add API endpoints either with a CustomResourceDefinition or through the API aggregation layer and an APIService object. Explain how aggregation works inside kube-apiserver and when you would choose it over a CRD.

level: seniorimportance: should knowfreq 30%

basics

~20 s

An APIService registers a group/version and points at a Service; kube-apiserver then proxies every request for that group/version to that extension server, which implements storage and semantics itself. Choose it when you need non-etcd storage or custom verbs.

open as a page

A client calls the Kubernetes API with `?watch=true&resourceVersion=<rv>` and after some idle time receives HTTP 410 Gone with "too old resource version". Explain the contract the API server's watch endpoint offers and what the client is required to do.

level: seniorimportance: should knowfreq 40%

basics

~20 s

A watch streams change events occurring after a given resourceVersion, which the server can only replay from a bounded in-memory window. Once that version ages out, the server returns 410 and the client must re-list to get a fresh resourceVersion and re-watch.

open as a page

A large cluster's control plane is degrading: API latency spikes, controllers time out, and kube-apiserver memory keeps climbing. How do you reason about its capacity, and how do you protect it from expensive clients?

level: principalimportance: nice to knowfreq 25%

basics

~20 s

Classify the load: expensive unpaginated LISTs dominate memory, watches dominate steady state. Add replicas for throughput, then constrain clients with pagination, selectors and informers, and enforce fairness with API Priority and Fairness so one client cannot starve the rest.

open as a page