In a Kubernetes control plane, what is etcd, what exactly is stored in it, and which components are allowed to talk to it?
answer
- single source of truth, only apiserver talks to it
- /registry/<resource>/<ns>/<name>, protobuf
- MVCC revision -> resourceVersion
- stores metadata, not logs/images/metrics
- lose etcd without snapshot = lose cluster
basics
~20 setcd is the cluster's only persistent database: a distributed, strongly consistent key-value store holding every API object (Pods, Deployments, Services, ConfigMaps, Secrets, RBAC, node state). Only kube-apiserver connects to it; everything else reads and writes through the API server.
solid answer
~50 setcd is a distributed key-value store that acts as Kubernetes' **single source of truth**. Everything the API exposes lives there: Pods, Deployments, ReplicaSets, Services, ConfigMaps, Secrets, ServiceAccounts, RBAC rules, CRDs and their custom resources, plus Node objects and Lease heartbeats. Key points: - **Only kube-apiserver connects to etcd.** Scheduler, controller-manager, kubelets and `kubectl` never speak etcd directly; the API server owns validation, admission, authorization and encoding. - **Keys are path-shaped**, roughly `/registry/<resource>/<namespace>/<name>`, values serialized as protobuf. - **Both spec and status** (desired and observed state) are stored — but not logs, metrics or images. - etcd is **MVCC-based**: every write bumps a global revision, and clients can watch a key prefix from a revision, which is how the API server powers `watch`. Since it is the only stateful control-plane component, losing etcd without a snapshot loses the cluster.
code
bash · 6 linesETCDCTL_API=3 etcdctl \
--endpoints=https://127.0.0.1:2379 \
--cacert=/etc/kubernetes/pki/etcd/ca.crt \
--cert=/etc/kubernetes/pki/etcd/server.crt \
--key=/etc/kubernetes/pki/etcd/server.key \
get /registry/pods/default/ --prefix --keys-onlygo deeper
Be able to say etcd is the cluster's key-value database, that it holds all API objects, and that only the API server talks to it.
Add the key layout, protobuf serialization, and the MVCC revision to resourceVersion link that makes watch and optimistic concurrency work.
Emphasize operational consequences: etcd is the only stateful component, its disk latency bounds API latency, and access is gated by mutual TLS plus API-server-level encryption at rest.
Frame it as the cluster's consistency boundary and blast radius: what a single strongly consistent store implies for scaling, multi-cluster topology, splitting Events, DR objectives, and why no component gets its own etcd credentials.
## What etcd is etcd is an open-source distributed key-value store designed for configuration data: small values, moderate write rates, very strong consistency. Kubernetes uses it as the backing store for the whole API. If an object is not in etcd, it does not exist as far as the cluster is concerned. "Strongly consistent" means a successful write is visible to every subsequent read on any member — there is no window where two clients see different cluster state. That is what lets controllers reason about the cluster as one coherent picture. It is achieved with the Raft consensus protocol: one elected leader, writes replicated to a majority before they are acknowledged. ## What is stored Everything served by the Kubernetes API: - Workload objects: Pods, Deployments, ReplicaSets, StatefulSets, DaemonSets, Jobs, CronJobs. - Networking and config: Services, EndpointSlices, Ingresses, ConfigMaps, Secrets. - Identity and policy: ServiceAccounts, Roles, RoleBindings, admission policy objects. - Cluster state: Node objects, Lease objects used for node heartbeats and leader election, Events (often given a separate etcd instance because they churn). - Extensions: CustomResourceDefinitions and every custom resource an operator creates. What is **not** stored: container images (registries), container logs (node disk), metrics (the metrics pipeline), and the contents of PersistentVolumes (the storage system). etcd holds the *metadata* describing those things. ## Key layout Keys are hierarchical paths under a prefix, conventionally `/registry`: - `/registry/pods/default/web-7d9f` — a namespaced object - `/registry/nodes/node-1` — a cluster-scoped object - `/registry/deployments/prod/api` Values are the serialized object, protobuf by default for built-in types (custom resources are stored as JSON). Because keys are prefixes, "list all Pods in namespace prod" becomes a range read over `/registry/pods/prod/`, and "watch all Pods" becomes a watch on the `/registry/pods/` prefix. This layout is an internal detail — you should never write to it directly — but it explains why the API server can serve list and watch efficiently. ## MVCC, revisions and watch etcd is multi-version: it does not overwrite in place. Every mutating transaction increments a cluster-wide **revision** counter and stores a new version of the key tagged with that revision. Old versions remain until compaction removes them. This is the mechanism behind Kubernetes' `resourceVersion`: the value you see on an object derives from the etcd revision. A client can say "watch this prefix starting at revision N" and etcd replays every change since N, then streams new ones. The API server uses exactly that to build its watch cache, which in turn feeds informers in controllers and kubelets. It is also how optimistic concurrency works: an update carrying a stale `resourceVersion` is rejected with a conflict, so two controllers cannot silently clobber each other. ## Why only the API server connects Routing all access through kube-apiserver means one place enforces authentication, authorization (RBAC), admission control, defaulting, validation, quota and audit logging. It also means the storage backend is swappable and clients are insulated from its encoding. A component with direct etcd credentials would bypass every one of those controls — which is why etcd is normally bound to localhost or a private control-plane network, protected by mutual TLS with a client certificate only the API servers hold, and why encryption-at-rest for Secrets is configured at the API server, not in etcd. ## Practical consequences Because etcd is the only stateful control-plane component, control-plane nodes are otherwise disposable: API server, scheduler and controller-manager can be rebuilt from config. etcd cannot — restoring it from a snapshot is the entire cluster disaster-recovery story. Its health also gates everything: when etcd is slow (usually disk fsync latency), API requests slow down, leader elections flap and controllers fall behind, so etcd latency is the first metric to check when a cluster "feels slow".
- Where does the resourceVersion on a Kubernetes object come from, and what is it used for?It derives from etcd's global MVCC revision at the time the object was last written. Clients use it for optimistic concurrency: an update carrying a stale resourceVersion is rejected with a 409 Conflict instead of overwriting a newer change. Watchers also use it as a resume point, asking to stream all changes since a given revision.
- If etcd stores Secrets, are they encrypted?Not by default — Secrets are stored base64-encoded, which is encoding, not encryption. You enable encryption at rest with an EncryptionConfiguration on kube-apiserver (for example aescbc or a KMS provider), which encrypts values before they reach etcd. Independently, etcd's disk should be encrypted and its client and peer traffic protected with mutual TLS.
etcd is the cluster's ledger and kube-apiserver is the only clerk allowed at the ledger — everyone else files a form with the clerk, who checks identity and rules before writing a line.
saying these in an interview costs you the question
- Saying kubelet or the scheduler connects to etcd directly — only kube-apiserver does
- Claiming container logs, metrics or images live in etcd
- Assuming Secrets in etcd are encrypted by default (they are only base64-encoded)
- Describing etcd as eventually consistent — it is linearizable via Raft
- Thinking etcd is just a cache that can be wiped and rebuilt from the nodes