skip to content

A scan finds etcd's client port 2379 on your Kubernetes control-plane nodes reachable from the worker subnet. What is at risk, and how do you lock it down?

level: seniorimportance: should knowfreq 42%

answer

  1. below the API server's checks
  2. two listeners, two TLS sets
  3. which CA does etcd trust
  4. one member at a time
  5. assume Secrets were read

basics

~20 s

Direct etcd access bypasses Kubernetes authentication, RBAC and admission, so a caller can read or rewrite every object, including Secrets. Require client and peer certificates signed by a dedicated etcd CA, and firewall 2379 and 2380 to control-plane hosts only.

solid answer

~40 s

etcd stores the whole cluster state. Anyone who can talk to it directly bypasses everything the API server enforces: no authentication, no RBAC, no admission, no API audit log. They can read every Secret, which is plaintext unless encryption at rest is configured, or write a privileged pod or a ClusterRoleBinding straight into storage. To fix it, etcd must run with `--client-cert-auth=true` and `--trusted-ca-file`, and with `--peer-client-cert-auth=true` and `--peer-trusted-ca-file` for member-to-member traffic. Both should trust a **dedicated etcd CA**, not the cluster CA, so a kubelet's certificate cannot open etcd. The API server connects using `--etcd-cafile`, `--etcd-certfile` and `--etcd-keyfile`. On top of that, firewall 2379 so only the API servers reach it, and 2380 so only the other etcd members do.

code

yaml · 25 lines
yaml
apiVersion: v1
kind: Pod
metadata:
  name: etcd-cp-2
  namespace: kube-system
spec:
  hostNetwork: true
  containers:
  - name: etcd
    image: registry.k8s.io/etcd:3.7.0-0
    command:
    - etcd
    - --name=cp-2
    - --data-dir=/var/lib/etcd
    - --listen-client-urls=https://127.0.0.1:2379,https://10.40.0.12:2379
    - --advertise-client-urls=https://10.40.0.12:2379
    - --listen-peer-urls=https://10.40.0.12:2380
    - --cert-file=/etc/kubernetes/pki/etcd/server.crt
    - --key-file=/etc/kubernetes/pki/etcd/server.key
    - --client-cert-auth=true
    - --trusted-ca-file=/etc/kubernetes/pki/etcd/ca.crt
    - --peer-cert-file=/etc/kubernetes/pki/etcd/peer.crt
    - --peer-key-file=/etc/kubernetes/pki/etcd/peer.key
    - --peer-client-cert-auth=true
    - --peer-trusted-ca-file=/etc/kubernetes/pki/etcd/ca.crt

go deeper

for a junior

Remember that etcd holds every cluster object and that only the API server should ever talk to it.

for a middle

Explain the client and peer ports, the certificate flags on each, and the API server's etcd client flags.

for a senior

Show the response to an exposure finding: close the network path, fix TLS one member at a time while keeping quorum, then rotate what may have been read.

for a principal

Own the CA layout and the etcd blast radius: separate CAs, snapshot handling and whether etcd runs on dedicated hosts or alongside the control plane.

## Why etcd is the real target **etcd** is the distributed key-value store that holds all Kubernetes objects, under keys starting with `/registry/`. The API server is the only component that should talk to it, and it enforces the rules on the way in: 1. **Authentication**: who is calling. 2. **Authorization**, with Node and RBAC: may they do this. 3. **Admission**, including Pod Security and policy webhooks: is this object acceptable. 4. **Audit**: record what happened. A client that talks to etcd **directly** skips all four. With plain read access it can list every Secret. If encryption at rest is not configured, those values are plaintext in storage. With write access it can insert a privileged pod, a ClusterRoleBinding to `cluster-admin`, or a changed ServiceAccount, and none of that appears in the API server's audit log. **Access to etcd amounts to full control of the cluster.** ## The two listeners | Port | Purpose | Who should reach it | etcd TLS flags | |---|---|---|---| | **2379** | client traffic | the API servers only | `--cert-file`, `--key-file`, `--client-cert-auth=true`, `--trusted-ca-file` | | **2380** | peer (member-to-member) traffic | the other etcd members only | `--peer-cert-file`, `--peer-key-file`, `--peer-client-cert-auth=true`, `--peer-trusted-ca-file` | Without `--client-cert-auth`, and unless etcd's own user authentication is turned on, etcd uses TLS for encryption but accepts any client that connects. Without the peer flags, members do not check that the other side of peer traffic is a real member of the cluster. kubeadm's local etcd sets all of these. External or hand-built etcd clusters often do not. ## The dedicated CA matters `--client-cert-auth` trusts **every** certificate signed by `--trusted-ca-file`. If that file is the **cluster CA**, then every kubelet client certificate, every admin kubeconfig certificate and every certificate issued through the cluster's signing API is also an etcd client certificate. A single compromised worker can then read etcd. kubeadm avoids this by using a separate **etcd CA** (`/etc/kubernetes/pki/etcd/ca.crt`). The API server gets its own client certificate from that CA: - `--etcd-servers`: the member URLs, over `https://` - `--etcd-cafile`: the etcd CA, used to verify etcd's server certificate - `--etcd-certfile` / `--etcd-keyfile`: the API server's etcd client identity The etcd CA should sign only etcd server, peer, health-check and API-server client certificates. ## Working the finding Take the scan result on a 3-control-plane, 27-worker cluster: 1. **Confirm exposure.** From a worker, check whether a TLS connection to `:2379` succeeds without a client certificate, and whether one succeeds with a kubelet's certificate. Either result is a critical finding. 2. **Close the network path first.** Host firewall or security-group rules should allow 2379 only from the three control-plane addresses and 2380 only between the three members. This does not require an etcd restart. 3. **Fix the TLS settings.** Update each member's flags, and its certificates if they come from the wrong CA, **one member at a time**. Check `etcdctl endpoint health` and wait for quorum (2 of 3) before moving to the next member. Losing two members at once loses quorum, and API server writes fail until it returns. 4. **Check what was already exposed.** If 2379 was open without client authentication, assume Secrets were read. Rotate every credential stored in Secrets, including any legacy ServiceAccount token Secrets and cloud credentials, and look for unexpected RBAC bindings or pods. Direct etcd writes never reached the API audit log, so look at the objects themselves. ## What encryption at rest does and does not give you With `--encryption-provider-config`, Secret values are encrypted before the API server writes them to etcd. So someone who reads etcd directly gets ciphertext for Secrets. **It does not protect anything else**: other resource types, and the ability to **write** objects, stay exposed unless they are also encrypted. Encryption at rest is extra protection behind client-certificate authentication, not a replacement for it. ## Other items - Take **etcd snapshots** with the same care: a snapshot file contains everything etcd stores. - Keep the etcd data directory owned by the etcd user with restrictive permissions. kube-bench's etcd and control-plane checks cover the flags and the file ownership. - Monitor for client connections to 2379 from anything other than API server addresses.

  • Why is pointing etcd's --trusted-ca-file at the cluster CA a problem even with client-cert-auth on?
    etcd accepts any certificate that CA signed. The cluster CA also signs kubelet client certificates and admin kubeconfigs, so a compromised worker could use its own certificate to read or write etcd directly. A dedicated etcd CA keeps the set of valid etcd clients limited to the API servers and the members themselves.
  • Does enabling encryption at rest make an exposed etcd acceptable?
    No. Encryption at rest covers only the resources listed in the `EncryptionConfiguration`, usually Secrets, and only for reads. An attacker with network access and no client-certificate check can still read other objects and write any object, such as a ClusterRoleBinding, without passing through RBAC or admission.

Direct etcd access is like walking into a bank's vault through the loading dock instead of the teller window: the ID check, approval rules and camera logs all belong to the teller window.

saying these in an interview costs you the question

  • Believes etcd access is still filtered by Kubernetes RBAC
  • Thinks TLS encryption alone authenticates etcd clients
  • Trusts the cluster CA for etcd client certificates
  • Treats encryption at rest as a substitute for etcd client authentication
  • Restarts all etcd members at once to apply new flags