skip to content

With the Kubernetes Node authorizer enabled, what does the NodeRestriction admission plugin additionally enforce, and which compromised-node attack does it blunt?

level: seniorimportance: nice to knowfreq 28%

answer

  1. authorizer cannot read the body
  2. own Node, own pods only
  3. mirror pods bound to self
  4. the protected label prefix
  5. labels steer the scheduler

basics

~20 s

NodeRestriction checks the content of a kubelet's writes. A kubelet may change only its own Node and pods bound to it, may not change its taints, and may not set node-restriction.kubernetes.io/ labels, so a stolen node cannot pull sensitive workloads onto itself.

solid answer

~50 s

The Node authorizer decides which API operations a kubelet (`system:node:<name>`) may attempt. It cannot look inside the object being written. **NodeRestriction** is the admission plugin that does. It rejects a kubelet's change to any Node other than its own. It allows a kubelet to create only **mirror pods**, bound to itself and owned by its own Node, and to update the status only of pods bound to it. It forbids changing taints after registration, and it forbids adding or changing labels under `node-restriction.kubernetes.io/`, as well as most other `kubernetes.io`/`k8s.io` labels. That blocks a known attack: a compromised node relabels itself to match a sensitive workload's `nodeSelector`, or removes its taint, so the scheduler places that workload, and its Secrets, on the attacker's machine. The fix is to put workload-isolation labels under the `node-restriction.kubernetes.io/` prefix and to keep the plugin enabled.

code

bash · 2 lines
bash
kubectl label node worker-19 node-restriction.kubernetes.io/workload-tier=restricted
kubectl get nodes -L node-restriction.kubernetes.io/workload-tier

go deeper

for a junior

Remember that kubelets have their own API identity and that one plugin limits what that identity can change.

for a middle

Explain why authorization cannot inspect a request body and what NodeRestriction checks on Node, label and pod writes.

for a senior

Walk through the relabel-and-attract attack and show the label-prefix fix. Be clear about what the plugin does not protect once a node is lost.

for a principal

Place node compromise in the cluster threat model. Decide which workloads need dedicated, admin-labelled node pools and how quickly a compromised node is cordoned and replaced.

## Two layers for one identity Each kubelet authenticates as the user `system:node:<nodeName>` in the group `system:nodes`. Two separate pieces of the API server limit what that identity can do: | Layer | Stage | What it can see | What it limits | |---|---|---|---| | **Node authorizer** | authorization | verb, resource, namespace, name | which objects a kubelet may read or write, such as only Secrets referenced by pods bound to it | | **NodeRestriction** | admission | the full object being written | what a kubelet may put **inside** the objects it is allowed to write | Authorization never sees the request body, so "this kubelet may update Node objects" cannot be narrowed to "only its own Node, and not its taints" there. **Admission** sees the body, and NodeRestriction is the built-in plugin that checks it. It is not in the kube-apiserver's default-on plugin set. You enable it with `--enable-admission-plugins=NodeRestriction`, which kubeadm does for you. ## What the plugin enforces For requests from a node identity, NodeRestriction checks: - **Node objects.** A kubelet may change only the Node whose name matches its own. It cannot change `spec.taints` after the Node exists, cannot change owner references, and cannot set forbidden labels. - **Labels.** It may not add, change or remove any label whose prefix is, or ends with, `node-restriction.kubernetes.io/`. It may not change other `kubernetes.io` / `k8s.io` labels either, except a short allowed list: hostname, OS, architecture, instance type, topology zone and region, plus the `kubelet.kubernetes.io/` and `node.kubernetes.io/` prefixes. - **Pods.** A kubelet may create only **mirror pods**: pods that carry the mirror-pod annotation, have `spec.nodeName` set to itself, and have one controller owner reference to its own Node. Those pods may not reference other API objects, such as Secrets or ConfigMaps. It may update status and evict only for pods bound to itself. - **Other node-scoped objects.** Related checks cover the node's Lease, its CSINode, PersistentVolumeClaim status updates, ServiceAccount token requests, ResourceSlices and certificate signing requests. ## The attack it blunts Say an attacker gets root on one worker in a 3-control-plane, 27-worker cluster. They now hold that kubelet's credentials. Without NodeRestriction: 1. They list the Nodes and see that the webhook-delivery dispatcher, which holds signing keys for customer callbacks, is pinned with `nodeSelector: workload-tier=restricted` and tolerates a `dedicated=restricted:NoSchedule` taint. 2. They add `workload-tier=restricted` to their own Node's labels, and could also remove taints from it. 3. When a dispatcher pod is rescheduled, the scheduler may place it on the compromised node. The Node authorizer then correctly allows that kubelet to read the pod's Secrets, because the pod is now bound to it. With NodeRestriction on, the taint change in step 2 is rejected, but adding a label like `workload-tier` still works, because that is a custom label outside the protected prefixes. That is why the second half of the fix is naming: **use `node-restriction.kubernetes.io/workload-tier=restricted`**, which a kubelet cannot set on itself. Only an administrator can apply it. ```yaml apiVersion: v1 kind: Pod metadata: name: webhook-dispatcher-0 spec: nodeSelector: node-restriction.kubernetes.io/workload-tier: restricted containers: - name: dispatcher image: registry.example.internal/webhook-dispatcher@sha256:3f1c0b2a9d4e5f60718293a4b5c6d7e8f90a1b2c3d4e5f60718293a4b5c6d7e8 ``` ## Limits worth stating - NodeRestriction limits what a stolen kubelet identity can **write through the API**. It does nothing about what root on the node can reach locally: containers already on that node, their mounted Secrets, and the runtime socket. - Labels applied through the kubelet's `--node-labels` flag at registration follow the same rules. A kubelet cannot give itself a protected label, so an administrator or a provisioning controller has to apply it. - The plugin covers only node identities. Other callers, for example a ServiceAccount allowed to patch Nodes, are controlled by RBAC alone.

  • Why can a kubelet still set a custom label like team=payments on its own Node when NodeRestriction is enabled?
    The plugin forbids only the `node-restriction.kubernetes.io/` prefix and `kubernetes.io`/`k8s.io` labels that are not on its allowed list. Custom prefixes are outside its check. Any label used to place sensitive workloads should therefore sit under the protected prefix.
  • Does NodeRestriction protect Secrets already mounted on a compromised node?
    No. Those Secrets belong to pods legitimately bound to that node, so the Node authorizer allows the kubelet to read them. Root on the node can read them from the mounts anyway. NodeRestriction only stops the node from attracting more sensitive workloads and from modifying other nodes' objects.

saying these in an interview costs you the question

  • Thinks the Node authorizer alone stops a kubelet relabelling its own Node
  • Believes NodeRestriction also stops root on the node reading mounted Secrets
  • Puts workload-isolation labels under an unprotected custom prefix
  • Assumes NodeRestriction limits ServiceAccounts that can patch Nodes