skip to content

Why does kube-bench have to run on the node with host PID and host file access?

level: seniorimportance: should knowfreq 44%

answer

  1. the flags, not the manifest
  2. no API object for a file mode
  3. hostPID exposes the node process table
  4. per node, so per DaemonSet
  5. detective only, it blocks nothing

basics

~20 s

Most Kubernetes benchmark checks read the flags a component was actually started with and the ownership and mode of files on disk. Neither is exposed by the Kubernetes API, so the runner needs the node's process table and filesystem.

solid answer

~50 s

The checks assert things like `--anonymous-auth=false` on the API server, `--read-only-port=0` on the kubelet, and specific ownership and modes on `/etc/kubernetes/manifests` and the kubelet config. None of that is a Kubernetes API object: you cannot ask the API server which flags it was launched with, and there is no resource describing a file's permission bits. So kube-bench runs where the evidence is - as a pod with `hostPID: true` so its `/proc` shows the node's processes, plus read-only `hostPath` mounts for the config and certificate paths, and per node, because every node has its own kubelet flags and its own file modes. That is also why the result is detective: it blocks nothing and mutates nothing, it reports `[PASS]`, `[FAIL]`, `[WARN]` and `[INFO]` after the fact. Without the host mounts the checks inspect the container's own empty view and the output is meaningless.

go deeper

for a junior

Know that this class of tool inspects a machine that is already running, not a change someone proposed, and that it needs access to the node's processes and files to do it.

for a middle

Explain which facts live only on the node - process arguments and file ownership and modes - and what hostPID and read-only hostPath mounts give the pod that a normal pod does not have.

for a senior

Show the operational judgment: per-node execution, WARN treated as unassessed rather than passed, control-plane sections marked provider-owned on managed clusters, and results routed to whoever owns the node.

for a principal

Own where the shared-responsibility line falls and how you evidence the half you do not operate, so the estate's compliance story is complete without your runner pretending to have checked a control plane it cannot see.

### What the CIS Kubernetes checks actually assert Most of the control-plane and node controls in the Kubernetes benchmark are assertions about two things on a specific machine: 1. **The command line of a running process.** Whether `kube-apiserver` was started with `--anonymous-auth=false`, what `--authorization-mode` it was given, whether the kubelet runs with `--read-only-port=0`, what `--client-ca-file` points at. 2. **The ownership and mode of files on disk.** The static pod manifests under `/etc/kubernetes/manifests`, the kubelet's config file and its kubeconfig, the certificate and key files, the etcd data directory. Neither of those is a Kubernetes API object. You cannot ask the API server what flags it was started with, and there is no resource to `kubectl get` for the mode of `/etc/kubernetes/manifests/kube-apiserver.yaml`. The only way to answer is to be on the machine and look. ### Why it runs as a privileged pod kube-bench is a binary plus benchmark definitions in YAML. Run on the host it simply reads `/proc` and the filesystem. Run the usual way — as a Job in the cluster it is assessing — it needs the host's view of both, so its pod spec typically sets: - `hostPID: true`, so the container's `/proc` shows the node's processes and kube-bench can read the API server's, controller-manager's, scheduler's and kubelet's actual command lines; - read-only `hostPath` mounts for the paths the checks care about — `/etc/kubernetes`, the kubelet config, the etcd data directory, systemd unit directories; - a node selector or a DaemonSet, because the answer is per-node: every node has its own kubelet flags and its own file permissions, and a single pod on one node tells you about one node. Strip those away and the checks do not quietly skip — they look at the container's own `/proc` and an empty filesystem and produce results that describe nothing. ### Manifest intent versus running reality This is the line that defines the whole category of tool. A file in git is a statement of intent. The process running on node 47 is the fact. They diverge for entirely ordinary reasons: someone edited a static pod manifest by hand during an incident and never put it back; a node bootstrap tool or a systemd drop-in appended flags the repository never saw; a package update replaced a config file and reset its mode; a node was built from an older image than the rest of the fleet. A profile runner exists precisely to measure the second thing, on each machine, as it is now. That is also why its results are *detective*: kube-bench blocks nothing and mutates nothing. It runs after the fact and produces a report. ### Reading the output kube-bench emits `[PASS]`, `[FAIL]`, `[WARN]` and `[INFO]` per check, grouped by benchmark section, with remediation text attached to the failures. The result to be careful about is `[WARN]`: it generally marks a check the tool cannot decide on its own — one that needs a human to confirm a value, or one whose target it could not inspect. Counting WARN as PASS inflates a compliance number with checks nobody performed. `--json` output makes the results ingestible; the checks are grouped by target, and the targets are the machine roles the benchmark defines — control-plane components, etcd, the node, and policy-level checks. ### The managed-cluster case On a managed cluster you do not get a shell on the control plane and there is no API server process on any node you can see. The control-plane sections of the benchmark are therefore **not applicable to you** rather than failing — the provider operates that layer, and evidence for it comes from the provider, not from your runner. What remains yours is the node role (kubelet flags, kubelet config file permissions, certificate rotation settings) and the policy-level checks, and kube-bench ships benchmark definitions for the managed offerings that scope the checks accordingly. Forcing the control-plane checks and reporting a page of failures is a misread of the shared-responsibility line, and an interviewer will notice. ### The interview point The question behind the question is whether you understand *what surface a control can be evaluated against*. Some assertions are answerable from a document; some are answerable only from a machine. "The API server must not enable anonymous auth" is the second kind once the cluster exists, and that is why this class of tool needs privileges on the target, runs per node, and produces a report rather than a verdict on a change.

  • On a managed cluster the control-plane checks cannot run. What do you conclude?
    That those checks are not applicable to you rather than failing. There is no API server process on any node you can reach, and the provider operates that layer, so evidence for it comes from the provider under the shared-responsibility split. You run the node and policy sections - kubelet flags, kubelet config file permissions, certificate settings - using the benchmark definitions scoped to the managed offering, and you report the control-plane section as provider-owned instead of as a page of red.
  • What does a WARN result mean, and why is it dangerous in a compliance number?
    It generally marks a check the tool cannot decide by itself - one needing a human to confirm a value, or one whose target it could not inspect. It is closer to 'not assessed' than to 'passed'. Rolling WARN into the pass column produces a compliance figure built partly from checks nobody performed, which is the fastest way to lose an auditor's confidence in the whole report.
  • Why is a per-node run necessary rather than one run per cluster?
    Because the assertions are per-machine facts. Node 47 can have a kubelet started with different flags, a config file with a different mode, or a node image a version behind the rest of the fleet, and one pod on one node tells you only about that node. Running as a DaemonSet, or once per node, is what makes the result a statement about the cluster rather than about a sample of one.

saying these in an interview costs you the question

  • Counts WARN results as passes
  • Assumes a pod without host mounts can read kubelet flags
  • Says the checks can be answered from the static pod manifests in git
  • Thinks the runner prevents the misconfiguration it finds
  • Expects control-plane checks to apply on a managed cluster

context