skip to content

You inherit a 3-control-plane, 27-worker self-managed Kubernetes cluster that fails many kube-bench control-plane and kubelet checks. How do you plan and sequence hardening without causing an outage?

level: principalimportance: should knowfreq 34%

answer

  1. rank by what an attacker gets
  2. every open door has users
  3. AlwaysAllow removal is the cliff
  4. waves with workload abort signals
  5. bake it into provisioning

basics

~20 s

Rank kube-bench failures by what an attacker gains, find who depends on each open setting before closing it, roll changes one control-plane node and small kubelet waves at a time, and bake the result into provisioning.

solid answer

~50 s

I treat kube-bench as a measurement and rank its FAILs by what an attacker gets: unauthenticated or `AlwaysAllow` access on the API server or kubelets and an exposed etcd first, then `NodeRestriction` and encryption at rest, then items like `--profiling`. Before closing any door I find its users: audit-log requests by `system:anonymous`, collectors on kubelet port 10255, and identities that only work because of `AlwaysAllow`. Moving to `Node,RBAC` is the riskiest step, so the bindings must exist and pass `kubectl auth can-i --as` first. I change one API server at a time, and kubelets in waves of three or four, with abort criteria from real workloads, such as the webhook-delivery dispatcher's p99 passing 310 ms. Then I bake every change into the node image and installer config and rerun kube-bench on a schedule, so upgrades and rebuilt nodes cannot undo it.

go deeper

for a junior

Remember that kube-bench reports on control-plane and node settings and that each failing item has a concrete flag or file behind it.

for a middle

Explain which findings allow unauthenticated control and why changing a flag can break existing users such as probes and collectors.

for a senior

Show the rollout: dependents found first, one control-plane node at a time, kubelet waves with checks, and quorum kept for etcd.

for a principal

Argue the prioritisation and the rebuild-or-fix decision, set abort criteria from real workloads, and make the baseline hold through upgrades and node replacement.

## Start from exposure, not from the report order **kube-bench** runs the CIS Kubernetes Benchmark checks against a node's component flags, config files and file permissions, and reports `PASS`, `FAIL`, `WARN` or `INFO`. On an inherited cluster it may return dozens of FAILs. Working through them in report order is the most common mistake. Rank them by **what an attacker gets** and **how easily they can reach it**: | Tier | Examples | Why it comes first | |---|---|---| | 1. Unauthenticated control | kube-apiserver `AlwaysAllow`, kubelet anonymous auth with `AlwaysAllow`, etcd without client-cert auth, open kubelet read-only port | Anyone on the network can take over a node or the cluster | | 2. Blast-radius limits | `NodeRestriction` missing, no encryption at rest, `--kubelet-certificate-authority` unset | Makes a single compromise worse | | 3. Defence in depth | `--profiling`, file ownership on manifests and PKI, audit settings | Real, but rarely how a cluster is lost | Items that kube-bench marks as manual, or that do not apply to your setup, go into an exceptions list with an owner and a reason. They are not silently ignored. ## Find the dependents before closing each door Every insecure setting that has been in place for a while has users. **Closing it is a contract change**, so find those users first: - **Anonymous access to the API server.** Look in the API audit log for requests by `system:anonymous`. Usually they are load-balancer and static-pod health probes, which can be kept by allowing anonymous access only on `/livez`, `/readyz` and `/healthz` through an `AuthenticationConfiguration`. - **Kubelet read-only port 10255.** Look at connection logs or `ss` on the workers. Usually it is an old metrics collector, which should move to 10250 with a ServiceAccount that has `nodes/metrics` and `nodes/stats`. - **`AlwaysAllow` on the API server.** This is the dangerous one. Every controller, CI job and person may be relying on permissions that no binding grants. Before the change: 1. Run with API audit logging for a representative period and list every distinct user, verb and resource. 2. Write RBAC bindings for each real need. 3. Check each identity with `kubectl auth can-i --list --as=<identity>`. 4. Change the flag on one API server and watch for 403s before doing the other two. ## Sequencing the rollout **Control plane (3 nodes).** Change one kube-apiserver manifest at a time. The kubelet restarts that static pod, and the load balancer uses the other two. Move on only when the restarted instance passes its health checks, shows no rise in 401 or 403 responses, and has been serving for a set time. For etcd, apply TLS changes one member at a time and never let quorum (2 of 3) drop. **Workers (27 nodes).** Change kubelet settings in **waves**: a canary node, then groups of three or four. Each wave: - cordon and drain if the change needs a restart that could disturb pods; - apply the `KubeletConfiguration` change and restart the kubelet; - check that an unauthenticated `curl` to port 10250 gets 401, that nothing listens on 10255, and that `kubectl logs` and `kubectl exec` still work through the API server; - uncordon, and watch workload signals before the next wave. **Abort criteria come from workloads, not only from components.** For example, the webhook-delivery dispatcher uses a Lease for leader election. If its p99 delivery latency goes above **310 ms** during a wave, or it loses its leader, stop and investigate before continuing. That catches failures that component health checks miss. ## Making it stay fixed A cluster fixed by hand drifts back: - Put the flags into the **installer's configuration**, such as kubeadm's `ClusterConfiguration` extra args and the kubelet config it distributes, so an upgrade does not overwrite hand edits. - Put the `KubeletConfiguration` into the **node image or provisioning**, so a replacement node starts hardened. - **Run kube-bench on a schedule** as a Job or DaemonSet and alert when a result changes from PASS to FAIL. Turning those results into formal compliance evidence is a separate practice. ## The tradeoffs a lead has to own - **Speed versus safety.** Tier 1 findings may justify a short, announced maintenance window instead of a slow rollout. Open kubelet APIs are routinely scanned for on reachable networks, so waiting weeks to avoid rollout risk is itself a risk. - **Rebuild versus fix in place.** If etcd or the kubelets were reachable without authentication for months, assume compromise. It can be cheaper to build a new cluster with the baseline and migrate workloads than to prove the old one is clean. - **Benchmark score versus actual risk.** Some benchmark items trade operability for little security gain. Record the decision and move on, rather than chasing a perfect score.

  • How do you find what depends on AlwaysAllow before switching the API server to Node,RBAC?
    Collect API audit events for a representative period and list each distinct user, group, verb and resource. For each one, write the least RBAC binding that covers it, then check it with `kubectl auth can-i --list --as=<user>`. Switch one API server first and watch for 403 responses from the identities on your list before changing the other two.
  • When would you rebuild the cluster instead of hardening it in place?
    When tier 1 exposure lasted long enough that compromise is plausible, for example etcd or kubelet 10250 reachable without authentication. Proving the cluster is clean is then harder than building a new one with the baseline from the start, restoring declarative workload config, and rotating every credential the old cluster held.

saying these in an interview costs you the question

  • Works through kube-bench findings strictly in report order
  • Switches AlwaysAllow to Node,RBAC on all API servers at once
  • Closes kubelet port 10255 without finding its current consumers
  • Treats a green kube-bench run as proof the cluster is secure
  • Fixes nodes by hand and does not change provisioning or the image
  • Watches only component health, not workload signals, during rollout waves