Why do cluster-wide Kubernetes admission policies exclude kube-system from their match scope?
answer
- scope is a decision, not leftover config
- who else writes objects in that namespace
- the engine is a workload too
- what happens after a full cluster restart
- excluded means unguarded, so gate it elsewhere
basics
~20 sThe rule was written for tenant workloads, not cluster infrastructure — and the policy engine's own Pods live in a system namespace, so an engine inside its own scope can block the Pods that would replace it.
solid answer
~50 sTwo reasons, and only one of them is about correctness. First, a rule aimed at application workloads is usually wrong for infrastructure: a rule such as "every Deployment with more than one replica must carry the labels a PodDisruptionBudget can select" says nothing sensible about a node agent DaemonSet or a control-plane component, so matching them buys no safety and breaks the cluster instead of a team's deploy. Second, and more serious, is the bootstrap loop: the policy engine is itself a workload, and if its own namespace is in scope then admitting the engine's replacement Pods requires the engine to be up. After a full restart, or when every replica is gone, the cluster cannot heal itself. So `kube-system`, the engine's own namespace, and the namespaces holding CNI, CSI and node agents are excluded by construction, and access to those namespaces is restricted instead.
go deeper
Be ready to say, in one breath, that a cluster-wide rule also catches system and infrastructure workloads, and that the engine enforcing the rule is itself a Pod that has to be admitted.
Explain the mechanism: namespaces are selected by labels, every namespace carries an automatic name label, and requests from controllers pass through admission exactly like a user's apply.
Show that you have thought about recovery. Walk through what happens after every engine replica is gone at once, and say which namespaces you exclude and what access control you put on them in exchange.
Own the opt-in versus opt-out framing: an exclusion list covers every namespace created tomorrow, a tenant-label match covers none of them, and which one you pick decides whether your coverage decays silently as the estate grows.
## What "scope" means here An admission policy does not automatically apply to everything. It carries a match block that answers three questions before any rule body runs: which resources and operations (the resource rules — API groups, versions, resources, verbs, cluster- or namespace-scoped), which namespaces (`namespaceSelector`, a label selector over the containing Namespace object), and which objects (`objectSelector`, a label selector over the object's own labels), plus optional CEL `matchConditions` evaluated last. A request that does not match is not "allowed by the policy" — the policy never sees it at all. Choosing what the engine is allowed to see is a design decision the platform team owns, not leftover configuration. ## Reason one: the rule is wrong outside its intended population Most guardrails are written with a picture of a tenant application in mind. Take a concrete one: *every Deployment with more than one replica must carry the labels a PodDisruptionBudget can select, so voluntary disruptions do not take the whole service down.* That is a reasonable ask for a team's HTTP service. Applied to cluster infrastructure it is meaningless or actively harmful — a node agent runs one Pod per node by design, and a control-plane add-on may be managed by an installer that will never add your label and will fight you if it does. Matching those objects gains nothing and converts a policy violation into a cluster outage, because the workloads you just rejected are the ones everything else depends on. There is a second population most authors forget: **the control plane's own controllers write objects too**. The ReplicaSet controller creates Pods, the garbage collector deletes them, the node lifecycle controller updates them. Those writes go through admission exactly like a developer's `kubectl apply`. A cluster-wide match therefore judges a large volume of machine-issued requests that no human is watching, and a rejection surfaces in a controller's event stream rather than on anyone's terminal. ## Reason two: the self-deadlock The engine evaluating your policies runs as Pods in the cluster. If those Pods are inside the policy's own scope, admitting them requires an engine that is already running. In steady state nobody notices. The failure appears when every replica is gone at once — a full cluster restart, a node pool replacement, an eviction storm. Now the API server tries to create the engine's replacement Pods, needs a decision, and the only thing that could give one is the workload it is trying to create. Whether that manifests as a hard stop or as an unguarded window depends on how the registration is configured to behave when the engine is unreachable, which is a separate decision — but the way to never depend on that decision is to keep the engine out of its own scope in the first place. The same applies to whatever the engine depends on: its certificate machinery, its ingress, its datastore. ## What the exclusion actually looks like Every namespace carries an automatically maintained `kubernetes.io/metadata.name` label, so a namespace can be excluded by name through a label selector: - select namespaces whose `kubernetes.io/metadata.name` is `NotIn` the list `[kube-system, <engine namespace>, <infrastructure namespaces>]`; or - invert the model and match only namespaces that carry an explicit `platform.example.com/tenant: "true"` label, so a namespace is in scope only when someone deliberately puts it there. The two differ in what a new namespace inherits. The exclusion list is opt-out: anything created tomorrow is covered by default, which is what you want for a security guardrail, at the cost of surprising a new infrastructure component. The tenant label is opt-in: nothing is covered until someone labels it, which is safe for the cluster and quietly useless for coverage. Whichever you pick, write down that it was a choice. ## The cost of the exclusion An excluded namespace is an unguarded namespace. Anyone who can create workloads in `kube-system` is outside every rule you have written, so the exclusion is only sound when paired with tight access control on those namespaces — the exclusion moves the control from admission to RBAC rather than removing it. And name the exclusions explicitly. An exclusion keyed on a label that tenants can set on their own namespaces is not an exclusion, it is an open door. One edge worth knowing: a static Pod is started by the kubelet from a file on the node, without asking the API server, so admission never gates the container. The kubelet then creates a mirror Pod object through the API server, and that object does pass through admission — rejecting it does not stop the workload, it only removes the cluster's view of it.
- Doesn't excluding kube-system just hand anyone a place to run whatever they like?Yes, and that is the trade you are making. The exclusion moves the control from admission to access control: the excluded namespaces have to be ones only the platform team and the control plane can write to. It also means the exclusion should name namespaces explicitly rather than keying on a label a tenant could apply to their own namespace, which would let anyone opt out of every rule at once.
- A control-plane component runs as a static Pod. Does your admission policy see it?Not in a way that stops it. The kubelet reads the manifest from disk and starts the container without consulting the API server, so no admission decision is ever requested for the workload itself. The kubelet then creates a mirror Pod object through the API server so the Pod is visible in `kubectl get pods`, and that create does go through admission — but rejecting it only hides the running workload from the cluster's view.
- Which namespaces besides kube-system usually belong in the exclusion?The policy engine's own namespace and anything it depends on, plus wherever cluster infrastructure lives: CNI and CSI drivers, node agents, log and metric collectors, the ingress controller. The test is not "is it a system namespace" but "if this rule rejected a Pod here, would the cluster be able to recover on its own?"
A fire-suppression system that also monitors the closet holding its own pump. Perfect coverage on paper; one fault and you cannot get in to fix the pump.
saying these in an interview costs you the question
- Thinks admission only sees human requests, not controller writes
- Treats exclusions as tidy-up rather than a deadlock guard
- Assumes the policy engine's own Pods are exempt automatically
- Cannot say who or what creates objects in kube-system
- Excludes namespaces by a label tenants can set themselves