A tenant's PDF-invoice renderer parses untrusted uploads in a shared 80-node Kubernetes cluster. What does moving it to a dedicated tainted node pool isolate, and what does it still share?
answer
- blast radius of a breakout
- kernel no longer shared
- Node authorizer bounds Secret reads
- control plane and CRDs still shared
- RBAC cannot police field values
basics
~20 sA dedicated node pool removes kernel, kubelet and node-resource sharing with other tenants, so an escape reaches only that tenant's pods and Secrets. The API server, etcd, cluster-scoped objects, cluster DNS and platform DaemonSets stay shared.
solid answer
~50 sParsing untrusted PDFs is exactly the workload that makes a kernel escape plausible. Moving it to its own tainted pool, with affinity pulling it there, means a breakout lands on nodes that host only that tenant. With the Node authorizer and `NodeRestriction`, a node's credential can read only Secrets referenced by pods bound to that node, so the blast radius stays inside the tenant. It also removes disk, PID and network contention with neighbours. What it does not change: every tenant still uses one API server and etcd, the same CRDs and other cluster-scoped objects, cluster DNS, and the platform DaemonSets that run on every node. Placement must be enforced, not voluntary. RBAC cannot restrict the value of `spec.tolerations`, so you need admission control, such as the `PodTolerationRestriction` and `PodNodeSelector` plugins or a policy engine. Add a sandboxed `RuntimeClass` or `hostUsers: false` for defence in depth.
code
yaml · 10 linesapiVersion: v1
kind: Namespace
metadata:
name: invoices
labels:
pod-security.kubernetes.io/enforce: restricted
annotations:
scheduler.alpha.kubernetes.io/node-selector: "tenant=invoices"
scheduler.alpha.kubernetes.io/defaultTolerations: '[{"key":"tenant","operator":"Equal","value":"invoices","effect":"NoSchedule"}]'
scheduler.alpha.kubernetes.io/tolerationsWhitelist: '[{"key":"tenant","operator":"Equal","value":"invoices","effect":"NoSchedule"}]'go deeper
Know that pods can be kept on their own nodes so a problem on one tenant's node does not touch others.
Explain that separate nodes mean separate kernels and kubelets, while the API server and cluster-wide objects stay common.
Show the Node authorizer blast-radius argument, enforce placement with admission rather than trust, and list what still leaks: control plane, CRDs, DNS, DaemonSets.
Weigh fragmented capacity and extra node groups against a separate cluster, and decide which untrusted workloads justify each step.
## The scenario The `invoices` tenant runs a **PDF-invoice renderer** that opens files uploaded by the public. Document parsers are a classic source of memory-corruption bugs, and a container is only kernel-level isolation. If the renderer is exploited and the attacker breaks out of the container, they are on the node. In a shared 80-node cluster that autoscales between 20 and 80 nodes, that node also runs other tenants' pods. This answer covers what a **dedicated, tainted node pool** buys, and what it cannot. ## What a dedicated pool isolates A taint on the pool keeps other tenants' pods off it, and node affinity keeps the renderer from landing elsewhere. Together they give: - **No shared kernel with other tenants.** An escape reaches only `invoices` workloads. - **Contained node credentials.** With the **Node authorizer** and the **`NodeRestriction`** admission plugin, which kubeadm enables, a kubelet's credential can `get` only the Secrets, ConfigMaps and volumes referenced by pods bound to that node, and can modify only its own Node object and its own pods. A compromised node leaks this tenant's Secrets, not everyone's. - **No noisy neighbours** on disk I/O, PIDs, conntrack entries or local ephemeral storage. - **Separate capacity accounting.** The pool can be sized and billed per tenant. ## What it still shares | Shared surface | Consequence | |---|---| | kube-apiserver and etcd | An API abuse or request storm still affects all tenants. API Priority and Fairness queues per user but shares each priority level | | Cluster-scoped objects | One set of CRDs, StorageClasses, webhooks and ClusterRoles for everyone | | Cluster DNS | Names in every namespace still resolve from the renderer | | Platform DaemonSets | Log shippers and CNI agents run on the dedicated nodes too, often privileged, and are part of the trust boundary | | Pod network | Still one network, so the default-deny NetworkPolicy is still required | This is why dedicated nodes are often described as **strong soft tenancy**. Only a separate control plane removes the first three rows. ## Making placement enforceable A taint only keeps pods *off* a node until someone writes a matching toleration. RBAC authorizes **verbs on resources**, not the **values of fields**, so any tenant allowed to create pods can add any toleration. Enforce placement at admission: 1. **`PodTolerationRestriction`**, an API server admission plugin that is off by default and enabled with `--enable-admission-plugins`. It reads the namespace annotations `scheduler.alpha.kubernetes.io/defaultTolerations` and `scheduler.alpha.kubernetes.io/tolerationsWhitelist`, adds the defaults and rejects pods whose tolerations fall outside the allowed list. 2. **`PodNodeSelector`**, also off by default. It merges the namespace annotation `scheduler.alpha.kubernetes.io/node-selector` into each pod and rejects a pod whose own node selector conflicts with it. 3. **A policy engine** can express the same rules when you would rather not enable those plugins. ## Defence in depth on the node - A **`RuntimeClass`** (`node.k8s.io/v1`) whose `handler` points at a sandboxed runtime puts a stronger boundary than namespaces and cgroups around the renderer. Its `scheduling.nodeSelector` keeps those pods on nodes that have the handler. - **Pod user namespaces** via `hostUsers: false` map container root to an unprivileged host UID. This needs runtime and kernel support. - The **`restricted` Pod Security Standard** on the namespace. ## Costs to weigh - **Fragmented capacity.** Spare room in the dedicated pool cannot absorb other tenants' load. If Cluster Autoscaler scales the pool, set a sensible minimum, or the renderer waits for a cold node. - **More node groups to patch and upgrade.** - **A false sense of completeness.** If the requirement is "no shared control plane", this is the wrong tool, and the answer is a separate cluster.
- Why is a taint on the dedicated Kubernetes node pool not enough to keep other tenants off it?A taint repels only pods that do not tolerate it, and any user who can create pods can write any toleration, because RBAC does not inspect field values. Unless admission control restricts which tolerations each namespace may carry, another tenant can simply opt in. PodTolerationRestriction or a policy engine closes that gap.
- The platform's log-shipping DaemonSet runs privileged on every node, including the dedicated pool. Why does that matter for this isolation?The DaemonSet shares the renderer's kernel and usually mounts host paths. An attacker who escapes onto the node can often read or tamper with it, and through its ServiceAccount reach whatever the DaemonSet may reach. Platform agents on tenant pools must be least-privileged, and their credentials must not grant cross-tenant access.
saying these in an interview costs you the question
- Dedicated nodes give a tenant its own control plane.
- A NoSchedule taint alone stops other tenants landing on the pool.
- RBAC can forbid a specific toleration value on pod create.
- A compromised kubelet can read every Secret in the cluster despite NodeRestriction.
- Once nodes are dedicated, NetworkPolicy is no longer needed.