You need a set of Kubernetes nodes reserved exclusively for one team's workload, so that nothing else lands there and that team's pods always run there. Describe exactly how you would configure it and why a toleration alone is not enough.
answer
- Two halves: taint keeps out, label+selector pulls in
- Same key on both sides is convention, not magic
- Taint at node-pool creation, not by hand after join
- DaemonSets need the custom toleration too
- Enforce with admission policy or PodTolerationRestriction
basics
~20 sTaint the nodes (for example dedicated=team-a:NoSchedule) so untolerating pods are kept out, label them (pool=team-a), then give the team's pods both a matching toleration and a nodeSelector or node affinity for that label. The taint provides exclusivity; the selector provides attraction.
solid answer
~50 sNode dedication needs **two independent halves**: 1. **Keep others out** — taint the nodes: `kubectl taint nodes n1 dedicated=team-a:NoSchedule`. Every pod without a matching toleration is filtered off those nodes. 2. **Pull yours in** — label the nodes (`dedicated=team-a` or `pool=team-a`) and put a `nodeSelector` / required node affinity on the team's pod template, plus the toleration. A toleration alone only removes the node from the exclusion list; it does not bias the scheduler towards it, so the pods would still spread across the ordinary nodes. Practical points: set the taint and label on the **node pool / node group definition**, not by hand after nodes join, otherwise every replacement node has a window where arbitrary pods can land. Remember DaemonSets and system agents need tolerations to keep working there. If you want the isolation enforced rather than merely conventional, add an admission policy (Gatekeeper/Kyverno, or the `PodTolerationRestriction` plugin) so other namespaces cannot simply add the toleration themselves.
code
bash · 3 lineskubectl label nodes n1 n2 n3 pool=team-a
kubectl taint nodes n1 n2 n3 dedicated=team-a:NoSchedule
kubectl get pods -A -o wide --field-selector spec.nodeName=n1go deeper
State the two halves — taint the node, and give the pod both the toleration and a nodeSelector — and show the two kubectl commands.
Explain why the toleration alone is insufficient, pick NoSchedule over PreferNoSchedule with a reason, and remember the DaemonSet agents.
Add the operational detail: taints/labels declared on the node pool so replacements are born correct, admission enforcement so the boundary is not merely conventional, and how you verify with field-selector queries.
Weigh dedication against fungible capacity — per-pool headroom and failure spares, stranded capacity, when soft separation is enough — and define who owns taint keys and how the policy is enforced cluster-wide.
## The two-sided nature of placement Kubernetes splits placement control into repulsion and attraction, and dedication needs both: - **Repulsion lives on the node**: taints. Only a taint can prevent pods that know nothing about your pool from landing on it. Node affinity cannot do this, because affinity is opt-in on the pod side and pods that don't declare it are unconstrained. - **Attraction lives on the pod**: `nodeSelector`, node affinity. Only these can make your pods actually prefer or require the pool. Miss the first and other teams' workloads bin-pack onto your expensive GPU nodes. Miss the second and your workload scatters across the general pool while the dedicated nodes sit empty — the classic "I added the toleration and nothing changed" bug report. ## The concrete recipe 1. **Label the nodes** so pods can select them: `kubectl label nodes n1 pool=team-a`. 2. **Taint the nodes** so nothing else can enter: `kubectl taint nodes n1 dedicated=team-a:NoSchedule`. 3. **In the workload's pod template** add both a toleration matching the taint and a `nodeSelector: {pool: team-a}` (or a `requiredDuringSchedulingIgnoredDuringExecution` node affinity if you need set-based matching such as `pool In [team-a, team-a-spare]`). By convention people reuse the same key on both sides (`dedicated=team-a` as label *and* as taint) — the API allows it and it keeps manifests readable, but they are two separate fields with two separate meanings. ## Choosing the effect Use `NoSchedule` for steady-state dedication. `NoExecute` additionally evicts anything already running that doesn't tolerate the taint — the right choice when you are converting an existing shared pool into a dedicated one and want the squatters gone immediately, but understand you are triggering an involuntary disruption that does not respect PodDisruptionBudgets. `PreferNoSchedule` is not dedication at all; it is a scheduling hint and other pods will still land there under pressure. ## Where to apply the taint On managed platforms, set the taint and label in the **node pool / managed node group / machine template** specification so nodes are born with them. If you taint by hand after a node registers, there is a race: between the node becoming Ready and your taint landing, the scheduler can bind arbitrary pending pods to it. That window reappears every time the pool scales out or replaces a node, so hand-tainting is not a stable design — it only looks like it works. ## Do not break the node's own agents A tainted node still needs its log shipper, CNI agent, node exporter and CSI node plugin. DaemonSet pods are given tolerations for the standard condition taints automatically, but **not** for your custom `dedicated` taint — you must add it to those DaemonSets, or use a blanket `operator: Exists` toleration in the agents you control. Forgetting this is how a dedicated pool ends up with no metrics and no logs. ## Making the boundary real As specified so far, dedication is a *convention*: any team that copies the toleration into their manifest gets access to the pool. If the pool exists for cost allocation or compliance, enforce it: - A validating admission policy (Kyverno, Gatekeeper, or a `ValidatingAdmissionPolicy`) that rejects pods carrying the `dedicated=team-a` toleration outside team A's namespaces. - The `PodTolerationRestriction` admission plugin, which can whitelist which tolerations a namespace may use and inject default ones. - ResourceQuotas in the owning namespace so the team cannot exceed the pool it paid for. ## Costs and tradeoffs to raise Exclusive pools fragment cluster capacity. Each pool needs its own headroom for spikes and its own spare node for failure, so N dedicated pools cost meaningfully more than one shared pool of the same total size, and a burst in one pool cannot borrow idle capacity from another. Dedication is justified by hardware (GPUs, local NVMe, high-memory shapes), licensing tied to cores, noisy-neighbour isolation for latency-critical services, or compliance boundaries — not by org-chart tidiness. When you only need a *tendency*, `PreferNoSchedule` plus preferred node affinity gives soft separation while keeping the capacity fungible. ## Verifying After rollout, check `kubectl get pods -o wide` that the workload actually landed on the pool, and check that nothing foreign is there: `kubectl get pods --all-namespaces --field-selector spec.nodeName=n1`. A pending pod's `describe` output will tell you plainly whether the taint or the selector is the blocking predicate.
- After you taint the pool, the team's metrics and logs disappear from those nodes. Why?The node-level agents run as DaemonSets, and DaemonSet pods only get automatic tolerations for the built-in condition taints such as not-ready, unreachable and the disk/memory pressure ones — not for your custom `dedicated` taint. Their pods are therefore filtered off the pool. The fix is to add a matching toleration (often a broad `operator: Exists`) to the agent DaemonSets that legitimately must run everywhere.
- Another team copies the toleration into their Deployment and starts using your dedicated nodes. How do you prevent that?Taints are an access-control mechanism only by convention, since any pod author can add the toleration. To make it real you gate it at admission: a Kyverno or Gatekeeper policy, or a ValidatingAdmissionPolicy, that rejects pods declaring the `dedicated=team-a` toleration unless they are in an allowed namespace. The older `PodTolerationRestriction` admission plugin does the same by whitelisting per-namespace tolerations.
A taint is the locked gate around a car park; the nodeSelector is the sign telling your drivers which car park to use. Handing out keys does not tell anyone to park there.
saying these in an interview costs you the question
- Adding only the toleration and expecting pods to move to the pool
- Using PreferNoSchedule and calling it dedicated capacity
- Tainting nodes manually after they join, leaving a race window on every scale-out or node replacement
- Forgetting DaemonSet agents need the custom toleration, so the pool loses logging/metrics/CSI
- Treating a taint as security isolation when any namespace can add the matching toleration