skip to content

Your Kubernetes cluster has grown to a dozen teams and several hardware shapes (GPU, high-memory, spot). How would you decide which node groups get taints, who is allowed to tolerate them, and what does that policy cost you?

level: principalimportance: nice to knowfreq 30%

answer

  1. Partitioning trades multiplexing for isolation
  2. Taint for hardware, failure semantics, compliance — not org chart
  3. Keys are platform-owned, set on the node group not by hand
  4. Tolerations are self-service → enforce at admission
  5. Every partition: own headroom, own failure spare, own DaemonSet tolerations

basics

~20 s

Taint only where sharing genuinely fails: scarce or expensive hardware, licensing, compliance boundaries, and disruption-prone nodes like spot. Own the taint keys centrally, enforce who may tolerate them at admission, and accept the cost — every exclusive pool needs its own spare capacity and headroom, so partitioned clusters are more expensive than shared ones.

solid answer

~60 s

My default is a **shared pool**; a taint has to earn its existence. Criteria for tainting a node group: - **Scarce/expensive hardware** — GPUs, local NVMe, high-memory shapes. Without a taint, ordinary pods bin-pack onto them and strand capacity you paid a premium for. - **Different failure semantics** — spot/preemptible nodes should be opt-in, so nothing lands there that cannot survive a two-minute termination notice. - **Compliance or tenancy boundaries** where node-level isolation is genuinely required. - **Noisy-neighbour isolation** for a small number of latency-critical services. Org-chart tidiness is *not* a criterion. Governance: taint keys are platform-owned and namespaced (`workload=gpu`, `lifecycle=spot`); teams do not invent them. Because any pod can add a toleration, the boundary is only real if admission enforces it — a Kyverno/Gatekeeper policy or `PodTolerationRestriction` mapping namespaces to permitted tolerations, plus quotas. Cost: each pool carries independent headroom and a failure spare, and idle capacity in one cannot serve a spike in another. Where a *tendency* suffices, use PreferNoSchedule plus preferred affinity and keep the capacity fungible.

go deeper

for a junior

Recognise the basic principle — taints reserve nodes, and reserving costs capacity — and that GPU pools are the standard example.

for a middle

Give concrete criteria (hardware, spot, compliance), note that taint plus label plus selector is the full recipe, and mention DaemonSets needing the new toleration.

for a senior

Argue the cost side quantitatively — headroom per pool, fragmentation, failure spares — and describe enforcement through admission policy plus quotas.

for a principal

Lead with the multiplexing tradeoff, define the taint-key vocabulary and ownership model, cover autoscaler and DaemonSet second-order effects, and set a review process so partitions do not accumulate.

## Framing: partitioning is a capacity decision, not a config decision Every taint you add carves the cluster into a smaller pool. Statistical multiplexing is the main economic argument for running a shared cluster at all: many uncorrelated workloads on one pool need less total headroom than the same workloads in separate pools, because their peaks do not coincide. Each partition you create gives some of that back. The question "should this node group be tainted?" is therefore really "is the isolation worth the stranded capacity?" ## When a taint is clearly justified **Scarce or premium hardware.** A GPU node is many times the price of a general node. Without `nvidia.com/gpu=true:NoSchedule` (or equivalent), the scheduler will happily place a stateless web pod there because it only sees CPU and memory fitting. That pod contributes nothing and can block a GPU job from fitting. Same argument for local-NVMe and high-memory shapes. **Different failure semantics.** Spot/preemptible nodes disappear on short notice. Tainting them `lifecycle=spot:NoSchedule` makes running there an explicit, informed choice — batch and stateless replicas opt in; the database and the single-replica control services do not. Without the taint, whether your critical singleton is on interruptible hardware is a lottery. **Licensing and compliance.** Per-core licensed software, or workloads whose data may only touch certain hardware or regions, need a boundary that cannot be crossed by omission. **Documented noisy-neighbour problems.** A latency-critical service that measurably suffers from co-tenancy. Demand evidence, because this reason is claimed far more often than it is true, and requests/limits plus CPU management policies often solve it more cheaply. ## When to say no Requests for a dedicated pool "because it's our team's budget" or "so we can reason about capacity" are usually better served by ResourceQuota and chargeback based on requests. Similarly, a preference rather than a requirement — "we'd rather not share with batch" — is served by `PreferNoSchedule` plus preferred node/pod anti-affinity, which keeps the nodes available as overflow when the cluster is tight. Soft separation degrades gracefully; hard partitioning fails as Pending pods next to idle nodes. ## Governance of the keys Taint keys are a cluster-wide namespace and need an owner. Practical rules: - The platform team owns the vocabulary and documents it: `workload=gpu`, `lifecycle=spot`, `tenancy=pci`. Keys are set on the node group / machine template so every node — including scale-out and replacements — is born tainted; hand-tainting leaves a race window each time the group grows. - Never invent keys under the reserved `node.kubernetes.io/` or `node.cloudprovider.kubernetes.io/` prefixes. - Pair each taint with a node label so workloads can also be *attracted*; the taint alone never places anything. ## Enforcement, because tolerations are self-service A toleration is three lines of YAML any team can copy. Left unenforced, dedication decays: within a quarter, several workloads tolerate the GPU taint "just in case". Make the boundary real with admission control — a Kyverno/Gatekeeper policy or `ValidatingAdmissionPolicy` that rejects pods declaring a restricted toleration outside the namespaces entitled to it, or the `PodTolerationRestriction` plugin which whitelists and defaults tolerations per namespace. Combine with quotas on the extended resource (`nvidia.com/gpu`) so entitlement is bounded, not just binary. Audit periodically: list pods carrying restricted tolerations and reconcile against the entitlement list. ## Second-order effects to think about - **Autoscaling.** Each tainted group scales independently; scaling logic must know the taints or it will grow the wrong group while pods stay Pending. Node-group configuration must advertise its taints and labels for simulated scheduling to work. - **DaemonSets.** Every new taint is a new toleration your logging, metrics, CNI and CSI DaemonSets need. Adding a taint without updating them silently blinds the pool. - **Bin-packing and fragmentation.** Small exclusive pools round badly: a pool needing 2.3 nodes' worth of capacity buys 3, and the 0.7 is stranded. - **Failure domains.** Every pool needs enough nodes to survive losing one, and to keep spreading replicas across zones. A three-node exclusive pool across three zones has no slack at all. - **Cognitive load.** Each taint is something every future engineer must learn before they can explain a Pending pod. Fewer, well-named partitions beat many ad-hoc ones. ## The recommendation to state Start with one shared pool plus a small number of hardware-driven tainted groups. Add a partition only with a written justification, a named owner, an admission policy, a DaemonSet update, and a review date. Prefer soft mechanisms whenever the requirement is a preference. Revisit the set of partitions periodically — pools created for a project that ended are pure stranded cost.

  • A team asks for a dedicated tainted node pool because their service is latency-sensitive. What do you ask before agreeing?
    I ask for evidence that co-tenancy is the cause: latency percentiles correlated with neighbours' activity, CPU throttling metrics, and whether requests and limits are set correctly in the first place. Guaranteed QoS, correct CPU requests, and in extreme cases the static CPU manager policy solve most real noisy-neighbour cases without partitioning. If the evidence holds, I would still ask whether PreferNoSchedule plus preferred anti-affinity is enough, so the nodes stay usable as overflow capacity.
  • How do you keep dedicated pools from being quietly colonised over time?
    Enforcement plus audit. Admission policy — Kyverno, Gatekeeper or a ValidatingAdmissionPolicy — rejects pods carrying a restricted toleration from namespaces that are not entitled to it, and quotas on the relevant extended resource bound how much an entitled team can take. On top of that I run a periodic report listing all pods with restricted tolerations and reconcile it against the entitlement list, because policies get exceptions and exceptions get forgotten.

Reserved parking spaces: a few for the delivery bay and the accessible bays are clearly worth it, but reserve a space per department and the car park is full of empty named spots while visitors circle.

saying these in an interview costs you the question

  • Treating taints as a security boundary when any namespace can add a toleration
  • Creating a pool per team by default rather than per genuine hardware or failure-semantics distinction
  • Ignoring that each partition needs its own headroom and failure spare, so partitioning costs money
  • Adding a new taint without updating the node-agent DaemonSets or the node-group configuration used for scaling
  • Using PreferNoSchedule where a hard guarantee is required, or a hard taint where a preference would keep capacity fungible

context