In a 38-node Kubernetes cluster, CPU-only pods keep landing on the six GPU nodes while a video-transcoding pod waits Pending for a GPU. How would you fence the GPU pool, and what does each mechanism guarantee?
answer
- keep out versus pull in
- taint key equals resource name
- an admission plugin writes the toleration
- NoSchedule leaves running pods alone
- tolerations are permission, not a lock
basics
~10 sTaint the GPU nodes (nvidia.com/gpu:NoSchedule) to keep CPU-only pods off, and label them so GPU pods can select them. Enabling the ExtendedResourceToleration admission plugin adds the matching toleration to any pod that requests nvidia.com/gpu.
solid answer
~40 sThere are two problems, handled by two mechanisms. First, **keep others out**: taint every GPU node, for example `nvidia.com/gpu=present:NoSchedule`, so only pods with a matching toleration can land there. Second, **pull GPU pods in**: a pod requesting `nvidia.com/gpu` can only fit on nodes that advertise it, and a node label plus `nodeSelector` or node affinity picks the right pool or card model. If you enable the `ExtendedResourceToleration` admission plugin (it is off by default), the API server adds a toleration with key `nvidia.com/gpu`, operator `Exists`, effect `NoSchedule` to every pod that requests that resource. Then teams don't hand-write tolerations. Taints are not a security boundary, though: anyone can add a toleration. So pair them with admission policy, and remember that DaemonSets and existing pods need their own handling. `NoSchedule` never evicts pods already running.
code
bash · 2 lineskubectl label nodes gpu-node-01 gpu-node-02 gpu-node-03 gpu-node-04 gpu-node-05 gpu-node-06 node-pool=gpu-transcode
kubectl taint nodes -l node-pool=gpu-transcode nvidia.com/gpu=present:NoSchedulego deeper
Recall that taints push pods away and labels with nodeSelector pull pods toward nodes; a GPU pool usually needs both.
Explain what ExtendedResourceToleration injects, why the taint key must equal the resource name, and why NoSchedule leaves running pods in place.
Diagnose stranded GPUs from CPU-only pods, plan the drain and DaemonSet tolerations, and add admission policy because a toleration is only a permission.
Decide how accelerator pools, tenancy and admission policy combine across teams, and what the stranding costs against the complexity of more pools.
## The symptom The cluster has 38 nodes in a mixed spot and on-demand pool. Six of them carry 4 GPUs each, 24 in total. The video-transcoding workers each request `nvidia.com/gpu: 1` and 350m of CPU. Ordinary CPU-only Deployments are also being scheduled onto the GPU nodes, because to kube-scheduler those nodes are just large machines with free CPU and memory. When the CPU-only pods fill a GPU node's CPU or memory, that node's GPUs become **stranded**: devices are free, but a new transcoding pod no longer fits on CPU or memory. The transcoder stays `Pending` while GPUs sit idle. ## Two directions, two mechanisms | Goal | Mechanism | What it guarantees | What it does not | |---|---|---|---| | Keep non-GPU pods off GPU nodes | Taint with effect `NoSchedule` | Pods without a matching toleration are never scheduled there | Does not evict pods already running; does not stop someone from adding a toleration | | Put GPU pods only on GPU nodes | The `nvidia.com/gpu` request itself | A pod is placed only where free GPUs are advertised | Does not choose between GPU models or pools | | Choose the right GPU pool | Node label plus `nodeSelector` or required node affinity | The pod lands only on nodes with that label | Does not keep other pods off those nodes | | Spare teams from writing tolerations | `ExtendedResourceToleration` admission plugin | Pods requesting the resource get the toleration automatically | Is off by default; matches only a taint whose key is the resource name | ## How to set it up 1. **Label the pool** with your own key, for example `node-pool=gpu-transcode`. Do this at node-provisioning time, so replacement nodes, including spot replacements, carry it. 2. **Taint the pool** with a key equal to the extended resource name: - `kubectl taint nodes gpu-node-01 nvidia.com/gpu=present:NoSchedule` - in production, set it in the node group's template or the kubelet's `--register-with-taints` flag, so it exists before any pod can land. 3. **Enable `ExtendedResourceToleration`** on kube-apiserver with `--enable-admission-plugins`. For every pod whose containers or init containers request an extended resource, it adds a toleration with `key: <resource name>`, `operator: Exists` and `effect: NoSchedule`. The taint key must therefore be exactly `nvidia.com/gpu`. A taint keyed `dedicated=gpu` would not be matched. 4. **Select the pool** in the transcoding Deployment with `nodeSelector: {node-pool: gpu-transcode}` when there are several GPU types. 5. **Handle what the taint does not.** CPU-only pods already running on GPU nodes stay there. Drain the nodes, or use a `NoExecute` taint deliberately. DaemonSets that must run on GPU nodes, such as the device plugin, log shippers and monitoring agents, need explicit tolerations. ## Why a toleration alone is not a fence A toleration only *permits* scheduling onto a tainted node. Any team can copy one into its manifest and put CPU-only pods on the GPU pool again. If exclusivity matters: - use an admission policy that rejects the GPU toleration on pods that request no GPU; - or scope GPU access by namespace with the `PodTolerationRestriction` admission plugin or a policy engine; - and put quotas on the extended resource, which is covered under ResourceQuota. ## Spot interplay If some GPU nodes are spot capacity, they already carry a capacity-type label and possibly an interruption taint. A transcoding pod that should avoid spot needs node affinity on that label as well. Fencing interruptible capacity is its own topic. The point here is that the GPU taint and the spot taint are separate keys, and a pod must tolerate each one it is allowed to run on. ## Verifying - `kubectl get nodes -l node-pool=gpu-transcode -o custom-columns=NAME:.metadata.name,TAINTS:.spec.taints` confirms every node is tainted. - `kubectl get pod <pod> -o jsonpath='{.spec.tolerations}'` on a new transcoding pod shows the injected toleration. - Listing the pods on each GPU node should show only GPU consumers and required DaemonSets.
- You enabled `ExtendedResourceToleration`, but transcoding pods still can't land on the GPU nodes. What do you check first?The taint key. The plugin adds a toleration whose key is the requested resource name (`nvidia.com/gpu`), with operator `Exists` and effect `NoSchedule`. If the nodes are tainted `dedicated=gpu:NoSchedule`, or use `NoExecute`, the injected toleration does not match. Then check that the plugin is really in kube-apiserver's `--enable-admission-plugins`.
- Why not rely on the GPU request alone and skip the taint?The request only pulls GPU pods toward GPU nodes. It does nothing to push CPU-only pods away. Without a taint, the scheduler happily fills GPU nodes' CPU and memory with ordinary pods, stranding their GPUs, so transcoding pods go Pending while devices sit free.
saying these in an interview costs you the question
- A nodeSelector on GPU pods keeps other pods off the GPU nodes
- Adding a NoSchedule taint evicts CPU-only pods already running there
- ExtendedResourceToleration is enabled on every cluster by default
- Any taint on a GPU node is matched by the auto-added toleration
- Taints stop other teams from ever scheduling onto the pool