skip to content

A Kubernetes pod stays Pending with the event '0/40 nodes are available: 12 Insufficient cpu, 28 node(s) had untolerated taint'. How do you read that message and drive the problem to resolution?

level: seniorimportance: must knowfreq 58%

answer

  1. 0/N = feasible nodes out of nodes examined; clauses = per-plugin counts
  2. Insufficient cpu = requests vs allocatable, never live utilisation
  3. Read 'Allocated resources' in kubectl describe node
  4. Condition taints = sick nodes; deliberate taints = reserved pool
  5. Toleration permits, nodeSelector attracts — usually need both

basics

~20 s

That line is the filter phase's aggregated verdict: each clause is one plugin's rejection count, and the counts sum to every node, so no node is feasible. Fix the dominant clause — right-size requests or add capacity for the fit failures, add a toleration or untaint for the rest — then let the pod be retried.

solid answer

~60 s

The message is kube-scheduler's **filter-phase summary**: `0/40 nodes are available` means zero feasible nodes out of forty examined, and each clause is the count of nodes rejected by one filter plugin. Here `NodeResourcesFit` rejected 12 and `TaintToleration` rejected 28 — together all 40, so every node failed for one reason or the other. Read it as **two independent problems**, not one. Fix in this order: 1. **Confirm what "Insufficient cpu" means** — fit is computed from **requests** versus allocatable minus existing requests, never live utilisation. So check `kubectl describe node` "Allocated resources", not a CPU graph. Either the pod's request is unrealistic, or the nodes are genuinely reserved out. 2. **Identify the taint** — `kubectl get nodes -o json | jq '.items[].spec.taints'`. It is usually a deliberate dedicated pool (GPU, spot) or a condition taint like `node.kubernetes.io/disk-pressure`. Add a toleration only if the pod belongs there. Also check whether a preemption line follows (`No preemption victims found`), whether the Cluster Autoscaler is refusing to scale (its events say why), and whether the pod is blocked by an unbound PVC, which shows up as a `VolumeBinding` clause instead.

code

bash · 6 lines
bash
kubectl describe pod checkout-6f8 | sed -n '/Events/,$p'
kubectl get events --field-selector involvedObject.name=checkout-6f8 --sort-by=.lastTimestamp
kubectl get pod checkout-6f8 -o jsonpath='{.spec.containers[*].resources.requests}{"\n"}'
kubectl describe node worker-03 | sed -n '/Allocated resources/,/Events/p'
kubectl get nodes -o custom-columns=NAME:.metadata.name,TAINTS:.spec.taints,CPU:.status.allocatable.cpu
kubectl -n kube-system get configmap cluster-autoscaler-status -o yaml

go deeper

for a junior

Parse the message: no node was available, some failed on resources and some on taints, and know the two obvious levers are smaller requests or a matching toleration.

for a middle

Explain that fit is computed from requests, show where to read allocated resources on a node, and distinguish deliberate pool taints from node condition taints.

for a senior

Run the full triage: event history, per-node allocation, largest-node feasibility, autoscaler refusal reasons, the preemption clause, and non-scheduler causes like scheduling gates.

for a principal

Turn the incident into policy — request right-sizing with VPA recommendations, node pools whose labels and taints encode intent, an SLO alert on pending duration, and a decision on autoscaling headroom versus cost.

## Anatomy of the message When the filter phase leaves zero feasible nodes, kube-scheduler emits a `FailedScheduling` event whose text is an *aggregation*: one clause per distinct rejection reason, with the number of nodes each reason accounted for. ``` 0/40 nodes are available: 12 Insufficient cpu, 28 node(s) had untolerated taint {workload=gpu: true}. preemption: 0/40 nodes are available: 40 No preemption victims found for incoming pod. ``` Three things to extract every time: 1. **The denominator** — 40 is the number of nodes considered. In big clusters `percentageOfNodesToScore` may stop the sweep early, so a denominator smaller than your fleet is normal, not a bug. 2. **Whether the clauses sum to the denominator.** If they do, you have the complete picture. If they don't, some nodes failed for reasons not shown (the message is truncated at a few reasons). 3. **The preemption verdict after the period.** `No preemption victims found` means even removing lower-priority pods would not help — typically because the blocking reason is a constraint (taint, affinity, volume) rather than capacity. Common clauses and what they actually mean: | Clause | Plugin | Real meaning | |---|---|---| | `Insufficient cpu` / `memory` / `<extended-resource>` | NodeResourcesFit | Sum of **requests** on the node plus this pod exceeds allocatable | | `node(s) had untolerated taint {k=v:Effect}` | TaintToleration | Deliberate pool, or a condition taint the node applied itself | | `didn't match Pod's node affinity/selector` | NodeAffinity | Label missing or mistyped | | `node(s) were unschedulable` | NodeUnschedulable | Cordoned | | `node(s) didn't have free ports for the requested pod ports` | NodePorts | hostPort collision | | `pod has unbound immediate PersistentVolumeClaims` | VolumeBinding | PVC not bound / no matching PV or zone mismatch | | `node(s) exceed max volume count` | volume limits | Per-node attachment cap (cloud disk limits) | | `node(s) didn't match pod topology spread constraints` | PodTopologySpread | Skew would exceed maxSkew with DoNotSchedule | | `node(s) didn't satisfy existing pods anti-affinity rules` | InterPodAffinity | An already-placed pod excludes this one | ## A working procedure **Step 1 — get the full text, not the summary.** `kubectl describe pod <p>` shows the latest event; `kubectl get events --field-selector involvedObject.name=<p> --sort-by=.lastTimestamp` shows the history, which matters because the reasons *change* as the cluster moves. **Step 2 — treat requests as the currency.** For `Insufficient cpu`, run `kubectl describe node <n>` and read the "Allocated resources" table: it lists requests and limits as a percentage of allocatable. A node at 4% CPU utilisation can be 100% requested. Then decide which is wrong: - The pod's request is inflated (a copy-pasted `cpu: "4"` for a service that uses 200m) → right-size it. - Neighbours' requests are inflated → a cluster-wide right-sizing problem; VPA in recommendation mode gives you the data. - The cluster is genuinely full → add capacity. Also check the pod is not asking for more than *any single node's* allocatable, which is unfixable by adding more of the same node type. Compare against `kubectl get nodes -o custom-columns=NAME:.metadata.name,CPU:.status.allocatable.cpu,MEM:.status.allocatable.memory`. **Step 3 — classify the taint.** Condition taints (`node.kubernetes.io/not-ready`, `unreachable`, `disk-pressure`, `memory-pressure`, `pid-pressure`, `network-unavailable`) mean the *nodes* are unhealthy — the pod is a symptom, and the fix is on the node, not in the manifest. Deliberate taints (`workload=gpu:NoSchedule`, `spot=true:NoSchedule`) mean someone reserved that pool; adding a toleration is only correct if this workload is supposed to be there, and it usually needs a matching `nodeSelector` too, since a toleration merely *permits* the node, it does not *attract* the pod. **Step 4 — check the autoscaler.** If a Cluster Autoscaler is present, an unschedulable pod should trigger scale-up. Its events (on the pod, and in the `cluster-autoscaler-status` ConfigMap) explain refusals: no node group can satisfy the pod's constraints, the group is at max size, or the pod cannot fit any node shape in the group. **Step 5 — mind the preemption clause.** If the pod carries a high priority and you expected it to evict others, `No preemption victims found` is the tell that the blocker is a constraint or that all pods on candidate nodes are of equal-or-higher priority. ## Things that look like scheduling failures but aren't - **Pending with no events at all** — often `spec.nodeName` set by hand to a nonexistent node, a `schedulerName` pointing at a scheduler that is not running, or a pod stuck behind **scheduling gates** (`spec.schedulingGates`), which hold a pod before it even enters the queue. - **Pending because of quota** — that fails at admission on the controller, not the scheduler; you will see it on the ReplicaSet, not the pod. - **ContainerCreating, not Pending** — the pod *was* scheduled; the problem is image pull, volume mount or CNI. ## Closing the loop Resolution is rarely "add a toleration": most of the time the durable fix is right-sized requests plus a node pool whose taints and labels express the intent. Leave the cluster with the pod's constraints matching a pool that exists, and add an alert on pods pending beyond your SLO so the next occurrence is noticed before a human reports it.

  • The pod requests 8 CPU and every node has 8 allocatable CPU, yet nothing schedules. Why, and what do you do?
    Allocatable is already reduced by system and kube reserved amounts, and every node also runs DaemonSet pods with their own requests, so the free remainder is always below the headline number. A pod requesting the full allocatable of a node effectively never fits. The fix is either a smaller request or a larger node shape; adding more nodes of the same size will not help.
  • How do you tell whether the Cluster Autoscaler will solve a Pending pod or not?
    Look at the autoscaler's events on the pod and the cluster-autoscaler-status ConfigMap. It simulates the pod against each node group's template, so it refuses when no group's shape or labels and taints can satisfy the pod's constraints, or when the group is already at max size. A pod blocked by a nodeSelector for a pool that has no node group will stay Pending forever regardless of capacity.
  • The event lists a taint like node.kubernetes.io/disk-pressure. Should you add a toleration?
    Almost never. That is a condition taint the node applied to itself because kubelet detected a resource problem, so the pod is reporting a node health issue rather than a placement mistake. The fix is on the node — reclaim disk, fix a runaway log or image cache, or replace the node. Tolerating it just schedules workloads onto a node that is about to evict them.

saying these in an interview costs you the question

  • Reading 'Insufficient cpu' as live CPU utilisation and dismissing it because dashboards look idle
  • Adding a broad toleration (or operator: Exists) to make the message go away, including for condition taints
  • Assuming a toleration alone will attract the pod to the pool, without a matching nodeSelector or affinity
  • Ignoring the preemption clause and expecting a high-priority pod to always find room
  • Confusing Pending with ContainerCreating, and debugging scheduling when the issue is image pull or volume mount

context