A pod has been stuck in the `Pending` phase for ten minutes. Walk through how you find out why, and explain what a `FailedScheduling` event message such as "0/6 nodes are available: 4 Insufficient cpu, 2 node(s) had untolerated taint" is telling you.
answer
- Pending = no node yet, or bound and still pulling
- describe pod → Events → FailedScheduling
- message is a tally: counts sum to nodes considered
- empty Node: field = scheduler; populated = kubelet
- no pod at all → check ReplicaSet FailedCreate
basics
~20 sPending means the pod is accepted but not bound to a node (or images are still pulling). Run kubectl describe pod and read the Events section: the scheduler's FailedScheduling message lists, per rejection reason, how many nodes it disqualified — that count tells you which constraint to fix.
solid answer
~50 s`Pending` means the pod exists in the API but has no node yet — the scheduler could not find one, or it is bound and still pulling images. First command: `kubectl describe pod <name>`. Check: - **Node:** field — empty means unscheduled; populated means it is a kubelet/image problem, not scheduling. - **Events** — the scheduler emits `FailedScheduling` with a summary. The message `0/6 nodes are available: 4 Insufficient cpu, 2 node(s) had untolerated taint` is a **tally of why each candidate node was rejected**, and the counts must add up to the fleet size. Here every node failed: four could not fit the pod's CPU **request**, two carry a taint the pod does not tolerate. To schedule, you must satisfy *all* predicates on at least one node — fixing only the CPU request still leaves the two tainted nodes unusable. The scheduler retries with backoff, so the event repeats and its `Age` shows the last attempt. `kubectl get events --sort-by=.lastTimestamp` gives the timeline.
code
bash · 5 lineskubectl get pod api-7c8b-4kd2 -o wide
kubectl describe pod api-7c8b-4kd2 | sed -n '/Events/,$p'
kubectl get events --field-selector reason=FailedScheduling --sort-by=.lastTimestamp
kubectl describe node ip-10-0-2-7 | grep -A12 'Allocated resources'
kubectl get nodes -o widego deeper
Know that Pending means unscheduled, that kubectl describe pod shows a FailedScheduling event, and how to read the reason and count.
Split Pending into unscheduled versus scheduled-but-not-started, map the common reason strings to their fixes, and confirm against kubectl describe node.
Work the tally quantitatively — which single change makes at least one node viable — and know the failure classes upstream of scheduling (quota, webhooks) that produce no pod at all.
Treat chronic Pending as a capacity/topology signal: are requests systematically inflated, are node pools too fragmented, is autoscaling reacting, and what admission policy prevents the class of failure recurring.
## What Pending actually means A pod's `phase` is `Pending` from the moment the API server accepts it until all its containers are created. Two very different situations live in that window: 1. **Unscheduled** — `spec.nodeName` is empty. The scheduler has not found a node that satisfies the pod's constraints. 2. **Scheduled but not started** — a node is assigned and the kubelet is pulling images, attaching volumes, or waiting on an init container. The first check therefore separates the two: `kubectl get pod X -o wide` shows the NODE column, or `kubectl describe pod X` shows the `Node:` field. Empty means it is a scheduling problem and the events will come from the `default-scheduler`. Populated means it is a kubelet problem and you should look at image pulls and volume attachment instead. ## Reading the FailedScheduling event The scheduler works in two phases. **Filtering** runs a set of predicates against every node — does it have enough allocatable CPU/memory for the pod's requests, does the pod tolerate its taints, does it match nodeSelector/nodeAffinity, can the pod's volumes be attached there, do pod affinity/anti-affinity and topology spread constraints hold. Nodes that pass go to **scoring**, which ranks the survivors. If filtering leaves zero nodes, the pod stays Pending and the scheduler records a `FailedScheduling` event. That event's message is deliberately a **tally**: ``` 0/6 nodes are available: 4 Insufficient cpu, 2 node(s) had untolerated taint {workload=gpu: }. preemption: 0/6 nodes are available: 6 No preemption victims found for incoming pod. ``` - `0/6 nodes are available` — six nodes were considered, none survived filtering. - Each clause is `<count> <reason>`. The counts sum to the number of nodes considered (a node is reported under the first predicate that rejected it). - The trailing `preemption:` clause explains whether evicting lower-priority pods could have made room. This structure is the whole diagnostic value: it tells you not just *a* reason but the **distribution** of reasons, which is how you know whether one fix is enough. If five of six nodes say `Insufficient memory` and one says `untolerated taint`, lowering the memory request is likely sufficient. If the counts are split evenly across unrelated reasons, you need more than one change — or more nodes. ## The common reason strings | Message fragment | Meaning | |---|---| | `Insufficient cpu` / `Insufficient memory` | The pod's **requests** exceed the node's remaining allocatable | | `node(s) had untolerated taint {k=v}` | Node is tainted `NoSchedule`/`NoExecute`; the pod lacks a toleration | | `node(s) didn't match Pod's node affinity/selector` | `nodeSelector` or `nodeAffinity` excluded it | | `node(s) had volume node affinity conflict` | The pod's bound PersistentVolume lives in another zone | | `pod has unbound immediate PersistentVolumeClaims` | A PVC is not bound and cannot wait for scheduling | | `node(s) didn't match pod topology spread constraints` | Placing here would exceed `maxSkew` | | `node(s) didn't satisfy existing pods anti-affinity rules` | An anti-affinity rule forbids co-location | | `node(s) were unschedulable` | The node is cordoned (`SchedulingDisabled`) | ## When there is no event at all If `kubectl describe pod` shows an empty Events section, either the events have aged out (they default to a one-hour TTL) or the pod is not being processed by the scheduler you expect — check `spec.schedulerName` and whether the scheduler is running. If the *pod itself* never appears, the problem is upstream of scheduling: inspect the controller (`kubectl describe replicaset` / `deployment`) for `FailedCreate` from a ResourceQuota, LimitRange, or admission webhook. ## A repeatable triage order 1. `kubectl get pod -o wide` — is a node assigned? 2. `kubectl describe pod` — read the `FailedScheduling` tally. 3. Map each clause to a fix: requests too large, missing toleration, selector too narrow, PVC unbound, spread constraint too strict, nodes cordoned. 4. `kubectl describe node` on a representative candidate to confirm allocatable capacity and taints. 5. `kubectl get events --sort-by=.lastTimestamp` for the timeline, since the scheduler retries with exponential backoff and old messages can mislead. ## What to say in an interview Do not jump straight to "the cluster is out of resources". Say that `Pending` splits into unscheduled versus not-yet-started, that `describe pod` gives a per-reason node tally, and that the tally tells you whether one change fixes it. Naming three or four concrete reason strings and their fixes is what makes the answer credible.
- `kubectl describe pod` shows a node assigned but the pod is still Pending. What now?Scheduling already succeeded, so stop looking at the scheduler. The kubelet is stuck earlier in startup: pulling an image (watch for `Pulling`/`Failed to pull`), attaching or mounting a volume (`FailedAttachVolume`, `FailedMount`), or resolving a missing ConfigMap or Secret referenced by the pod. The pod events on that node name the specific step.
- A Deployment reports fewer replicas than desired, but there is no Pending pod at all. Where do you look?The pod was never created, so scheduling is irrelevant. Describe the ReplicaSet — the failure surfaces as a `FailedCreate` event naming a ResourceQuota that would be exceeded, a LimitRange violation, a Pod Security admission rejection, or a validating/mutating webhook error.
saying these in an interview costs you the question
- Assuming Pending always means insufficient resources
- Never running `kubectl describe pod` and guessing from `kubectl get pods` alone
- Ignoring that the FailedScheduling counts sum to the fleet and so reveal whether one fix suffices
- Looking at scheduler logs before reading the event the scheduler already published on the pod
- Confusing Pending with CrashLoopBackOff — a crashing pod has been scheduled and started