skip to content

The Kubernetes scheduler reports `Insufficient cpu` for a pod even though `kubectl top nodes` shows the cluster is only 30% utilised. Explain why, and what actually determines whether a pod fits on a node.

level: middleimportance: must knowfreq 68%

answer

  1. scheduler sums REQUESTS, not usage
  2. allocatable = capacity − kube/system-reserved − eviction reserve
  3. limits ≠ scheduling; limits = throttle / OOMKill
  4. describe node → Allocated resources table
  5. fragmentation: free total ≠ fits on one node

basics

~20 s

The scheduler sums pods' CPU/memory requests, not actual usage, against the node's allocatable (capacity minus kube-reserved, system-reserved and eviction thresholds). A cluster idle at 30% can be fully reserved if requests are inflated. Fix by right-sizing requests or adding capacity.

solid answer

~60 s

Scheduling is a **reservation** system, not a usage system. - **Requests** are what the scheduler adds up. A node fits the pod only if `sum(requests of existing pods) + pod's request <= node allocatable` for every resource. - **Allocatable** = `capacity − kube-reserved − system-reserved − hard eviction thresholds`. A 4-vCPU/16Gi node typically advertises noticeably less. - **Limits are irrelevant to scheduling**; they are runtime ceilings enforced by cgroups (CPU throttling, memory OOMKill). So a node whose pods request 3.8 of 3.86 allocatable CPU is full, no matter that they idle at 5%. `kubectl top` measures usage; `kubectl describe node` shows the *Allocated resources* table with requests and their percentages — that is the number the scheduler uses. Remedies: right-size requests from observed p95 usage (this is usually the real bug — copy-pasted `cpu: 1` everywhere), check `kubectl describe node` for the largest reservers, add nodes or enable the cluster autoscaler, and note that fragmentation matters — one big pod may not fit even when the aggregate free total looks sufficient.

code

bash · 7 lines
bash
kubectl describe node ip-10-0-2-7 | grep -A10 'Allocated resources'
kubectl get node ip-10-0-2-7 -o jsonpath='{.status.allocatable}'
kubectl top node ip-10-0-2-7
kubectl get pods -A -o custom-columns=\
NS:.metadata.namespace,NAME:.metadata.name,\
CPUREQ:.spec.containers[*].resources.requests.cpu \
  --field-selector spec.nodeName=ip-10-0-2-7

go deeper

for a junior

Know that the scheduler adds up requests rather than usage, and that a node can be fully reserved while sitting idle.

for a middle

Define allocatable versus capacity, state that limits do not affect placement, and read the Allocated resources table to find the biggest reservers.

for a senior

Right-size requests from observed p95/p99, recognise fragmentation, understand the compressible/incompressible asymmetry between CPU and memory, and connect this to cluster autoscaler behaviour and cost.

for a principal

Set the platform policy: LimitRange defaults, quota per team, VPA recommendations feeding request hygiene, node-pool shapes matched to workload sizes, and an explicit target for allocation-versus-utilisation efficiency.

## Requests versus limits versus usage Three different numbers get confused constantly: - **Request** — a reservation. The scheduler guarantees the node has this much unallocated. For CPU it also sets the cgroup weight (`cpu.shares`) that decides how contended CPU is shared. - **Limit** — a ceiling enforced at runtime. Exceeding a CPU limit means **throttling** (`cpu.max` quota); exceeding a memory limit means the container is **OOMKilled**. - **Usage** — what the process actually consumes right now, which is what `kubectl top` and your metrics stack report. **Only requests participate in scheduling.** This single fact explains the whole puzzle: a cluster can be 100% *allocated* and 20% *utilised* at the same time. The scheduler is not trying to bin-pack real load; it is honouring promises, because a pod that requested 2 CPU must still get 2 CPU when its traffic arrives at 3 a.m. ## Capacity versus allocatable A node does not offer all its hardware to pods: ``` allocatable = capacity − kube-reserved (kubelet, container runtime) − system-reserved (sshd, systemd, kernel overhead) − eviction-hard (the reserve kept so the kubelet can act before the node dies) ``` On a managed 4 vCPU / 16 GiB node, allocatable is commonly around 3.9 vCPU and 14 GiB. `kubectl describe node` prints both Capacity and Allocatable, and beneath them an *Allocated resources* table listing the summed **requests** and their percentage of allocatable. That table is the scheduler's worldview. There is also a hard cap on pods per node (default 110, often lower on managed platforms). A node with plenty of free CPU can still report `Too many pods`. ## Why `kubectl top` misleads `kubectl top nodes` reports CPU and memory *usage* from metrics-server. A fleet showing 30% usage with 95% of CPU requested is the classic pattern in organisations where every team copied `requests: {cpu: "1", memory: "2Gi"}` from a template. The cluster is full of reservations nobody uses, so the bill is high, the scheduler refuses new work, and everything looks idle. The correct comparison is `kubectl describe node` (requests) against your metrics (usage), per workload. Any pod whose p95 usage is a small fraction of its request is over-requesting. ## Fragmentation Even when the aggregate free capacity is large, a single pod may not fit. Six nodes with 300m free each total 1.8 CPU, but a pod requesting 1 CPU fits nowhere — the scheduler places whole pods on single nodes. Symptoms: `Insufficient cpu` on every node while cluster-wide free capacity looks ample. Remedies are fewer/larger nodes for large pods, dedicated node pools, or a cluster autoscaler that can add an appropriately sized node. ## CPU is compressible, memory is not CPU is a rate: a container over its request is throttled but survives, so modest CPU over-commit is tolerable and CPU limits are often deliberately omitted to avoid throttling latency-sensitive services. Memory is not reclaimable: exceeding a limit means the process dies. This asymmetry is why memory requests should be close to real peak usage while CPU requests can track something like p90 with room to burst. ## Diagnosing and fixing 1. `kubectl describe node <node>` — read *Allocated resources*; compare CPU/memory requests as a percentage of allocatable. 2. List the biggest reservers: sort pods on the node by `spec.containers[*].resources.requests.cpu`. 3. Compare each workload's request to its observed p95/p99 usage. Cut the ones that are wildly over-provisioned. A Vertical Pod Autoscaler in recommendation mode can generate these numbers for you. 4. Check for fragmentation — is the free capacity spread thinly across nodes? 5. If the requests are honest, the cluster genuinely needs more capacity: add nodes or let the cluster autoscaler do it (note it also reasons about *requests*, so it scales up exactly when pods are Pending for insufficiency). 6. Consider `LimitRange` defaults per namespace so pods without explicit requests do not get an unhelpful default, and `ResourceQuota` to make teams' reservations visible and bounded. ## Interview framing Lead with "the scheduler sums requests, not usage," define allocatable versus capacity, mention that limits do not affect placement, and finish with the two real-world remedies — right-size requests from observed usage, and recognise fragmentation when aggregate free capacity is misleading. That is a complete, senior-sounding answer even at middle level.

  • If you remove a pod's CPU limit but keep its request, what changes for scheduling and at runtime?
    Scheduling is unaffected — placement only ever considered the request. At runtime the container is no longer throttled by a CFS quota, so it can burst into idle CPU on the node; under contention the kernel shares CPU proportionally to requests, so it still gets at least its reservation. The tradeoff is less predictable neighbours and noisier latency for co-located pods.
  • Every node reports `Insufficient memory` but the cluster has 20 GiB free in total. Why might the pod still not schedule?
    Fragmentation. The scheduler places a whole pod on one node, so 20 GiB spread as 2 GiB across ten nodes cannot host a pod requesting 4 GiB. You need a node with 4 GiB of contiguous free allocatable — achieved by adding a larger node, consolidating workloads, or letting the cluster autoscaler provision a node from a bigger instance type.

saying these in an interview costs you the question

  • Claiming the scheduler looks at actual CPU/memory usage
  • Thinking limits, not requests, determine whether a pod fits
  • Comparing the pod's request against node Capacity instead of Allocatable
  • Missing fragmentation — assuming aggregate free capacity implies a fit
  • Treating `kubectl top` output as proof that a cluster has schedulable room

context