The Kubernetes scheduler reports `Insufficient cpu` for a pod even though `kubectl top nodes` shows the cluster is only 30% utilised. Explain why, and what actually determines whether a pod fits on a node.
answer
- scheduler sums REQUESTS, not usage
- allocatable = capacity − kube/system-reserved − eviction reserve
- limits ≠ scheduling; limits = throttle / OOMKill
- describe node → Allocated resources table
- fragmentation: free total ≠ fits on one node
basics
~20 sThe scheduler sums pods' CPU/memory requests, not actual usage, against the node's allocatable (capacity minus kube-reserved, system-reserved and eviction thresholds). A cluster idle at 30% can be fully reserved if requests are inflated. Fix by right-sizing requests or adding capacity.
solid answer
~60 sScheduling is a **reservation** system, not a usage system. - **Requests** are what the scheduler adds up. A node fits the pod only if `sum(requests of existing pods) + pod's request <= node allocatable` for every resource. - **Allocatable** = `capacity − kube-reserved − system-reserved − hard eviction thresholds`. A 4-vCPU/16Gi node typically advertises noticeably less. - **Limits are irrelevant to scheduling**; they are runtime ceilings enforced by cgroups (CPU throttling, memory OOMKill). So a node whose pods request 3.8 of 3.86 allocatable CPU is full, no matter that they idle at 5%. `kubectl top` measures usage; `kubectl describe node` shows the *Allocated resources* table with requests and their percentages — that is the number the scheduler uses. Remedies: right-size requests from observed p95 usage (this is usually the real bug — copy-pasted `cpu: 1` everywhere), check `kubectl describe node` for the largest reservers, add nodes or enable the cluster autoscaler, and note that fragmentation matters — one big pod may not fit even when the aggregate free total looks sufficient.
code
bash · 7 lineskubectl describe node ip-10-0-2-7 | grep -A10 'Allocated resources'
kubectl get node ip-10-0-2-7 -o jsonpath='{.status.allocatable}'
kubectl top node ip-10-0-2-7
kubectl get pods -A -o custom-columns=\
NS:.metadata.namespace,NAME:.metadata.name,\
CPUREQ:.spec.containers[*].resources.requests.cpu \
--field-selector spec.nodeName=ip-10-0-2-7go deeper
Know that the scheduler adds up requests rather than usage, and that a node can be fully reserved while sitting idle.
Define allocatable versus capacity, state that limits do not affect placement, and read the Allocated resources table to find the biggest reservers.
Right-size requests from observed p95/p99, recognise fragmentation, understand the compressible/incompressible asymmetry between CPU and memory, and connect this to cluster autoscaler behaviour and cost.
Set the platform policy: LimitRange defaults, quota per team, VPA recommendations feeding request hygiene, node-pool shapes matched to workload sizes, and an explicit target for allocation-versus-utilisation efficiency.
## Requests versus limits versus usage Three different numbers get confused constantly: - **Request** — a reservation. The scheduler guarantees the node has this much unallocated. For CPU it also sets the cgroup weight (`cpu.shares`) that decides how contended CPU is shared. - **Limit** — a ceiling enforced at runtime. Exceeding a CPU limit means **throttling** (`cpu.max` quota); exceeding a memory limit means the container is **OOMKilled**. - **Usage** — what the process actually consumes right now, which is what `kubectl top` and your metrics stack report. **Only requests participate in scheduling.** This single fact explains the whole puzzle: a cluster can be 100% *allocated* and 20% *utilised* at the same time. The scheduler is not trying to bin-pack real load; it is honouring promises, because a pod that requested 2 CPU must still get 2 CPU when its traffic arrives at 3 a.m. ## Capacity versus allocatable A node does not offer all its hardware to pods: ``` allocatable = capacity − kube-reserved (kubelet, container runtime) − system-reserved (sshd, systemd, kernel overhead) − eviction-hard (the reserve kept so the kubelet can act before the node dies) ``` On a managed 4 vCPU / 16 GiB node, allocatable is commonly around 3.9 vCPU and 14 GiB. `kubectl describe node` prints both Capacity and Allocatable, and beneath them an *Allocated resources* table listing the summed **requests** and their percentage of allocatable. That table is the scheduler's worldview. There is also a hard cap on pods per node (default 110, often lower on managed platforms). A node with plenty of free CPU can still report `Too many pods`. ## Why `kubectl top` misleads `kubectl top nodes` reports CPU and memory *usage* from metrics-server. A fleet showing 30% usage with 95% of CPU requested is the classic pattern in organisations where every team copied `requests: {cpu: "1", memory: "2Gi"}` from a template. The cluster is full of reservations nobody uses, so the bill is high, the scheduler refuses new work, and everything looks idle. The correct comparison is `kubectl describe node` (requests) against your metrics (usage), per workload. Any pod whose p95 usage is a small fraction of its request is over-requesting. ## Fragmentation Even when the aggregate free capacity is large, a single pod may not fit. Six nodes with 300m free each total 1.8 CPU, but a pod requesting 1 CPU fits nowhere — the scheduler places whole pods on single nodes. Symptoms: `Insufficient cpu` on every node while cluster-wide free capacity looks ample. Remedies are fewer/larger nodes for large pods, dedicated node pools, or a cluster autoscaler that can add an appropriately sized node. ## CPU is compressible, memory is not CPU is a rate: a container over its request is throttled but survives, so modest CPU over-commit is tolerable and CPU limits are often deliberately omitted to avoid throttling latency-sensitive services. Memory is not reclaimable: exceeding a limit means the process dies. This asymmetry is why memory requests should be close to real peak usage while CPU requests can track something like p90 with room to burst. ## Diagnosing and fixing 1. `kubectl describe node <node>` — read *Allocated resources*; compare CPU/memory requests as a percentage of allocatable. 2. List the biggest reservers: sort pods on the node by `spec.containers[*].resources.requests.cpu`. 3. Compare each workload's request to its observed p95/p99 usage. Cut the ones that are wildly over-provisioned. A Vertical Pod Autoscaler in recommendation mode can generate these numbers for you. 4. Check for fragmentation — is the free capacity spread thinly across nodes? 5. If the requests are honest, the cluster genuinely needs more capacity: add nodes or let the cluster autoscaler do it (note it also reasons about *requests*, so it scales up exactly when pods are Pending for insufficiency). 6. Consider `LimitRange` defaults per namespace so pods without explicit requests do not get an unhelpful default, and `ResourceQuota` to make teams' reservations visible and bounded. ## Interview framing Lead with "the scheduler sums requests, not usage," define allocatable versus capacity, mention that limits do not affect placement, and finish with the two real-world remedies — right-size requests from observed usage, and recognise fragmentation when aggregate free capacity is misleading. That is a complete, senior-sounding answer even at middle level.
- If you remove a pod's CPU limit but keep its request, what changes for scheduling and at runtime?Scheduling is unaffected — placement only ever considered the request. At runtime the container is no longer throttled by a CFS quota, so it can burst into idle CPU on the node; under contention the kernel shares CPU proportionally to requests, so it still gets at least its reservation. The tradeoff is less predictable neighbours and noisier latency for co-located pods.
- Every node reports `Insufficient memory` but the cluster has 20 GiB free in total. Why might the pod still not schedule?Fragmentation. The scheduler places a whole pod on one node, so 20 GiB spread as 2 GiB across ten nodes cannot host a pod requesting 4 GiB. You need a node with 4 GiB of contiguous free allocatable — achieved by adding a larger node, consolidating workloads, or letting the cluster autoscaler provision a node from a bigger instance type.
saying these in an interview costs you the question
- Claiming the scheduler looks at actual CPU/memory usage
- Thinking limits, not requests, determine whether a pod fits
- Comparing the pod's request against node Capacity instead of Allocatable
- Missing fragmentation — assuming aggregate free capacity implies a fit
- Treating `kubectl top` output as proof that a cluster has schedulable room