skip to content

On a Kubernetes cluster, why can a pod requesting 2.6 GiB of memory stay unplaceable while free allocatable memory across all nodes totals 11.3 GiB, and how do you reduce that stranded capacity?

level: seniorimportance: should knowfreq 44%

answer

  1. a pod lands on one node
  2. largest free block, not sum
  3. one resource exhausts, others strand
  4. scoring changes only future placements
  5. repack, standardise shapes

basics

~20 s

Pods must fit whole on one node, so free capacity summed across nodes means nothing. Many small leftover gaps, or free CPU stuck next to exhausted memory, strand capacity. Reduce it by repacking pods, standardising request shapes and scoring toward fuller nodes.

solid answer

~50 s

The scheduler places a pod on **one** node, so what matters is the largest free block on a single node, not the cluster total. If nine nodes each have 1.1–1.4 GiB left (Allocatable minus requested), the cluster has 11.3 GiB free, but no node can take a 2.6 GiB request. That is **fragmentation**. The multi-resource version is **stranding**: a node with six free cores and 400Mi of free memory cannot use those cores, because its memory is gone. To fix it, first measure free Allocatable per node and per resource. Then consolidate by evicting smaller pods off one or two nodes so they re-land elsewhere, by hand or with a descheduler compaction policy. Make request sizes and CPU-to-memory ratios match the node shape. Change scoring (for example `MostAllocated`) so future placements fill nodes instead of spreading, knowing it does nothing for pods already running.

go deeper

for a junior

Remember that a pod must fit entirely on one node, so free space spread across many nodes may be unusable.

for a middle

Explain fragmentation and multi-resource stranding with a concrete example, and show how to compute free Allocatable per node from requests.

for a senior

Diagnose from per-node, per-resource free blocks, pick between repacking, compaction and scoring changes, and plan around single-replica stateful pods and PDBs when moving things.

for a principal

Decide fleet-level standards, such as request size classes and node shapes matched to workload ratios, and choose how much deliberate headroom to keep for large pods.

## Why a cluster total lies Kubernetes schedules against **per-node** budgets. For each node, free capacity for a resource is its **Allocatable** minus the sum of the **requests** of pods bound there. A pod fits only if one node has enough free capacity for **every** resource the pod requests. The cluster-wide sum of free capacity is therefore a poor predictor of what can still be placed. Take a 9-node bare-metal cluster running an internal wiki with an embedded database. A new wiki replica requests `2.6Gi` of memory. Free memory per node is: | Node | Free memory (Allocatable − requested) | |---|---| | n1, n5 | 1.1 GiB each | | n3, n8 | 1.2 GiB each | | n6 | 1.25 GiB | | n4 | 1.3 GiB | | n7 | 1.35 GiB | | n2, n9 | 1.4 GiB each | | **Total** | **11.3 GiB** | The largest free block is 1.4 GiB, so the pod stays Pending. The cluster "has room" for four such pods on paper and room for none in practice. ## Two shapes of waste - **Fragmentation (one resource):** leftovers scattered across nodes, each smaller than the pods you need to place. It builds up naturally, because the default `LeastAllocated` scoring spreads pods to the emptiest node and deletions leave holes behind. - **Stranding (several resources):** one resource runs out first and the others on that node become unusable. A node with 6 free CPUs and 400Mi of free memory has stranded CPU. A node full on CPU requests has stranded memory. This happens when the CPU-to-memory ratio of the workloads differs from the node's ratio. - **Count stranding:** the `pods` resource (from `maxPods`, default 110) can run out on nodes crowded with tiny pods while CPU and memory are still free. ## Diagnosing it 1. For each node, compute Allocatable minus requested for CPU, memory and pods. The **Allocated resources** section of `kubectl describe node` shows requested as a percentage of Allocatable. 2. Look at the **largest free block** per resource, not the sum, and compare it with the request sizes of the pods you actually need to place. 3. Check which resource runs out first on each node. If memory is at 97% and CPU at 40% across the fleet, the problem is the shape of the requests, not the total amount. 4. Confirm the pending pod's effective request (including init containers and overhead) is really what you think it is. ## Reducing it | Lever | What it changes | Limitation | |---|---|---| | Manual repack (cordon + drain one or two lightly used nodes) | frees whole-node blocks now | disruptive; limited by PodDisruptionBudgets and single-replica pods | | Descheduler compaction (`HighNodeUtilization`) | evicts pods from under-used nodes so they re-land on fuller ones | documented to need `MostAllocated` scoring, otherwise pods spread straight back | | Scoring toward fuller nodes (`MostAllocated`, or `RequestedToCapacityRatio`) | future placements leave whole nodes free | does nothing for pods already running; concentrates blast radius | | Standard request sizes | leftovers become reusable | teams must accept a small set of sizes | | Match the CPU:memory ratio of requests to the node shape | less stranding | may require different node hardware or separate pools | | Keep a free block for large pods | the big pod always has somewhere to land | capacity deliberately kept idle | Some points to weigh: - **Repacking is eviction.** The wiki's database pod may be a single replica with local data. Moving it causes downtime, so move the small, replicated pods instead. - **Changing scoring alone is not a fix.** It only affects where new pods go. Existing fragmentation stays until pods churn or something evicts them. - **On bare metal there is no extra node to add**, so consolidation and request shape are the main tools. Adding nodes belongs to node autoscaling, where it exists at all. - **Watch the pods count on nodes full of sidecars and small jobs.** "Too many pods" is stranding too. ## What to report A useful fleet dashboard shows, per resource, requested vs Allocatable **and** the largest free block per node. Utilisation from metrics sits next to it, but it answers a different question. If you track only the cluster-wide requested percentage, the cluster can look fine while the next large pod has nowhere to go.

  • Does switching the scheduler to MostAllocated scoring fix fragmentation that already exists?
    No. Scoring runs only when a pod is being placed, so it shapes where new and rescheduled pods land. Pods already bound stay where they are. To recover existing holes, evict or drain pods so they are placed again, by hand or through a descheduler compaction policy, and let the new scoring pack them.
  • How does a mismatch between workload CPU:memory ratio and node shape create stranded capacity?
    If nodes have 4 GiB per core but workloads request 8 GiB per core, memory runs out when only half the cores are requested, and the rest of the CPU is stranded. Fixes are node shapes closer to the workload ratio, separate pools for memory-heavy workloads, or requests brought closer to real need so the ratio narrows.

A car park with twenty scattered single gaps cannot fit a bus that needs three adjacent bays, even though twenty bays are "free".

saying these in an interview costs you the question

  • The cluster has 11.3 GiB free, so the 2.6 GiB pod must fit somewhere
  • Free capacity means memory that kubectl top shows as unused
  • Switching to MostAllocated scoring repacks the pods already running
  • Only memory can be stranded; CPU is always usable
  • The pods count per node never limits placement