skip to content

Allocatable & Bin Packing

A node advertises less capacity than it physically has - kube-reserved, system-reserved and eviction thresholds come off the top - and the scheduler then packs pods against requests, never usage. This is why a fleet can be 95% requested and 20% busy.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

4

In Kubernetes, why does a node's Allocatable differ from its Capacity, and how does the kubelet compute Allocatable?

level: middleimportance: must knowfreq 62%

answer

  1. two lists in node status
  2. three terms come off the top
  3. eviction term: memory and disk only
  4. stock kubelet still loses 100Mi
  5. filter sums requests, not usage

basics

~20 s

Capacity is what the machine has; Allocatable is what the kubelet offers to pods: Capacity minus kube-reserved, minus system-reserved, minus the hard eviction threshold. The scheduler fits pod requests against Allocatable, never against Capacity or live usage.

solid answer

~50 s

A Node reports two numbers: `status.capacity`, what the kubelet discovered on the machine, and `status.allocatable`, what it will let pods request. The kubelet computes Allocatable as Capacity minus `kubeReserved` (kubelet, container runtime), minus `systemReserved` (OS daemons), minus the hard eviction threshold (`evictionHard`), and that last term only applies to memory and ephemeral storage. The reservations default to nothing, but the default `memory.available<100Mi` threshold means memory Allocatable is lower than Capacity even on a stock kubelet. The scheduler's `NodeResourcesFit` filter then checks that the sum of pod **requests** on the node, plus the new pod's requests, stays within Allocatable. It never looks at actual usage. With the default `enforceNodeAllocatable: ["pods"]`, the kubelet also caps the pods cgroup at Capacity minus the two reservations. The eviction threshold is left out of that cap so eviction can fire before the kernel OOM killer.

code

bash · 2 lines
bash
kubectl get node wiki-node-07 -o jsonpath='{.status.capacity}{"\n"}{.status.allocatable}{"\n"}'
kubectl describe node wiki-node-07 | grep -A 12 'Allocated resources'

go deeper

for a junior

Know that a node has Capacity and a smaller Allocatable, and that pods are scheduled against Allocatable using their requests. Be able to find both numbers with kubectl describe node.

for a middle

Recite the formula and which resources each term reduces, including that the eviction threshold hits only memory and ephemeral storage. Work a subtraction and a how-many-pods-fit example out loud.

for a senior

Explain what goes wrong when reservations are undersized, how enforceNodeAllocatable differs from reserving, and why the pods cgroup cap leaves the eviction threshold out.

for a principal

Treat reservation sizing as fleet policy: measure daemon usage per node shape, standardise the kubelet configuration, and weigh the capacity it costs across every node against eviction storms.

## Two numbers on every Node object Every Kubernetes **Node** object carries two resource lists in its status: - **`status.capacity`**: what the kubelet discovered on the machine. That covers CPU cores, memory, ephemeral storage on the root filesystem, a `pods` count taken from the kubelet's `maxPods` setting (default **110**), and any extended resources a device plugin advertises. - **`status.allocatable`**: the part of that capacity the kubelet offers to pods. This is the number scheduling works against. `kubectl describe node` prints both. It also prints an **Allocated resources** section that sums the requests and limits of the pods on that node and shows them as percentages of Allocatable, not of Capacity. ## The subtraction The kubelet computes: **Allocatable = Capacity − kube-reserved − system-reserved − hard eviction threshold** | Term | Configured by | Meant for | Resources it reduces | |---|---|---|---| | kube-reserved | `kubeReserved` / `--kube-reserved` | the kubelet, the container runtime, node-level Kubernetes agents | cpu, memory, ephemeral-storage | | system-reserved | `systemReserved` / `--system-reserved` | OS daemons such as systemd, sshd and the journal | cpu, memory, ephemeral-storage | | hard eviction threshold | `evictionHard` / `--eviction-hard` | a floor the kubelet defends by evicting pods | memory (`memory.available`) and ephemeral storage (`nodefs.available`) only | Defaults matter here: - `kubeReserved` and `systemReserved` are **empty by default**. The upstream kubelet reserves nothing unless told to, though many distributions and managed services set their own values. - On Linux, the default `evictionHard` is `memory.available<100Mi`, `nodefs.available<10%`, `nodefs.inodesFree<5%` and `imagefs.available<15%`. So even a stock kubelet reports memory Allocatable at least 100Mi below Capacity. - CPU has no hard-eviction signal, so **CPU Allocatable is reduced only by the two reservations**. ## A worked example Take one node of a 9-node bare-metal cluster that reports 16 CPUs and `65536Mi` of memory in Capacity. Its kubelet configuration file sets: ```yaml apiVersion: kubelet.config.k8s.io/v1beta1 kind: KubeletConfiguration kubeReserved: cpu: "340m" memory: "1300Mi" systemReserved: cpu: "410m" memory: "1700Mi" evictionHard: memory.available: "500Mi" nodefs.available: "10%" nodefs.inodesFree: "5%" imagefs.available: "15%" enforceNodeAllocatable: - pods ``` 1. CPU Allocatable = 16000m − 340m − 410m = **15250m**. No eviction term applies to CPU. 2. Memory Allocatable = 65536Mi − 1300Mi − 1700Mi − 500Mi = **62036Mi**. 3. An internal wiki pod with an embedded database requests `2.6Gi` of memory, which is 2662.4Mi. Ignoring DaemonSet pods, 23 copies fit (61235.2Mi). A 24th copy (63897.6Mi) is rejected even if the other 23 are using a fraction of what they requested. ## How the scheduler uses it The `NodeResourcesFit` filter plugin in kube-scheduler adds up the **requests** of every pod already bound to the node, adds the incoming pod's requests, and compares the total with Allocatable for each resource, including the `pods` count. The incoming pod's figure is its **effective request**. That is the larger of its biggest init container and the sum of its app and sidecar containers, plus any pod `overhead` from its RuntimeClass. Actual consumption plays no part in this check. That is why a node can be "full" to the scheduler while `kubectl top node` shows it mostly idle, and why a fleet can be 95% requested and 20% busy. ## Reserved is not the same as enforced Reserving capacity only shrinks what the scheduler hands out. Enforcement is a separate kubelet setting: - With the default `enforceNodeAllocatable: ["pods"]`, the kubelet caps the top-level **pods cgroup** at Capacity minus kube-reserved minus system-reserved. In the example that is 62536Mi. The eviction threshold is left out of the cap on purpose: pods can push the node into the eviction zone, and the kubelet can react before the kernel OOM killer does. - Adding `kube-reserved` or `system-reserved` to that list also caps those daemons' own cgroups. That requires `--kube-reserved-cgroup` / `--system-reserved-cgroup`, and it is risky: a capped kubelet or sshd that exceeds its reservation can be starved or OOM-killed. - If the reservations are too **small**, system daemons eat into memory the scheduler already promised to pods. The node then hits memory pressure and evicts pods, even though every pod stayed within its request. ## Why it matters Allocatable is the denominator for every node-level capacity figure: requested percentage, how many pods of a given size fit, and how much headroom survives a node loss. Treating Capacity as the budget overstates the cluster. Treating live usage as the budget misunderstands what the scheduler actually checks.

  • What happens when kube-reserved and system-reserved are set far below what the node's daemons really use?
    The scheduler still hands out the full Allocatable, so pods can request memory the daemons are already using. Under load, real free memory falls to the `memory.available` threshold and the kubelet starts evicting pods, even though each pod stayed within its request. The fix is to size the reservations from measured daemon usage. Enforcing caps on the daemons' own cgroups is not the answer, because that risks starving the kubelet itself.
  • Which request figure does the NodeResourcesFit filter use for a pod that has init containers, a sidecar and a RuntimeClass overhead?
    It uses the pod's effective request. That is the larger of its biggest regular init container and the sum of its app containers plus sidecars (restartable init containers), plus the RuntimeClass `overhead`. The same figure is added to the node's requested total once the pod is bound, so a large one-off init container can block placement even if it runs for only seconds.
  • Why does the kubelet cap the pods cgroup at capacity minus the reservations rather than at Allocatable?
    Leaving the hard eviction threshold outside the cgroup cap gives pods room to push free memory below `memory.available`. The kubelet can then notice and evict a chosen pod. If the cap included the threshold, the kernel's cgroup OOM killer would act first, and it picks victims without the kubelet's ranking.

A hotel with 200 rooms keeps some for staff and some empty as an emergency buffer, and only the rest go on the booking site. The booking site counts reservations, not whether guests are actually in their rooms.

saying these in an interview costs you the question

  • Allocatable is the memory currently free on the node
  • Allocatable equals Capacity unless you configure kube-reserved or system-reserved
  • The scheduler compares a pod's requests with the node's current usage
  • The hard eviction threshold also reduces CPU Allocatable
  • kube-reserved is enforced as a cgroup limit on the kubelet by default
open as a page

On a Kubernetes cluster, why can a pod requesting 2.6 GiB of memory stay unplaceable while free allocatable memory across all nodes totals 11.3 GiB, and how do you reduce that stranded capacity?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Pods must fit whole on one node, so free capacity summed across nodes means nothing. Many small leftover gaps, or free CPU stuck next to exhausted memory, strand capacity. Reduce it by repacking pods, standardising request shapes and scoring toward fuller nodes.

open as a page

Your 9-node bare-metal Kubernetes cluster is 95% requested but only 20% busy, and finance wants two nodes back. How do you respond as the platform owner?

level: principalimportance: should knowfreq 30%

basics

~10 s

Not yet: the scheduler budgets requests, and at 95% requested the cluster cannot even absorb one node loss. Reduce requests first, repack, keep an N+1 ceiling near 89% requested, and only then return hardware.

open as a page

Why does kube-scheduler never move running pods when Kubernetes nodes become unbalanced, and what does the descheduler do about it?

level: middleimportance: nice to knowfreq 33%

basics

~20 s

kube-scheduler decides only for unbound pods; once bound, a pod stays put. The descheduler is a separate add-on that evicts selected pods through the Eviction API so their controllers recreate them and the scheduler places them again.

open as a page