skip to content

kube-scheduler's NodeResourcesFit plugin scores nodes with a LeastAllocated strategy by default, and can be configured to MostAllocated. What does each choice do to a cluster, and how would you decide between them?

level: principalimportance: nice to knowfreq 26%

answer

  1. Scoring strategy of NodeResourcesFit; ranking only, filter unchanged
  2. LeastAllocated = spread (headroom, small blast radius, idle spend)
  3. MostAllocated = pack (autoscaler can delete empty nodes, cheaper, noisier)
  4. Packs requests, not usage — right-sizing is a prerequisite
  5. Two profiles + spec.schedulerName; keep BalancedAllocation and topology spread

basics

~20 s

LeastAllocated favours the emptiest node, spreading pods for headroom and blast-radius safety. MostAllocated packs pods onto the fullest node that still fits, leaving nodes empty enough to be removed and cutting cost. Choose spread for latency-sensitive services, packing for batch on autoscaled or spot capacity.

solid answer

~60 s

Both are scoring strategies of the **`NodeResourcesFit`** score plugin, set per scheduler **profile** in `KubeSchedulerConfiguration`. They change ranking only — feasibility is still decided by the filter phase. - **`LeastAllocated`** (default) scores a node higher the more *unrequested* capacity it has. Effect: pods spread across nodes; each pod gets more real headroom for bursts above its requests; a node failure takes out a smaller share of any one workload; but the cluster carries fragmentation and idle nodes nobody can remove. - **`MostAllocated`** scores fuller nodes higher, bin-packing. Effect: fewer nodes carry the load, so the Cluster Autoscaler can drain and delete the empties — real money saved. Costs: noisier neighbours, larger blast radius per node, and more disruption when a packed node is drained. - **`RequestedToCapacityRatio`** lets you shape the curve between them and weight resources unevenly (useful for GPUs, where you want *full* GPU nodes). Decide by workload: latency-sensitive services on fixed capacity → spread, plus `NodeResourcesBalancedAllocation` and topology spread constraints for availability. Batch or spot-backed capacity where node count is the bill → pack. Run both as two profiles and let workloads opt in via `spec.schedulerName`.

code

yaml · 23 lines
yaml
apiVersion: kubescheduler.config.k8s.io/v1
kind: KubeSchedulerConfiguration
profiles:
  - schedulerName: default-scheduler
  - schedulerName: bin-packing
    pluginConfig:
      - name: NodeResourcesFit
        args:
          scoringStrategy:
            type: RequestedToCapacityRatio
            resources:
              - name: nvidia.com/gpu
                weight: 10
              - name: cpu
                weight: 1
              - name: memory
                weight: 1
            requestedToCapacityRatio:
              shape:
                - utilization: 0
                  score: 0
                - utilization: 100
                  score: 10

go deeper

for a junior

Know the direction of each: LeastAllocated prefers emptier nodes, MostAllocated prefers fuller ones, and both only rank nodes that already fit.

for a middle

Explain the tradeoff in terms of burst headroom and blast radius versus node count and cost, and that it is configured per scheduler profile.

for a senior

Connect it to Cluster Autoscaler scale-down behaviour, fragmentation and large-pod feasibility, and pair packing with topology spread and accurate requests.

for a principal

Decide per workload class rather than cluster-wide, quantify the change with node count, throttling, eviction and latency metrics, and recognise node shape, PDBs and request accuracy as the constraints that decide whether the strategy pays.

## What the strategies actually compute The filter phase has already reduced the cluster to nodes where the pod *fits*. Scoring only decides which of those is preferred. `NodeResourcesFit` computes, per resource, a ratio of **requested** to **allocatable** on the node *including* the incoming pod, then converts it to a 0–100 score: - `LeastAllocated`: score rises as the requested fraction falls — the emptiest feasible node wins. - `MostAllocated`: score rises as the requested fraction rises — the fullest feasible node wins. - `RequestedToCapacityRatio`: you supply a shape (a set of utilisation→score points) and per-resource weights, so you can express "prefer nodes between 60% and 80% utilised" or "pack GPUs hard but spread CPU". Remember the currency is **requests**, not observed usage. Bin-packing by requests packs *reservations*; if requests are wildly larger than reality you pack empty air, and if they are smaller than reality you pack contention. Any packing strategy is downstream of request accuracy — which is why VPA recommendations, or at least a periodic right-sizing pass, are a prerequisite rather than a nice-to-have. ## The case for spreading (LeastAllocated) - **Burst headroom.** A container may exceed its CPU request whenever the node has spare cycles; CFS only throttles at the *limit*. On an emptier node, bursts are actually servable, so p99 latency improves. Memory is harsher: exceeding a limit is an OOM kill, and on a packed node a memory-pressure eviction can take neighbours with it. - **Blast radius.** If each service has 3 replicas across 20 lightly loaded nodes, losing a node costs a fraction of one service. On a packed cluster the same node might hold every replica's worth of several services. - **Simpler operations.** Draining a node for an upgrade moves fewer pods, so PodDisruptionBudgets are easier to satisfy. The cost is money and fragmentation: nodes hover at 40% requested, none is empty enough for the autoscaler to remove, and you pay for the gaps. ## The case for packing (MostAllocated) - **The Cluster Autoscaler removes empty nodes, not half-empty ones.** It scales down a node when its pods can be rescheduled elsewhere and utilisation is below a threshold for a period. Spreading actively fights this: every node holds a few pods, so none qualifies. Packing concentrates load and *creates* removable nodes. On an autoscaled cluster the scoring strategy is one of the few direct levers on node count, and node count is the bill. - **Large-pod feasibility.** Fragmentation is the enemy of a pod that needs 8 CPU: twenty nodes with 4 CPU free each cannot host it. Packing preserves contiguous space. - **Spot/preemptible fleets** are cheaper per node but expected to churn; concentrating disposable batch work on few nodes matches the model. The cost is interference (CPU cache, memory bandwidth, disk and network I/O are not isolated by requests), larger per-node blast radius, and heavier drains. ## Deciding The decision is not global — it is per workload class, which is why **profiles** exist: ```yaml profiles: - schedulerName: default-scheduler # services: spread - schedulerName: bin-packing # batch: pack pluginConfig: - name: NodeResourcesFit args: scoringStrategy: type: MostAllocated ``` Workloads opt in with `spec.schedulerName`. A useful default split: customer-facing services stay on the spreading default and rely on **topology spread constraints** for availability; batch, CI and analytics use the packing profile on their own node pool. Other knobs that belong in this conversation: - **`NodeResourcesBalancedAllocation`** prefers nodes where CPU and memory end up *evenly* consumed, avoiding the classic stranded-resource shape (a node with 12 GB free but no CPU). Keep it enabled under either strategy; it attacks fragmentation from a different angle. - **Topology spread constraints** are the correct tool for availability, not the scoring strategy. Packing plus a `maxSkew: 1` spread over `topology.kubernetes.io/zone` gives you cost *and* survivability — the spread constraint filters or scores for distribution while `MostAllocated` chooses densely *within* the allowed set. - **Node shape matters more than the strategy.** Very large nodes amplify blast radius and make packing decisions coarse; very small ones make DaemonSet overhead (one copy per node) a real tax. Sizing the pool is often the higher-leverage decision. - **PodDisruptionBudgets** determine whether the autoscaler can actually act on the empty-ish nodes packing creates. A batch tier with restrictive PDBs blocks scale-down and negates the saving. ## How to validate the choice Instrument before and after: node count over time, cluster-wide requested-versus-allocatable ratio, p99 latency and CPU throttling for the service tier, OOM-kill and eviction rates, and autoscaler scale-down event counts. If packing does not reduce node count, the binding constraint is elsewhere — usually PDBs, local storage, or pods that cannot be rescheduled — and you have taken on interference risk for nothing. ## The honest summary There is no universally right answer, which is exactly why it is configurable. Spreading buys latency stability and small blast radius with money; packing buys money with interference risk and larger failure domains. On fixed on-prem capacity the money argument mostly evaporates and spreading usually wins; on an elastic cloud cluster with accurate requests and disposable batch work, packing pays for itself.

  • Why does spreading pods with LeastAllocated make an autoscaled cluster more expensive?
    The Cluster Autoscaler removes a node only when its pods can be rescheduled elsewhere and it sits below a utilisation threshold. Spreading puts a few pods on every node, so no node ever becomes empty enough to qualify and the fleet stays at its high-water mark. Packing concentrates the same workload onto fewer nodes and deliberately creates the removable, near-empty nodes scale-down needs.
  • If you bin-pack, how do you still protect availability?
    Use topology spread constraints — for example maxSkew 1 over topology.kubernetes.io/zone or kubernetes.io/hostname with DoNotSchedule for the replicas that matter. Those constrain the feasible or preferred set, and MostAllocated then packs densely within it, so you get cost savings and distribution together. Availability should come from explicit spread rules and PodDisruptionBudgets, not as a side effect of the scoring strategy.
  • What must be true about resource requests before bin-packing is safe?
    Requests must approximate real usage, because packing works on reservations rather than observed load. If requests are far above reality you pack empty space and save nothing; if they are far below it you concentrate real contention and hit throttling, memory pressure and evictions. Gather VPA recommendations or historical usage and right-size first, then change the scoring strategy.

Seating a restaurant: spread everyone out and diners are comfortable but you can never close a section; fill tables from the fullest section first and you can send half the staff home — at the cost of elbows.

saying these in an interview costs you the question

  • Saying the scoring strategy can place a pod on a node that failed the filter phase
  • Treating MostAllocated as a way to oversubscribe nodes rather than a ranking preference
  • Bin-packing without first right-sizing requests, so you pack reservations that bear no relation to usage
  • Relying on LeastAllocated for availability instead of topology spread constraints and PodDisruptionBudgets
  • Assuming the choice is cluster-wide when profiles allow per-workload strategies

context