skip to content

Walk through how kube-scheduler decides which node an unscheduled pod lands on, naming the phases it goes through and what each one produces.

level: middleimportance: must knowfreq 66%

answer

  1. Watch pods with empty spec.nodeName → priority queue
  2. Filter = hard yes/no per node (predicates) → feasible set
  3. Score = 0-100 per plugin × weight, highest wins, random tiebreak
  4. Requests, not actual usage, drive NodeResourcesFit
  5. Reserve/Permit → assume → async binding cycle; kubelet re-admits

basics

~20 s

The scheduler watches for pods with no node assigned, then runs two phases: filter, which reduces all nodes to the feasible ones (fit, taints, affinity, volumes), and score, which ranks the survivors 0-100 by weighted plugins. The best-scoring node wins, ties broken randomly, then the pod is bound.

solid answer

~50 s

kube-scheduler watches for pods whose `spec.nodeName` is empty and processes them one at a time from a priority-ordered queue. For each pod it runs a **scheduling cycle**: 1. **PreFilter** — validates the pod and precomputes state (e.g. total requests). 2. **Filter** (the old "predicates") — every node is tested by plugins such as `NodeResourcesFit`, `NodeAffinity`, `NodeName`, `TaintToleration`, `NodePorts`, `VolumeBinding`, `PodTopologySpread`, `InterPodAffinity`. Output: the set of *feasible* nodes. If it is empty, `PostFilter` may attempt preemption and the pod stays `Pending`. 3. **PreScore/Score** (the old "priorities") — each feasible node gets 0–100 from each score plugin (`NodeResourcesFit` least-allocated by default, `ImageLocality`, `PodTopologySpread`, `InterPodAffinity`, `NodeResourcesBalancedAllocation`); scores are normalised, multiplied by plugin weights and summed. Highest total wins; ties are broken randomly. 4. **Reserve/Permit** — the assignment is cached optimistically ("assume"). Then an asynchronous **binding cycle** (PreBind → Bind → PostBind) writes the Binding object; kubelet on that node sees the pod and starts it. In large clusters `percentageOfNodesToScore` lets filtering stop early once enough feasible nodes are found.

code

bash · 4 lines
bash
kubectl get pod api-7c9d -o wide
kubectl describe pod api-7c9d | sed -n '/Events/,$p'
kubectl get events --field-selector reason=FailedScheduling --sort-by=.lastTimestamp
kubectl get pod api-7c9d -o jsonpath='{.spec.schedulerName}{"\n"}'

go deeper

for a junior

Name the two phases in order — filter narrows to nodes that can run the pod, score ranks them — and know that kubelet, not the scheduler, starts the containers.

for a middle

Name concrete plugins on each side, explain that fit is computed from requests, mention weights and random tiebreaking, and note pods are handled one at a time.

for a senior

Add the reserve/permit/assume mechanics, the asynchronous binding cycle and kubelet re-admission, percentageOfNodesToScore in large clusters, and how profiles configure plugin weights.

for a principal

Discuss the placement policy the scoring configuration encodes — spread versus bin-pack, its interaction with autoscaling and cost — and when a custom profile or plugin is justified over expressing intent with standard constraints.

## The scheduler's job kube-scheduler is a control-plane component that watches the API server for **pods with an empty `spec.nodeName`**. For each one it picks a node and writes a `Binding`. It does not start containers — kubelet on the chosen node notices the pod is assigned to it and does that. Scheduling is therefore a *placement decision recorded in the API*, nothing more. Pods are pulled from a priority queue: an *active* queue ordered by pod priority (then by the time they became unschedulable), plus *backoff* and *unschedulable* queues that hold pods waiting to be retried when relevant cluster events occur. Pods are handled **one at a time** in the scheduling cycle, which is why the scheduler is a single-threaded decision maker with a concurrent tail. ## Phase 1 — Filter (predicates) The filter phase answers a yes/no question per node: *can this pod run here at all?* It is a hard, non-negotiable pass. Representative plugins: - **NodeResourcesFit** — do the pod's CPU/memory/ephemeral-storage/extended-resource **requests** fit in the node's allocatable minus the requests of pods already there? Note: *requests*, not actual usage. A node at 5% real CPU can still be "full". - **NodeUnschedulable** — skips nodes cordoned with `spec.unschedulable`. - **NodeName / NodeAffinity** — honours `spec.nodeName`, `nodeSelector`, and `requiredDuringSchedulingIgnoredDuringExecution` affinity. - **TaintToleration** — rejects nodes with `NoSchedule`/`NoExecute` taints the pod does not tolerate. - **NodePorts** — rejects nodes where a requested `hostPort` is already taken. - **VolumeBinding, VolumeRestrictions, node volume limits** — can the PVs be bound/attached here, and is the per-node attachment limit respected? - **PodTopologySpread, InterPodAffinity** — the *required* variants act as filters. The output is the **feasible node list**. If it is empty, the pod is unschedulable: `PostFilter` runs (preemption lives there), the pod goes back to the queue, and its event carries the aggregated reason — `0/40 nodes are available: 12 Insufficient cpu, 28 node(s) had untolerated taint {...}`. In large clusters the scheduler does not filter every node. `percentageOfNodesToScore` (default: adaptive, roughly 50% shrinking toward a 5% floor as the cluster grows) stops the sweep once *enough* feasible nodes are found, resuming from where it stopped last time so nodes are visited fairly. This trades a little placement quality for latency, and it means "the best node in the cluster" is not guaranteed — only the best among those examined. ## Phase 2 — Score (priorities) Every feasible node is scored 0–100 by each enabled score plugin; scores are normalised, multiplied by the plugin's **weight**, and summed. Common plugins: - **NodeResourcesFit** with a scoring strategy — `LeastAllocated` by default, which spreads pods toward emptier nodes; `MostAllocated` bin-packs; `RequestedToCapacityRatio` lets you shape the curve. - **NodeResourcesBalancedAllocation** — prefers nodes where CPU and memory utilisation end up balanced, avoiding a node with free memory but no CPU. - **ImageLocality** — favours nodes that already have the container image, cutting startup time. - **InterPodAffinity / NodeAffinity** — `preferredDuringScheduling…` terms contribute here rather than filtering. - **PodTopologySpread** — `ScheduleAnyway` constraints score here; it carries a higher default weight because spreading matters for availability. - **TaintToleration** — `PreferNoSchedule` taints reduce a node's score instead of excluding it. The highest total wins; equal totals are broken **randomly** (reservoir sampling), which is why identical pods do not all pile onto the same node. ## Reserve, Permit, and the binding cycle After picking a node the scheduler runs **Reserve** (plugins claim resources in their own caches, e.g. volume binding) and **Permit** (a plugin may approve, deny, or *wait* — this is how gang/co-scheduling extensions hold a pod until its peers are ready). It then **assumes** the pod is on that node in its internal cache so the next pod in the queue sees the capacity as taken, and hands off to the **binding cycle**, which runs **asynchronously** and in parallel with the next pod's scheduling cycle: `PreBind` (e.g. actually bind the PVC), `Bind` (write the Binding to the API), `PostBind`. That asynchrony is the source of a class of subtle behaviour. If binding fails — API error, a `PreBind` volume error — the scheduler runs **Unreserve**, drops the assumption and requeues the pod. And because the decision was made against a cached snapshot, the node's state can change in between; kubelet performs its own **admission check** when the pod arrives and may reject it with `OutOfcpu`/`OutOfmemory`, producing a failed pod on the node rather than a pending one. ## Extensibility All of the above is the **scheduling framework**: filter and score are just plugin extension points, configured by a `KubeSchedulerConfiguration` with named **profiles**. A pod selects a profile via `spec.schedulerName`, so one binary can serve several policies, and you can enable/disable plugins or change their weights per profile without forking the scheduler.

  • Does the filter phase use a node's actual CPU usage or the sum of pod requests?
    It uses requests. NodeResourcesFit compares the pod's requests against node allocatable minus the requests of pods already assigned there, so a node running at 5% real CPU can still be unschedulable if its pods have reserved everything. That is why pods with no requests schedule almost anywhere and then fight for real resources at runtime, and why right-sizing requests is the main lever on packing density.
  • Two nodes end up with the same total score. How does the scheduler choose?
    Randomly, using reservoir sampling over the nodes tied at the maximum. This is deliberate: deterministic tiebreaking would make a batch of identical pods all land on the same node. It also means placement is not reproducible run to run, so if you need a specific distribution you must express it with topology spread constraints or affinity rather than relying on scoring behaviour.
  • Why can a pod be bound to a node and then fail with OutOfcpu?
    The scheduling cycle decides against the scheduler's cached snapshot and the binding cycle is asynchronous, so the node's real state can move between the decision and the pod arriving. Kubelet runs its own admission check when it sees a pod assigned to it and rejects the pod if the resources are not actually available. The pod ends up failed on that node rather than pending, and its controller creates a replacement.

Hiring: the filter phase is the pass/fail requirements screen (must have the visa, must be in the country), the score phase is ranking the shortlist that survived — and only the offer letter (the Binding) makes it real.

saying these in an interview costs you the question

  • Saying the scheduler places pods based on live CPU/memory utilisation rather than requests
  • Believing filter and score are one combined ranking, or that a low-scoring node can still be chosen when a higher one exists
  • Claiming the scheduler starts the container — it only writes a Binding; kubelet starts it
  • Assuming every node is evaluated in a large cluster, ignoring percentageOfNodesToScore
  • Thinking scheduling is atomic with no window between decision and kubelet admission

context