In a Kubernetes cluster, what does the kube-scheduler actually do when a new Pod is created, and what does it not do?
answer
- empty spec.nodeName → scheduler's inbox
- filter, score, bind
- binding is the only write
- kubelet watches its own nodeName
- Pending + FailedScheduling event
basics
~20 skube-scheduler watches for Pods with an empty spec.nodeName, picks a suitable node, and writes that choice back to the API server (a binding). It does not start the container — the kubelet on the chosen node does that.
solid answer
~50 sWhen you create a Pod (usually indirectly, via a Deployment), the object lands in etcd with `spec.nodeName` empty. **kube-scheduler** watches the API server for such unscheduled Pods. For each one it filters nodes that cannot host the Pod (insufficient CPU/memory, unmatched node selector, untolerated taint, unavailable ports/volumes) and scores the survivors, then picks the best. Its only *action* is writing a **Binding** — effectively setting `spec.nodeName` — through the API server. After that the scheduler is done. The **kubelet** on the named node sees a Pod bound to itself, pulls images and asks the container runtime to start containers, then reports status back. So the scheduler is a placement decision-maker, not an executor, and it never talks to nodes directly. If no node fits, the Pod stays `Pending` with a `FailedScheduling` event, which is why `kubectl describe pod` is the first debugging step.
go deeper
Know the one-liner: scheduler picks the node and writes it down; kubelet runs the container. Mention Pending Pods and kubectl describe pod.
Add the filter/score/bind sequence, that decisions use requests not usage, and that spec.nodeName bypasses scheduling entirely.
Discuss failure behaviour (running Pods unaffected), no automatic rebalancing, preemption as the only scheduler-driven deletion, and how you debug chronic Pending.
Frame it as the placement/execution split that keeps the control plane restartable and idempotent, and reason about when a second or custom scheduler (via schedulerName) is worth operating.
## The problem the scheduler solves Kubernetes separates *what should run* from *where it runs*. When something creates a Pod, the Pod object is persisted with `spec.nodeName` unset. Nothing yet knows which machine will host it. Assigning that machine is the scheduler's entire job. ## Where it sits kube-scheduler is a control-plane process (a static Pod on control-plane nodes in most installs). Like every other control-plane component it talks **only** to kube-apiserver — never to kubelets, never directly to etcd. It maintains a watch on Pods and on cluster state (nodes, other Pods' resource requests, PersistentVolumes, CSI limits) so it can decide without a round trip per query. ## The loop 1. A Pod appears with `spec.nodeName == ""` and `spec.schedulerName` matching this scheduler (default `default-scheduler`). 2. It enters the scheduling queue, ordered by priority. 3. **Filtering**: nodes that cannot possibly host the Pod are removed — not enough allocatable CPU/memory for the Pod's *requests*, node selector or affinity mismatch, a taint the Pod does not tolerate, a required hostPort already in use, a volume that cannot attach in that zone. 4. **Scoring**: remaining feasible nodes get a numeric score (spread across nodes, image locality, affinity preferences) and the highest wins, with ties broken randomly. 5. **Binding**: the scheduler POSTs a Binding (or updates `spec.nodeName`) to the API server. (The internals of filter and score plugins are a separate subject; what matters architecturally is that the output is one field written through the API.) ## What happens next — and why the split matters Every kubelet watches Pods filtered by `spec.nodeName == <its own name>`. The moment the binding is written, the kubelet on that node sees the Pod, creates the sandbox, pulls images, starts containers, runs probes, and patches `status`. Two consequences follow: - **The scheduler is not on the critical path after binding.** If kube-scheduler crashes, already-bound Pods keep running and restarting; only *new* Pods stop being placed. This is the classic interview follow-up. - **You can bypass it.** Setting `spec.nodeName` yourself in the manifest makes the Pod skip scheduling entirely — the kubelet picks it up directly. That is how DaemonSet Pods historically worked and how you can pin a debug Pod. It also means the resource checks the scheduler would have done never happen, so you can overcommit a node. ## Scheduling is a decision, not a guarantee The scheduler decides based on *requests*, not actual usage, and based on a snapshot of state. Between decision and start, the node may fill up; the kubelet can then reject the Pod (status `OutOfcpu`/`OutOfmemory`) and it is rescheduled. Conversely the scheduler never moves a running Pod: there is no rebalancing. If a node becomes crowded or a better node appears later, nothing migrates. Rescheduling only happens because a Pod died and its controller created a *new* Pod object, which is then scheduled fresh. Eviction/preemption of lower-priority Pods to make room is the one case where the scheduler causes deletion, and even then it deletes victims and retries rather than moving anyone. ## Pending Pods When no node passes filtering, the Pod stays `Pending`. The scheduler emits a `FailedScheduling` event listing why each node was rejected — "insufficient memory", "node(s) had untolerated taint", "didn't match Pod's node affinity". Reading that message is the fastest diagnosis path, and it is the standard reason a Pod sits Pending forever while the cluster looks healthy. ## Relationship to the controller manager Controllers create Pod *objects* to satisfy a desired count; the scheduler places them; kubelets run them. Three separate loops, each reading and writing the same API objects, each idempotent and restartable. That decomposition — no component commands another, all coordinate through declarative state — is the point of the design.
- What happens to the cluster if kube-scheduler is down for an hour?Everything already running keeps running, because kubelets act on bindings that were already written and restart containers locally. What stops is placement: newly created Pods pile up in Pending with no node assigned, so rollouts, scale-ups and replacements for Pods on a failed node stall. The moment the scheduler returns it drains the backlog, since it works from current API state rather than a queue it lost.
- How can you get a Pod onto a node without the scheduler?Set `spec.nodeName` explicitly in the manifest. The Pod is then already bound, so the scheduler ignores it and the named kubelet picks it up directly. The cost is that none of the scheduler's checks run — no resource fit, no taint or affinity evaluation — so you can push a node past its allocatable capacity and trigger evictions.
Like a seating host at a restaurant: they assign your table and write it on the sheet, but the kitchen and waiter — not the host — actually serve you.
saying these in an interview costs you the question
- Saying the scheduler starts containers or talks to the node — it only writes a binding; the kubelet starts containers.
- Claiming the scheduler continuously rebalances running Pods across nodes; it never moves a running Pod.
- Thinking scheduling uses actual CPU/memory usage rather than the Pod's declared requests.
- Believing a scheduler outage stops running workloads.