In a Kubernetes pod spec, what is the difference between setting spec.nodeName and setting spec.nodeSelector, and what are the consequences of each?
answer
- Empty spec.nodeName = unscheduled; the scheduler's job is to fill it
- nodeSelector = hard label filter in the filter phase, AND semantics only
- nodeName = skip the scheduler; kubelet on that node admits or rejects
- nodeName failure = OutOfcpu Failed pod, not Pending
- Cordon/taints/spread/preemption all bypassed by nodeName
basics
~20 snodeSelector is a hard filter the scheduler applies: the pod goes to any node whose labels match. nodeName bypasses the scheduler entirely — kubelet on that exact node picks the pod up. If that node is full, missing or cordoned, the pod fails outright instead of waiting.
solid answer
~50 s**`nodeSelector`** is a map of label key/values that the scheduler enforces during the **filter phase**: only nodes carrying all those labels are feasible, and normal scoring then picks among them. The pod still benefits from everything else — resource fit, taints, spreading, retry while `Pending`. **`nodeName`** is the field the scheduler *writes* when it binds a pod. Setting it yourself means the pod is already "scheduled": kube-scheduler ignores it, no filters or scores run, and kubelet on that named node sees the pod and tries to admit it. The consequences are asymmetric: - If the named node lacks capacity, kubelet rejects the pod with `OutOfcpu`/`OutOfmemory` — it **fails** rather than waiting for room. - If the node does not exist, is cordoned, or is `NotReady`, the pod just sits with nothing happening; cordoning does not protect it, and taints are not evaluated. - You lose spreading, preemption and rescheduling. Use `nodeSelector` (or node affinity for richer expressions) for real workloads; reserve `nodeName` for debugging and bootstrap pods.
code
yaml · 22 linesapiVersion: v1
kind: Pod
metadata: {name: inference}
spec:
nodeSelector:
node.example.com/pool: gpu
kubernetes.io/arch: amd64
containers:
- name: app
image: registry.example.com/inference:3.1
resources:
requests: {cpu: "2", memory: 8Gi}
---
apiVersion: v1
kind: Pod
metadata: {name: debug-on-worker-07}
spec:
nodeName: worker-07
containers:
- name: shell
image: registry.example.com/netshoot:1.0
command: ["sleep", "3600"]go deeper
State the core difference clearly: nodeSelector lets the scheduler choose among labelled nodes, nodeName skips the scheduler and pins to one host.
Add the failure modes — OutOfcpu rejection by kubelet, taints and cordon bypassed — and that nodeSelector is AND-only equality with node affinity as the richer form.
Explain it in terms of the lifecycle (empty nodeName is what defines unscheduled), why pinning breaks under node replacement and autoscaling, and pair selectors with taints for dedicated pools.
Frame placement as expressing node properties rather than identities, so intent survives node churn, and treat any nodeName in production manifests as a defect to be designed out.
## Two different points in the lifecycle A pod is "unscheduled" precisely when `spec.nodeName` is empty. kube-scheduler watches for exactly that condition, chooses a node, and writes the name via a `Binding`. Kubelets, in turn, watch for pods whose `nodeName` equals their own node and run them. That single fact explains everything: - **`nodeSelector` participates in the choice.** It is input to the scheduler's filter phase, implemented by the `NodeAffinity`/node-selector filter plugin. A node is feasible only if it carries every listed label with the exact value. - **`nodeName` replaces the choice.** A pod created with it set is indistinguishable, to the scheduler, from a pod it already bound. It is skipped entirely. ## nodeSelector in practice ```yaml spec: nodeSelector: kubernetes.io/os: linux node.example.com/pool: gpu ``` Semantics: - **AND only.** Every pair must match. There is no OR, no `In` set, no `Exists`, no negation. For those you need `spec.affinity.nodeAffinity`, whose `requiredDuringSchedulingIgnoredDuringExecution` form is the general version of the same filter and additionally supports `preferredDuringScheduling…` soft terms that feed the score phase. - **Hard.** No matching node means the pod stays `Pending` with `didn't match Pod's node affinity/selector` in its event, and it is retried when the cluster changes — for example when a node gets the label or the Cluster Autoscaler adds one from a matching group. - **Only labels.** Common built-ins are `kubernetes.io/os`, `kubernetes.io/arch`, `topology.kubernetes.io/zone`, `node.kubernetes.io/instance-type`, plus whatever labels you or your node pools apply. - **Ignored during execution.** If the label is removed from a node later, running pods stay put; the constraint is only evaluated at scheduling time. - **Not isolation.** A selector says "put me there"; it does not stop other pods landing on the same node. Keeping others *off* is what **taints** do — the two are complementary, and dedicated pools normally use both. ## nodeName in practice ```yaml spec: nodeName: worker-07 ``` What you give up and what can go wrong: - **No resource fit check by the scheduler.** Kubelet does its own admission when the pod arrives; if the node cannot fit the requests the pod is rejected with status `OutOfcpu`, `OutOfmemory` or `OutOfpods`. Note the failure mode: the pod ends up **Failed on that node**, not `Pending` waiting for capacity. A controller will keep making replacements that keep failing. - **No taint evaluation, no affinity, no topology spread, no preemption.** None of those plugins run. - **Cordoning does not help.** `kubectl cordon` sets `spec.unschedulable`, which the `NodeUnschedulable` *filter* plugin reads — and no filters run for a `nodeName` pod. Draining still evicts it, but new ones keep arriving. - **A typo or a deleted node means silence.** No kubelet is watching for that name, so the pod sits indefinitely with no events explaining why. - **Node names are not stable** in autoscaled or immutable-infrastructure clusters, so anything pinned by name breaks on node replacement. Legitimate uses are narrow: an ad-hoc debug pod you need on one specific host, a manifest placed in kubelet's static-pod directory (those are inherently node-bound), or bootstrap components that must run before a scheduler exists. Even for debugging, `kubectl debug node/<name>` is usually the better tool. ## Choosing between them, and what else exists The ladder of increasing expressiveness: 1. `nodeSelector` — simple equality on labels; enough for most "run on the GPU pool" cases. 2. `nodeAffinity` required terms — same power plus `In`, `NotIn`, `Exists`, `Gt`, `Lt`, and multiple alternative term sets (OR). 3. `nodeAffinity` preferred terms — a *soft* wish contributing to the score, so the pod still schedules elsewhere when the preferred nodes are full. 4. `nodeName` — not a constraint at all, but an override; it is a way of saying "do not schedule this". A practical rule: if you can express the requirement in terms of a *property* of nodes (pool, zone, hardware, OS), use labels and a selector — it survives node replacement and cooperates with autoscaling. Reach for a node's identity only when the identity itself is the point.
- What happens to a pod with spec.nodeName set to a node that does not have enough free CPU?The scheduler never sees it, so no fit check happens there. Kubelet on that node runs its own admission check when the pod appears, finds the requests do not fit, and rejects it — the pod goes to a Failed state with reason OutOfcpu. Unlike a normal unschedulable pod it does not wait in Pending for capacity, and a controller will keep creating replacements that fail the same way.
- When would you use node affinity instead of nodeSelector?When the rule needs more than equality — set membership with In/NotIn, existence checks, numeric Gt/Lt comparisons, or alternative groups of terms that behave like OR. Node affinity also offers preferredDuringSchedulingIgnoredDuringExecution, a soft form that contributes to the score phase so the pod still lands somewhere when the preferred nodes are full, which plain nodeSelector cannot express.
nodeSelector is booking 'a window seat in economy'; nodeName is walking onto the plane and sitting in seat 14A — if it's taken, you are not rebooked, you are removed.
saying these in an interview costs you the question
- Thinking nodeName is just a stricter nodeSelector that still goes through the scheduler
- Believing a cordoned node will refuse a pod that has nodeName set
- Expecting a nodeName pod to wait in Pending for capacity instead of failing with OutOfcpu
- Treating nodeSelector as isolation that keeps other pods off the node — that requires taints
- Assuming nodeSelector supports OR, negation or set operators