skip to content

Scaling and Scheduling

Where pods land and how capacity follows: the scheduler's filter-then-score cycle, placement APIs like affinity and taints, node Allocatable and bin packing, and the HPA, VPA and node autoscalers. Requests decide placement while real usage decides the bill.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

explore

questions

page 1 of 2

What does node affinity do for a Kubernetes Pod, and how does it differ from the simpler nodeSelector field in the Pod spec?

level: juniorimportance: must knowfreq 62%

answer

  1. nodeSelector = exact match, AND-ed, mandatory
  2. affinity = matchExpressions: In / NotIn / Exists / Gt / Lt
  3. required = filter phase; preferred = score phase, weight 1-100
  4. nodeSelectorTerms OR-ed, matchExpressions AND-ed
  5. IgnoredDuringExecution = never evicts

basics

~20 s

Both steer a Pod onto nodes carrying particular labels. nodeSelector is exact key=value matching and every entry must match. Node affinity adds an expression language (In, NotIn, Exists, Gt, Lt), OR-ed alternatives, and a soft preferred form the scheduler may ignore.

solid answer

~50 s

Node affinity is a Pod-side placement rule written against **node labels**. `spec.nodeSelector` is the primitive form: a flat map of exact key=value pairs, all of which must match, with no alternatives and no fallback. `spec.affinity.nodeAffinity` has two fields: - `requiredDuringSchedulingIgnoredDuringExecution` — a hard filter. It holds `nodeSelectorTerms`; terms are OR-ed, and the `matchExpressions` inside one term are AND-ed. Operators are `In`, `NotIn`, `Exists`, `DoesNotExist`, `Gt`, `Lt`, so you get sets and negation. - `preferredDuringSchedulingIgnoredDuringExecution` — soft. A list of `{weight: 1-100, preference}`; matching nodes gain that weight during scoring, and the Pod still schedules if nothing matches. `IgnoredDuringExecution` in both names means the rule is enforced only at placement time — relabelling a node afterwards never evicts the running Pod. Use nodeSelector for one trivial exact match; use node affinity when you need a set of acceptable values, a negation, or a preference with a fallback such as "prefer the spot pool, accept on-demand".

code

yaml · 29 lines
yaml
apiVersion: v1
kind: Pod
metadata:
  name: worker
spec:
  affinity:
    nodeAffinity:
      requiredDuringSchedulingIgnoredDuringExecution:
        nodeSelectorTerms:
        - matchExpressions:
          - key: kubernetes.io/arch
            operator: In
            values: ["amd64", "arm64"]
      preferredDuringSchedulingIgnoredDuringExecution:
      - weight: 100
        preference:
          matchExpressions:
          - key: pool
            operator: In
            values: ["spot"]
      - weight: 1
        preference:
          matchExpressions:
          - key: pool
            operator: In
            values: ["on-demand"]
  containers:
  - name: app
    image: registry.example.com/app:1.4.0

go deeper

for a junior

Know that both match node labels, that nodeSelector is exact-match only, and that affinity has a hard required form and a soft preferred form.

for a middle

Add the operator list, the OR/AND semantics of terms versus expressions, weights 1-100 in the scoring phase, and what IgnoredDuringExecution really means.

for a senior

Talk about diagnosing FailedScheduling events, the prefer-spot-fallback-on-demand pattern, and making sure node-group templates carry the labels so autoscaling can react.

for a principal

Frame it as a labelling contract: which labels are platform-owned versus team-owned, whether hard rules should be allowed at all, and how placement policy is enforced through admission rather than left to each Deployment.

## The problem The Kubernetes scheduler assigns every new Pod to a node. Left alone it picks any node with enough free CPU and memory. That is often not good enough: a Pod needs a GPU, must run on an ARM machine, or should land on a cheap spot-instance pool. Node labels plus node affinity are how you express that from the Pod side. ## Node labels Every Node object carries labels — key/value strings. The kubelet and cloud provider populate well-known ones: `kubernetes.io/os`, `kubernetes.io/arch`, `node.kubernetes.io/instance-type`, `topology.kubernetes.io/zone`, `topology.kubernetes.io/region`, `kubernetes.io/hostname`. Operators add their own, e.g. `pool=spot` or `hardware=gpu`. Placement rules are always written against labels, never node names. ## nodeSelector — the simple form `spec.nodeSelector` is a flat map. `{disktype: ssd}` means "only nodes whose `disktype` label equals `ssd`". Multiple entries are AND-ed. That is the entire feature: exact string equality, always mandatory, no OR, no negation, no preference. If no node matches, the Pod stays `Pending` forever. ## nodeAffinity — the expressive form Under `spec.affinity.nodeAffinity` there are two fields. **requiredDuringSchedulingIgnoredDuringExecution** is a hard rule evaluated in the scheduler's filter phase. It contains `nodeSelectorTerms`, a list. The list is OR-ed; the `matchExpressions` inside a single term are AND-ed. So you can express "(zone in [eu-west-1a, eu-west-1b]) OR (pool = fallback)". **preferredDuringSchedulingIgnoredDuringExecution** is a soft rule evaluated in the scoring phase. It is a list of `{weight, preference}` where weight is 1–100 and `preference` holds `matchExpressions`. Every node matching a preference gets that weight added to its score; the highest-scoring feasible node wins. Nothing matching is fine — the Pod still schedules. Each expression is `{key, operator, values}`. Operators: `In`, `NotIn`, `Exists`, `DoesNotExist`, `Gt`, `Lt`. `NotIn` and `DoesNotExist` give you node-level anti-affinity. `Gt`/`Lt` compare a single integer-valued label. There is also `matchFields` for a small set of non-label fields such as `metadata.name`. ## Reading the long field names Split each name in two. The first half — `requiredDuringScheduling` or `preferredDuringScheduling` — says when the rule applies: at placement. The second half — `IgnoredDuringExecution` — says what happens afterwards: nothing. If someone relabels the node so the Pod no longer matches, the running Pod stays put. A `RequiredDuringExecution` variant that would evict was proposed and never shipped; taints with the `NoExecute` effect are the mechanism that actually evicts running Pods, and that is a separate feature working from the node side. ## Choosing between them nodeSelector is fine and very readable when the rule is one exact match you never want a fallback for. Reach for node affinity when you need a set of acceptable values, a negation, an OR of alternatives, or — the most valuable case — a soft preference. A common production shape is: prefer `pool=spot` with weight 100, prefer `pool=on-demand` with weight 1, require neither. Workloads then drift onto cheap capacity but never go Pending when spot capacity disappears. The two can be combined on one Pod; both must be satisfied. Required node affinity and nodeSelector together are AND-ed. ## Failure mode and diagnosis A hard rule no node satisfies leaves the Pod `Pending`. `kubectl describe pod` shows a `FailedScheduling` event reading roughly "0/12 nodes are available: 12 node(s) didn't match Pod's node affinity/selector". The usual causes are a typo in the label key, a label applied to the wrong nodes, or a value written as an integer in YAML where the label is a string. ## Cost Node affinity is cheap: one pass over candidate nodes comparing label maps, both in filtering and scoring. It costs nothing like inter-pod affinity, which must inspect the Pods already placed across whole topology domains. ## Interaction with cluster autoscaling Cluster Autoscaler simulates node-group templates, and those templates carry labels, so a Pending Pod requiring `pool=gpu` can trigger a scale-up of that pool — provided the node group's template genuinely advertises the label. Labels bolted on by a boot script after the node registers are invisible to that simulation, and the Pod stays Pending.

  • If someone removes the label a running Pod's required node affinity depends on, what happens to that Pod?
    Nothing — it keeps running. All current affinity rules end in IgnoredDuringExecution, so they are only evaluated when the scheduler places the Pod. The rule would only bite again if the Pod were recreated, at which point it might not schedule at all. To actively evict Pods from a node you need a taint with the NoExecute effect, which is a different mechanism.
  • Inside nodeAffinity, how do multiple nodeSelectorTerms combine, and how do multiple matchExpressions inside one term combine?
    The list of nodeSelectorTerms is OR-ed: a node is feasible if it satisfies any one term. Within a single term, all matchExpressions must hold, so they are AND-ed. That asymmetry is the usual source of bugs — putting two expressions in separate terms accidentally loosens the rule instead of tightening it.

nodeSelector is a job ad saying "must live in Berlin". Node affinity is "must live in Germany or Austria (hard), Berlin preferred (weight 100), Munich acceptable (weight 10)" — and once you are hired, moving house does not get you fired.

saying these in an interview costs you the question

  • Saying node affinity evicts Pods when node labels change — every current form is IgnoredDuringExecution.
  • Claiming preferred affinity is a guarantee; it only adds weight during scoring and can be ignored entirely.
  • Confusing node affinity with taints and tolerations — affinity attracts a Pod to nodes, tolerations only permit a Pod onto tainted nodes and never attract anything.
  • Believing matchExpressions across separate nodeSelectorTerms are AND-ed (they are OR-ed).
  • Thinking nodeSelector is deprecated — it is not, it is simply the limited form.

context

open as a page

In Kubernetes, how does a pod request a GPU, and why can it request 0.35 CPU but not 0.35 GPU?

level: juniorimportance: must knowfreq 68%

basics

~20 s

A pod requests a GPU as an extended resource, such as nvidia.com/gpu: 1, in its container limits. Extended resources are whole units the scheduler counts, not shares. So Kubernetes rejects fractions, requires the request to equal the limit, and never overcommits them.

open as a page

What does a Kubernetes HorizontalPodAutoscaler object do, what must be true of the workload for it to work, and what does it deliberately not do?

level: juniorimportance: must knowfreq 70%

basics

~20 s

It watches a metric (usually CPU) for a Deployment or StatefulSet and rewrites that workload's replica count, clamped between minReplicas and maxReplicas, to keep the metric near a target. Needs metrics-server plus resource requests on the pods. It adds and removes pods only — it never resizes a pod and never adds nodes.

open as a page

What is KEDA in a Kubernetes cluster, and what does a KEDA ScaledObject declare about the workload it scales?

level: juniorimportance: must knowfreq 62%

basics

~20 s

KEDA is a Kubernetes operator that scales workloads on external event signals such as Kafka lag, queue depth or a Prometheus query. A ScaledObject names the target workload, its replica bounds and its triggers. KEDA turns that into an HPA it owns.

open as a page

In a Kubernetes cluster with a spot node pool, which workloads belong on those nodes, and how do you keep everything else off them?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Spot nodes suit replicated, stateless or checkpointing work that can lose a node at short notice. Label the pool by capacity type and taint it, so only pods that tolerate the taint and select the label land there.

open as a page

In Kubernetes, what is a node taint and what is a pod toleration, and what do the NoSchedule, PreferNoSchedule and NoExecute taint effects each do?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A taint marks a node so it repels pods; a matching toleration on a pod lets that pod ignore the taint. NoSchedule blocks new pods, PreferNoSchedule is a soft preference to schedule elsewhere, and NoExecute also evicts already-running pods that do not tolerate it.

open as a page

How would you use Kubernetes podAntiAffinity to keep the replicas of a Deployment off the same machine, and what exactly does the topologyKey field control?

level: middleimportance: must knowfreq 58%

basics

~20 s

Add podAntiAffinity whose labelSelector matches the Deployment's own Pod labels, with topologyKey kubernetes.io/hostname. topologyKey names a node label; nodes sharing a value form one domain, and the scheduler refuses to place a second matching Pod into a domain that already has one.

open as a page

In Kubernetes, why does a node's Allocatable differ from its Capacity, and how does the kubelet compute Allocatable?

level: middleimportance: must knowfreq 62%

basics

~20 s

Capacity is what the machine has; Allocatable is what the kubelet offers to pods: Capacity minus kube-reserved, minus system-reserved, minus the hard eviction threshold. The scheduler fits pod requests against Allocatable, never against Capacity or live usage.

open as a page

How does the Kubernetes Cluster Autoscaler decide to add a node, and what conditions must hold before it removes one?

level: middleimportance: must knowfreq 56%

basics

~20 s

Scale-up: it watches for Pending Pods the scheduler could not place, simulates whether a new node from some node group would fit them, and grows that group. Scale-down: a node whose requested resources stay under a threshold for a sustained period, and all of whose Pods can move elsewhere, is cordoned, drained, and deleted.

open as a page

A Kubernetes HorizontalPodAutoscaler is set to an average CPU utilization target of 50%. Walk through exactly how the controller turns the observed metric into a new replica count.

level: middleimportance: must knowfreq 72%

basics

~20 s

desiredReplicas = ceil(currentReplicas × currentMetric ÷ targetMetric). currentMetric is the mean across ready pods of (usage ÷ request) as a percentage. If the ratio is within a tolerance of 1 (10% by default) nothing happens, and the result is clamped to minReplicas/maxReplicas.

open as a page

How does a KEDA ScaledObject take a Kubernetes Deployment to zero replicas and back, and which fields govern each transition?

level: middleimportance: must knowfreq 58%

basics

~20 s

KEDA's own loop handles the zero boundary. It activates the target to one replica when a trigger exceeds its activation threshold, and returns it to zero after cooldownPeriod with every trigger inactive. The generated HPA scales between one and maxReplicaCount.

open as a page

What is a Kubernetes PodDisruptionBudget, and which kinds of pod loss does it actually protect against?

level: middleimportance: must knowfreq 60%

basics

~20 s

A PodDisruptionBudget declares the minimum availability a set of pods must keep during voluntary disruptions — evictions from node drains, cluster-autoscaler scale-down, upgrades. The API server rejects an eviction that would break it. It cannot stop involuntary loss: node crashes, kernel panics, kubelet node-pressure eviction, or a direct pod delete.

open as a page

In Kubernetes, what does a PriorityClass object do, how does a pod end up with a priority value, and what is the globalDefault field for?

level: middleimportance: must knowfreq 58%

basics

~20 s

A PriorityClass is a cluster-scoped object mapping a name to an integer value. A pod names one in spec.priorityClassName; admission copies the number into spec.priority. Higher priority means earlier scheduling attempts and the right to preempt lower-priority pods. globalDefault marks the one class applied to pods that name none.

open as a page

Walk through how kube-scheduler decides which node an unscheduled pod lands on, naming the phases it goes through and what each one produces.

level: middleimportance: must knowfreq 66%

basics

~20 s

The scheduler watches for pods with no node assigned, then runs two phases: filter, which reduces all nodes to the feasible ones (fit, taints, affinity, volumes), and score, which ranks the survivors 0-100 by weighted plugins. The best-scoring node wins, ties broken randomly, then the pod is bound.

open as a page

When a cloud provider signals that a Kubernetes spot node will be reclaimed, what should happen before the machine disappears, and what makes it happen?

level: middleimportance: must knowfreq 58%

basics

~20 s

An interruption handler catches the provider's notice, cordons the node and evicts its pods through the Eviction API, so replacements schedule elsewhere. Each pod's shutdown must fit inside the notice, or it is killed when the machine goes.

open as a page

You need a set of Kubernetes nodes reserved exclusively for one team's workload, so that nothing else lands there and that team's pods always run there. Describe exactly how you would configure it and why a toleration alone is not enough.

level: middleimportance: must knowfreq 54%

basics

~20 s

Taint the nodes (for example dedicated=team-a:NoSchedule) so untolerating pods are kept out, label them (pool=team-a), then give the team's pods both a matching toleration and a nodeSelector or node affinity for that label. The taint provides exclusivity; the selector provides attraction.

open as a page

In a Kubernetes pod spec, what does the `tolerationSeconds` field on a toleration do, and how does it interact with a taint whose effect is NoExecute?

level: middleimportance: must knowfreq 58%

basics

~20 s

tolerationSeconds makes a toleration temporary: with a NoExecute taint the pod is allowed to keep running for that many seconds after the taint appears, then it is evicted. Omitted means tolerate forever; zero or negative means evict immediately. It is ignored for other effects.

open as a page

What do Kubernetes `topologySpreadConstraints` in a pod spec do, and what do the `maxSkew`, `topologyKey`, `labelSelector` and `whenUnsatisfiable` fields each mean?

level: middleimportance: must knowfreq 52%

basics

~20 s

They tell the scheduler to spread matching pods evenly across failure domains. topologyKey is the node label defining a domain (zone, hostname); labelSelector picks which pods are counted; maxSkew is the largest allowed difference in count between domains; whenUnsatisfiable says whether to block scheduling or place anyway.

open as a page

In a Kubernetes topology spread constraint, what is the difference between `whenUnsatisfiable: DoNotSchedule` and `whenUnsatisfiable: ScheduleAnyway`, and how do you decide which to use?

level: middleimportance: must knowfreq 46%

basics

~20 s

DoNotSchedule is a hard filter: if no domain can take the pod without exceeding maxSkew, the pod stays Pending. ScheduleAnyway is a scoring preference: the scheduler favours the emptiest domain but always places the pod. Choose hard when imbalance is worse than unavailability, soft otherwise.

open as a page

A Kubernetes HorizontalPodAutoscaler is not scaling and `kubectl get hpa` prints `<unknown>` in the TARGETS column. How do you diagnose it, and what are the usual root causes?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Check whether metrics exist at all with kubectl top pods. If that fails, metrics-server is missing or unhealthy (often TLS/kubelet-certificate or APIService issues). If top works but the HPA is unknown, the pods lack a resource request for the metric, the selector matches no ready pods, or a custom-metrics adapter is down. kubectl describe hpa names the failing condition.

open as a page

Cluster Autoscaler grew a 140-node Kubernetes cluster for a burst, but a week later nodes sit mostly idle and none are removed. How do you find what blocks scale-down?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Check that scale-down is running, then that each node is really under the 0.5 threshold, which is calculated from requests rather than usage. Then find the pod pinning it: local storage, unbudgeted kube-system pods, bare pods, safe-to-evict false, or a zero-disruption PDB.

open as a page

Explain the mechanism by which `kubectl drain` and a PodDisruptionBudget interact — which API is involved, what response signals a blocked eviction, and which pod removals bypass the whole mechanism.

level: seniorimportance: must knowfreq 46%

basics

~20 s

drain cordons the node, then calls the pods/eviction subresource for each pod. The disruption controller checks the matching PDB and returns HTTP 429 TooManyRequests when disruptionsAllowed is zero; drain retries until it succeeds or times out. Direct pod deletion, kubelet node-pressure eviction, and scheduler preemption all bypass eviction and ignore PDBs.

open as a page

A Kubernetes pod cannot be scheduled because no node has room for it. Walk through, step by step, what kube-scheduler's preemption logic does and what happens to the pods it removes.

level: seniorimportance: must knowfreq 52%

basics

~20 s

After all nodes fail filtering, the scheduler's PostFilter preemption plugin looks for one node where deleting some strictly-lower-priority pods would let the pending pod fit. It picks the least disruptive victim set, sets nominatedNodeName on the pending pod, and deletes victims gracefully. The freed node is not guaranteed to the pod.

open as a page

A Kubernetes pod stays Pending with the event '0/40 nodes are available: 12 Insufficient cpu, 28 node(s) had untolerated taint'. How do you read that message and drive the problem to resolution?

level: seniorimportance: must knowfreq 58%

basics

~20 s

That line is the filter phase's aggregated verdict: each clause is one plugin's rejection count, and the counts sum to every node, so no node is feasible. Fix the dominant clause — right-size requests or add capacity for the fit failures, add a toleration or untaint for the rest — then let the pod be retried.

open as a page

In a Kubernetes pod spec, what is the difference between setting spec.nodeName and setting spec.nodeSelector, and what are the consequences of each?

level: juniorimportance: should knowfreq 44%

basics

~20 s

nodeSelector is a hard filter the scheduler applies: the pod goes to any node whose labels match. nodeName bypasses the scheduler entirely — kubelet on that exact node picks the pod up. If that node is full, missing or cordoned, the pod fails outright instead of waiting.

open as a page

What are the update modes of the Kubernetes Vertical Pod Autoscaler, and what does each one actually do to a running Pod?

level: middleimportance: should knowfreq 44%

basics

~20 s

Off only publishes recommendations for humans to read. Initial applies them when a Pod is created and never again. Auto (and Recreate) additionally evicts running Pods so they are recreated with new resource requests. VPA changes requests per Pod; it never changes replica count.

open as a page

How does a Kubernetes device plugin make a node's GPUs visible to kube-scheduler and wire an allocated GPU into a container?

level: middleimportance: should knowfreq 48%

basics

~20 s

The plugin registers with the kubelet over a gRPC socket and streams its healthy devices through ListAndWatch. The kubelet publishes the count in Node allocatable for the scheduler, then calls Allocate to get the device files, mounts and environment variables to inject.

open as a page

What does the `cluster-autoscaler.kubernetes.io/safe-to-evict` annotation on a Kubernetes pod do, and when would you set it to `true` or `false`?

level: middleimportance: should knowfreq 44%

basics

~10 s

It is a per-pod override for Cluster Autoscaler's scale-down check: true lets the autoscaler remove the pod's node even if the pod would normally block it, and false pins the node against scale-down.

open as a page

A Kubernetes PodDisruptionBudget accepts either minAvailable or maxUnavailable. How do the two differ in practice, how are percentages resolved, and when does each choice go wrong?

level: middleimportance: should knowfreq 50%

basics

~20 s

minAvailable sets an absolute floor of healthy pods; maxUnavailable sets how many may be missing relative to the expected replica count. Percentages are resolved against the controller's replica count and round up. minAvailable as an integer is dangerous when replicas shrink; maxUnavailable: 0 and minAvailable equal to the replica count both block every drain.

open as a page

What does setting preemptionPolicy: Never on a Kubernetes PriorityClass change about how pods using that class are scheduled, and when would you want it?

level: middleimportance: should knowfreq 34%

basics

~20 s

Pods in that class keep their high priority for queue ordering — they get scheduling attempts before lower-priority pods — but they never evict anyone. If nothing fits they stay Pending. Use it for important-but-not-urgent work that should jump the queue without disrupting running workloads.

open as a page

showing 1–30 of 59