skip to content

Scaling and Scheduling

Where pods land and how capacity follows: the scheduler's filter-then-score cycle, placement APIs like affinity and taints, node Allocatable and bin packing, and the HPA, VPA and node autoscalers. Requests decide placement while real usage decides the bill.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

explore

questions

page 2 of 2

Why does the Kubernetes documentation warn against inter-Pod affinity and anti-affinity in large clusters, and how would you keep placement rules from slowing scheduling down?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Inter-Pod rules are evaluated per candidate node against every already-placed matching Pod and its node's topology label, so cost scales roughly with nodes times matching Pods, unlike node affinity's single label comparison. Limit selectors, prefer hostname topology, and prefer soft rules.

open as a page

On a Kubernetes cluster, why can a pod requesting 2.6 GiB of memory stay unplaceable while free allocatable memory across all nodes totals 11.3 GiB, and how do you reduce that stranded capacity?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Pods must fit whole on one node, so free capacity summed across nodes means nothing. Many small leftover gaps, or free CPU stuck next to exhausted memory, strand capacity. Reduce it by repacking pods, standardising request shapes and scoring toward fuller nodes.

open as a page

Why is running the Kubernetes Vertical Pod Autoscaler in Auto mode alongside a Horizontal Pod Autoscaler that targets CPU utilization considered unsafe, and what can you do instead?

level: seniorimportance: should knowfreq 40%

basics

~20 s

CPU-based horizontal scaling measures usage divided by the request. VPA moves the request, so the same real load changes the measured utilization and makes the horizontal controller scale in or out for no workload reason. The two form a feedback loop. Split them: VPA on memory only, horizontal scaling on CPU or on custom metrics.

open as a page

In a 38-node Kubernetes cluster, CPU-only pods keep landing on the six GPU nodes while a video-transcoding pod waits Pending for a GPU. How would you fence the GPU pool, and what does each mechanism guarantee?

level: seniorimportance: should knowfreq 45%

basics

~10 s

Taint the GPU nodes (nvidia.com/gpu:NoSchedule) to keep CPU-only pods off, and label them so GPU pods can select them. Enabling the ExtendedResourceToleration admission plugin adds the matching toleration to any pod that requests nvidia.com/gpu.

open as a page

A Kubernetes deployment's replica count oscillates all day — up 20, down 8, up 15. Which HorizontalPodAutoscaler settings control that, and how do the scaleUp and scaleDown stabilization windows actually work?

level: seniorimportance: should knowfreq 44%

basics

~20 s

The behavior block. behavior.scaleDown.stabilizationWindowSeconds (default 300) makes the controller take the highest recommendation from the last window before shrinking; the scaleUp window (default 0) takes the lowest. Policies under each direction cap how many pods or what percentage may change per period. Widen the scale-down window and add a percent/pod policy to damp flapping.

open as a page

CPU is a poor proxy for load on a queue-consuming service. How do you drive a Kubernetes HorizontalPodAutoscaler from a non-CPU signal such as queue depth or requests per second, and what does that require in the cluster?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Use the autoscaling/v2 metric types Pods, Object or External instead of Resource. Those are served by custom.metrics.k8s.io and external.metrics.k8s.io, which core Kubernetes does not implement — you install an adapter (Prometheus Adapter, KEDA, a cloud provider's) that translates a query into that API. Use External with AverageValue for queue depth so the target means backlog per pod.

open as a page

A KEDA-scaled fraud-rules engine on Kubernetes sits at zero replicas, and after a Kafka burst the first event waits 47 seconds before processing. Where does that time go, and what would you change?

level: seniorimportance: should knowfreq 38%

basics

~20 s

The delay is a chain: up to one pollingInterval before KEDA notices, then scheduling, start-up and readiness, then the consumer joining its group. Measure each stage; for a tight budget, keep minReplicaCount at 1 instead.

open as a page

Karpenter keeps replacing nodes under a latency-sensitive Kubernetes workload. How do NodePool consolidation and drift work, and how do you limit the disruption they cause?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Consolidation deletes or replaces nodes whose pods fit elsewhere more cheaply, and drift replaces nodes that no longer match their NodePool or NodeClass. Limit both with consolidateAfter, disruption budgets scoped by reason and schedule, PDBs, and the do-not-disrupt annotation.

open as a page

A node upgrade is stuck: the drain has been retrying for an hour with "Cannot evict pod as it would violate the pod's disruption budget". Walk through your diagnosis and the options for unblocking it.

level: seniorimportance: should knowfreq 42%

basics

~20 s

Find which PDB has disruptionsAllowed = 0 with kubectl get pdb -A, then check why: replicas equal to minAvailable, maxUnavailable 0, pods failing readiness so they never count as healthy, a selector spanning several workloads, or overlapping PDBs. Fix the cause — scale up, correct the budget, fix readiness — rather than deleting the PDB.

open as a page

Kubernetes ships two built-in PriorityClasses named system-cluster-critical and system-node-critical. What are they for, how do they differ, and what rules should govern their use?

level: seniorimportance: should knowfreq 30%

basics

~20 s

They are reserved classes for infrastructure: system-cluster-critical (2000000000) for control-plane-level add-ons like CoreDNS, and system-node-critical (2000001000, the highest) for per-node agents like the CNI and kube-proxy. Both sit above the one-billion user ceiling. Never put application workloads in them.

open as a page

A payments-authorization API on Kubernetes runs 7 replicas, 5 on spot nodes. One reclaim wave killed 4 at once and cut its 13-minute node drain short. How would you redesign its placement?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Guarantee an on-demand floor that can serve peak alone, let only the surplus replicas use spot, spread both across zones, and cut shutdown time to fit the reclaim notice, because a disruption budget cannot stop a reclaim.

open as a page

Kubernetes automatically adds taints such as `node.kubernetes.io/not-ready`, `node.kubernetes.io/unreachable` and `node.kubernetes.io/disk-pressure` to nodes. Explain what adds them, what they do to running pods, and how you would use them when operating a cluster.

level: seniorimportance: should knowfreq 42%

basics

~20 s

The node lifecycle controller taints nodes from their reported conditions. not-ready and unreachable are NoExecute, so pods are evicted once their injected 300-second toleration expires; pressure taints such as disk-pressure and memory-pressure are NoSchedule and only keep new pods away. DaemonSets tolerate them so node agents keep running.

open as a page

A Kubernetes Deployment declares a topology spread constraint with `maxSkew: 1` across availability zones, yet after a rollout its pods sit 5/2/1 across three zones. Give the likely causes and how you would investigate.

level: seniorimportance: should knowfreq 36%

basics

~20 s

Most likely: the constraint is ScheduleAnyway so it only scores; or the domains were unequal at placement time (a zone was full or unavailable) and nothing rebalances afterwards; or nodes lack the zone label and are excluded; or the labelSelector is counting the wrong pod set. Check the constraint, the node labels, and the actual per-zone counts.

open as a page

Kubernetes offers both `topologySpreadConstraints` and pod anti-affinity for keeping replicas apart. What are the practical differences, and when would you still reach for anti-affinity?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Anti-affinity is binary per domain — required anti-affinity allows at most one matching pod per domain, which caps replicas at the number of domains. Spread constraints express a degree of evenness via maxSkew, so many pods per domain are fine. Use anti-affinity when you truly need at most one, or to express repulsion from different pods.

open as a page

Your 9-node bare-metal Kubernetes cluster is 95% requested but only 20% busy, and finance wants two nodes back. How do you respond as the platform owner?

level: principalimportance: should knowfreq 30%

basics

~10 s

Not yet: the scheduler budgets requests, and at 95% requested the cluster cannot even absorb one node loss. Reduce requests first, repack, keep an N+1 ceiling near 89% requested, and only then return hardware.

open as a page

You own node autoscaling for a 140-node Kubernetes cluster shared by 22 product teams. How do you set scale-down policy: its aggressiveness, who may block it, and what that costs?

level: principalimportance: should knowfreq 29%

basics

~20 s

Set scale-down aggressiveness per pool, not globally. Allow vetoes only as reviewed exceptions: budgets that allow a disruption, and time-limited pins on a tainted pool. Report stranded node-hours per team so each veto's cost is visible.

open as a page

As the owner of a shared Kubernetes platform, how would you decide how much capacity runs on spot nodes, and what rules would every team have to follow?

level: principalimportance: should knowfreq 34%

basics

~20 s

Let the workload mix set the spot share, not the discount: only interruption-tolerant work goes there, critical services keep an on-demand floor that serves peak alone, spot pools are diversified, and admission policy enforces who may tolerate them.

open as a page

Why does kube-scheduler never move running pods when Kubernetes nodes become unbalanced, and what does the descheduler do about it?

level: middleimportance: nice to knowfreq 33%

basics

~20 s

kube-scheduler decides only for unbound pods; once bound, a pod stays put. The descheduler is a separate add-on that evicts selected pods through the Eviction API so their controllers recreate them and the scheduler places them again.

open as a page

When would you use a KEDA ScaledJob instead of a ScaledObject to process a queue on Kubernetes, and how do they scale differently?

level: middleimportance: nice to knowfreq 28%

basics

~10 s

A ScaledObject scales a long-running Deployment through a generated HPA. A ScaledJob creates Kubernetes Jobs that run to completion. Use a ScaledJob for long tasks that scale-in must never interrupt.

open as a page

When several Cluster Autoscaler node groups could each fit a Pending Kubernetes pod, how does it pick one to grow, and what do its expanders optimise?

level: middleimportance: nice to knowfreq 31%

basics

~20 s

Cluster Autoscaler simulates the Pending pods against a template node for each node group, then an expander picks among the groups that fit. The default expander, least-waste, picks the group that leaves the least CPU and memory unused.

open as a page

How do low-priority placeholder pods running the pause image give a Kubernetes cluster headroom for pods displaced by spot reclaims?

level: middleimportance: nice to knowfreq 30%

basics

~10 s

Placeholder pods with a very low PriorityClass reserve spare capacity. Displaced pods preempt them and start at once, and the now-Pending placeholders make the Cluster Autoscaler add a node to restore the headroom.

open as a page

What does Kubernetes Dynamic Resource Allocation add over the device-plugin extended-resource model for GPUs, and when would you move to it?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Dynamic Resource Allocation replaces an anonymous device count with published device inventories (ResourceSlices) and claims (ResourceClaims) that select devices by attributes using CEL. The scheduler picks specific devices, and claims can be shared or given fallbacks. It is GA since Kubernetes 1.34.

open as a page

Describe the extension points of the Kubernetes scheduling framework and how you would customise placement without forking kube-scheduler.

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

kube-scheduler is a plugin host with ordered extension points: QueueSort, PreFilter, Filter, PostFilter, PreScore, Score, Reserve, Permit, then PreBind, Bind, PostBind. You customise by writing a KubeSchedulerConfiguration profile that enables, disables or reweights plugins, or by compiling in your own plugin and selecting it via spec.schedulerName.

open as a page

A latency-sensitive service and its in-cluster cache should run close together, while the service's own replicas must stay spread for availability. How would you design that placement with Kubernetes affinity rules, and what would you refuse to guarantee?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Use soft podAffinity from the service to the cache at zone topology to cut cross-zone latency and cost, plus podAntiAffinity on the service's own label for spreading. Keep both soft — hard co-location plus hard spreading over-constrains the scheduler and produces Pending Pods.

open as a page

When would you choose just-in-time node provisioning such as Karpenter over the node-group-based Kubernetes Cluster Autoscaler, and what do you give up by doing so?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Cluster Autoscaler resizes pre-defined, fixed-shape node groups. Karpenter reads Pending Pods and launches individually chosen instances from a broad type list, then consolidates. Choose it for heterogeneous workloads, faster provisioning and better bin-packing; you give up predictability, a cloud-agnostic component, and stable long-lived nodes.

open as a page

You are defining the pod priority tiers for a shared Kubernetes cluster running production services, internal tooling and opportunistic batch jobs. How do you choose the tiers and their values, and what failure modes do you design against?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Use few, widely spaced tiers you can explain during an incident — roughly platform, production, internal, batch, preemptible — set the global default low, grant preemption rights only where start time must be bounded, and control who may use each tier with scoped quotas. Alert on preemption rate as a capacity signal.

open as a page

kube-scheduler's NodeResourcesFit plugin scores nodes with a LeastAllocated strategy by default, and can be configured to MostAllocated. What does each choice do to a cluster, and how would you decide between them?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

LeastAllocated favours the emptiest node, spreading pods for headroom and blast-radius safety. MostAllocated packs pods onto the fullest node that still fits, leaving nodes empty enough to be removed and cutting cost. Choose spread for latency-sensitive services, packing for batch on autoscaled or spot capacity.

open as a page

Your Kubernetes cluster has grown to a dozen teams and several hardware shapes (GPU, high-memory, spot). How would you decide which node groups get taints, who is allowed to tolerate them, and what does that policy cost you?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Taint only where sharing genuinely fails: scarce or expensive hardware, licensing, compliance boundaries, and disruption-prone nodes like spot. Own the taint keys centrally, enforce who may tolerate them at admission, and accept the cost — every exclusive pool needs its own spare capacity and headroom, so partitioned clusters are more expensive than shared ones.

open as a page

You are setting a default multi-zone spreading policy for every workload in a shared Kubernetes cluster. What would you make the default, what would you require teams to opt into, and what does the policy cost in capacity and incident behaviour?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Default to soft zone and hostname spreading for everything (ScheduleAnyway, small maxSkew) via the scheduler's default constraints, and require an explicit opt-in for hard DoNotSchedule constraints, reserved for quorum and compliance workloads. Hard defaults strand pods during zone outages and scale-ups; soft defaults cost some imbalance but never block capacity.

open as a page

showing 31–59 of 59