skip to content

Describe the extension points of the Kubernetes scheduling framework and how you would customise placement without forking kube-scheduler.

level: seniorimportance: nice to knowfreq 28%

answer

  1. QueueSort (one only) → PreFilter → Filter → PostFilter → PreScore → Score → Reserve → Permit
  2. PostFilter runs only when zero feasible nodes — preemption lives there
  3. Permit can approve/deny/WAIT — the gang-scheduling hook
  4. Binding cycle async: PreBind → Bind → PostBind; Unreserve is the rollback
  5. Profiles + schedulerName: many policies, one binary

basics

~20 s

kube-scheduler is a plugin host with ordered extension points: QueueSort, PreFilter, Filter, PostFilter, PreScore, Score, Reserve, Permit, then PreBind, Bind, PostBind. You customise by writing a KubeSchedulerConfiguration profile that enables, disables or reweights plugins, or by compiling in your own plugin and selecting it via spec.schedulerName.

solid answer

~50 s

The framework turns scheduling into a plugin pipeline with named extension points, run per pod: **Scheduling cycle (serial):** `QueueSort` (one plugin only; orders the pending queue) → `PreEnqueue` (scheduling gates) → `PreFilter` (validate + precompute, may reject early) → `Filter` (feasibility per node) → `PostFilter` (runs only when nothing is feasible — preemption lives here) → `PreScore` → `Score` + `NormalizeScore` → `Reserve` (claim state, with `Unreserve` as the rollback) → `Permit` (approve / deny / **wait**, the hook gang-scheduling uses). **Binding cycle (concurrent):** `PreBind` → `Bind` → `PostBind`. Customisation ladder: 1. **Reweight or toggle plugins** in a `KubeSchedulerConfiguration` **profile** — e.g. switch `NodeResourcesFit` to `MostAllocated` for bin-packing, or raise `PodTopologySpread`'s weight. 2. **Multiple profiles in one binary**, each with its own `schedulerName`; pods opt in via `spec.schedulerName`. 3. **Out-of-tree plugins** compiled into a scheduler built on the framework (the scheduler-plugins project ships coscheduling, capacity scheduling, network-aware placement). 4. **A second scheduler** running alongside, or the legacy webhook **Extender** — slower, and rarely the right answer today.

code

yaml · 15 lines
yaml
apiVersion: kubescheduler.config.k8s.io/v1
kind: KubeSchedulerConfiguration
profiles:
  - schedulerName: default-scheduler
    plugins:
      score:
        enabled:
          - name: PodTopologySpread
            weight: 4
  - schedulerName: bin-packing
    pluginConfig:
      - name: NodeResourcesFit
        args:
          scoringStrategy:
            type: MostAllocated

go deeper

for a junior

Know that the scheduler is plugin-based and that filter and score are extension points rather than fixed code.

for a middle

List the main points in order, explain that a KubeSchedulerConfiguration profile enables, disables and weights plugins, and that spec.schedulerName selects a profile.

for a senior

Add Reserve/Unreserve and Permit semantics, the async binding cycle, how coscheduling is built, and the operational tradeoffs of a second scheduler versus a second profile.

for a principal

Judge when custom scheduling is warranted at all — genuinely new semantics versus preferences the built-ins model — and account for the maintenance cost of tracking upstream Go APIs in a hot-path component.

## Why the framework exists Before Kubernetes 1.15-ish, changing placement meant either a scheduler "policy" file with a fixed vocabulary of predicates and priorities, or an HTTP **Extender** the scheduler called out to — a network hop in the middle of a hot loop, with no access to scheduler state. The **scheduling framework** replaced that with an in-process plugin architecture: every built-in behaviour is itself a plugin registered at one or more extension points, so third-party logic is a first-class citizen rather than a bolt-on. ## The extension points, in order **Queue and entry** - **PreEnqueue** — decides whether a pod may enter the active queue at all. `SchedulingGates` lives here: a pod with `spec.schedulingGates` set is held out of scheduling entirely until a controller removes the gates. This is how you implement "do not schedule until quota is approved / peers exist". - **QueueSort** — exactly **one** plugin may be enabled; it defines the ordering of pending pods. The default sorts by pod priority descending, then by the timestamp at which the pod became unschedulable. **Scheduling cycle (one pod at a time, serial)** - **PreFilter** — pod-level checks and precomputation whose result is shared with later points via a `CycleState` object. It may reject the pod outright ("no node can ever satisfy this"), saving a full node sweep. - **Filter** — called per node; returns feasible/infeasible with a reason string. This is where `NodeResourcesFit`, `NodeAffinity`, `TaintToleration`, `NodePorts`, `VolumeBinding`, required `PodTopologySpread` and `InterPodAffinity` live. Runs in parallel across nodes internally, and stops early once `percentageOfNodesToScore` worth of feasible nodes are found. - **PostFilter** — invoked **only** if the feasible set is empty. `DefaultPreemption` implements victim selection here. A PostFilter plugin can also simply produce a better failure message. - **PreScore** — precompute for scoring (e.g. build the topology map once instead of per node). - **Score** — per feasible node, returns an integer; **NormalizeScore** rescales a plugin's raw range into 0–100; the framework multiplies by the plugin **weight** and sums. Highest total wins, random tiebreak among equals. - **Reserve / Unreserve** — the plugin records that the pod is taking something (the volume binder reserves PVC bindings here). `Unreserve` is the rollback path if anything later fails, which makes Reserve the place to hold *stateful* claims rather than doing it in Score. - **Permit** — the most powerful hook: it can **approve**, **deny**, or return **wait** with a timeout. A waiting pod holds its reservation while the plugin decides. Coscheduling (gang scheduling for MPI/Spark jobs) is built here: each pod of the gang waits at Permit until enough peers have arrived, then all are released together — or all time out and are unreserved. **Binding cycle (asynchronous, concurrent with the next pod's scheduling cycle)** - **WaitOnPermit** → **PreBind** (do the real external work, e.g. finalise volume binding — failures here trigger Unreserve) → **Bind** (write the Binding; the first plugin to handle it wins, so a custom Bind plugin can integrate an external placement system) → **PostBind** (cleanup/notification, cannot fail the scheduling). ## Configuring without writing code Most "custom scheduling" needs are met by a `KubeSchedulerConfiguration`: ```yaml apiVersion: kubescheduler.config.k8s.io/v1 kind: KubeSchedulerConfiguration profiles: - schedulerName: default-scheduler plugins: score: enabled: - name: NodeResourcesBalancedAllocation weight: 5 disabled: - name: ImageLocality pluginConfig: - name: PodTopologySpread args: defaultingType: List ``` A **profile** is a named bundle of plugin enablement, weights and args. Several profiles can live in one scheduler process; a pod picks one with `spec.schedulerName`. That is how a cluster offers, say, a spread-oriented default and a bin-packing profile for batch, without running two binaries. ## Writing a plugin An out-of-tree plugin implements the Go interfaces for the points it cares about (`FilterPlugin`, `ScorePlugin`, …), registers itself in a `main.go` built on `k8s.io/kubernetes/cmd/kube-scheduler/app`, and is deployed as *your* scheduler image — either replacing the default scheduler or running as a second deployment with a distinct `schedulerName`. Practical constraints: the plugin runs in the hot path, so it must be fast and use informer caches rather than API calls; it must be deterministic enough to be debuggable; and it takes on the framework's Go version and API compatibility. The upstream `kubernetes-sigs/scheduler-plugins` repository packages the common ones (coscheduling, capacity scheduling, node-resource-topology, trimaran load-aware scoring), which is usually a better starting point than a blank file. ## Choosing the right level Running a **second scheduler** is fine but puts you in charge of the split-brain risk: two schedulers making decisions against independent caches can both place a pod's worth of resources on the same node, and only kubelet admission catches it. The old **Extender** webhook still exists for filter/prioritise/bind, but it costs a network round trip per pod and cannot see scheduler state — reach for it only to integrate an external system you cannot link into Go. And before any of this: most placement intent is already expressible with node affinity, taints, topology spread constraints and resource requests. Custom plugins earn their keep for genuinely different *semantics* — gang scheduling, quota-aware placement across queues, real-utilisation-aware scoring — not for preferences the built-ins already model.

  • Which extension point would you use to implement gang scheduling, and why?
    Permit. It runs after a node has been chosen and reserved but before binding, and it can return a wait result with a timeout, so each pod of the gang holds its reservation while the plugin counts how many peers have arrived. When the quorum is reached all waiting pods are approved together; on timeout they are unreserved and requeued. Doing it earlier would not hold resources, and doing it at bind time would be too late to roll back cleanly.
  • What are the risks of running a second scheduler alongside the default one?
    Each scheduler maintains its own cache and makes decisions independently, so two schedulers can commit the same node capacity at the same time; only kubelet's admission check catches the overlap, producing failed pods with OutOfcpu. You also need every pod to set spec.schedulerName correctly, since a mistake silently routes work to the wrong policy, and you now operate and upgrade two components that must track the cluster's API version.
  • When is changing plugin weights in a profile enough, versus writing a plugin?
    Weights and scoring strategies cover preferences the built-ins already model — spread versus bin-pack, balancing CPU against memory, favouring image locality. You need a plugin when the semantics are genuinely absent, such as scheduling a set of pods atomically, placing against real utilisation rather than requests, or enforcing cross-namespace queue quotas. Start with the upstream scheduler-plugins project rather than writing from scratch.

A conveyor line with labelled stations: you can change the station settings (weights), swap a station out (disable a plugin), or bolt on your own machine at a defined mount point — without rebuilding the conveyor.

saying these in an interview costs you the question

  • Believing filter and score are hardcoded rather than plugins configured per profile
  • Thinking PostFilter runs on every pod instead of only when no node is feasible
  • Assuming multiple QueueSort plugins can be enabled at once
  • Confusing profiles (multiple policies in one binary) with running multiple scheduler deployments
  • Reaching for a custom plugin when affinity, taints and topology spread already express the requirement

context