skip to content

Scheduling & Placement Constraints

How a scheduler picks a host: ruling out nodes that cannot fit the workload, scoring the rest, then honouring rules that spread or pin copies. Asked because stuck, unscheduled workloads are routine.

on this pageshow

questions

5

When a cluster scheduler places one new replica, what are the two stages it runs over the candidate hosts?

level: middleimportance: must knowfreq 62%

answer

  1. two stages, not one
  2. ruled out, then ranked
  3. hard rules first, preferences second
  4. empty candidate list means pending
  5. score only picks among hosts that fit

basics

~20 s

A cluster scheduler first filters: it drops every host that cannot satisfy the workload's reservation, required attributes or hard placement rules. It then scores the survivors against weighted preferences and places the workload on the highest-scoring host.

solid answer

~50 s

Placement is two stages that differ in kind. **Filtering** is a yes/no feasibility pass: for each host, is there enough unreserved room for this workload's reservation, does the host carry the attributes the workload requires, do the hard affinity, anti-affinity and spread rules permit it, and is the host free of a marker that repels work the workload has no exemption for. Hosts that fail drop out. **Scoring** ranks the hosts that survived, using weighted preferences such as how much room is left after placement, how the placement affects spread across failure domains, and whether the image is already on the host. The highest score wins, with ties broken arbitrarily. The split explains the two symptoms you actually see: an unplaced workload failed the filter stage, while a workload on a host you did not expect merely lost the scoring round.

code

pseudocode · 19 lines
pseudocode
candidates = every host in the cluster

for each host in candidates:
    if unreserved_room(host) < workload.reservation:        drop host
    if host.labels does not match workload.required_labels: drop host
    if host.repels_work and workload has no exemption:      drop host
    if host.agent is not reporting:                         drop host

if candidates is empty:
    workload stays pending          # nothing below this line runs
    record a rejection reason per host
    retry when cluster state changes

for each host in candidates:
    score[host] = w1 * room_left_after_placement(host)
                + w2 * spread_gain(host, workload.copies)
                + w3 * image_already_present(host, workload.image)

place workload on the host with the highest score   # ties broken arbitrarily

go deeper

for a junior

Remember the shape: the scheduler first throws out hosts that cannot work, then ranks what is left and picks one. A workload that never starts failed the first step.

for a middle

Be able to name the conditions in the filter stage — unreserved room, required host attributes, hard affinity and spread rules, repel markers — and say what scoring typically weighs among the hosts that survive.

for a senior

Show that you read the per-host rejection reasons before touching anything, and that you know a requirement shrinks the candidate list while a preference only moves a score, so the two produce different incidents.

for a principal

Frame the scoring weights as policy: packing tight raises utilisation and makes large reservations unplaceable, while spreading wastes room but keeps big shapes fitting. Decide which one your fleet's workload mix should buy.

## One workload, many hosts, one decision A cluster scheduler is handed one unplaced unit of work and a list of hosts, and it answers a single question: which host, if any. It answers in two stages that differ in kind. **Filtering** is boolean and decides feasibility — a host either can take this workload or it cannot. **Scoring** is numeric and decides preference — among the hosts that can, which one do we want. Almost everything else in a placement discussion is a detail of one of those two stages. ## Stage one: ruling hosts out The scheduler walks the candidate list and drops any host that fails a hard condition: - **Unreserved room.** Sum the reservations of everything already placed on the host, subtract that from the host's capacity, and compare what is left against this workload's reservation for every resource it declares. Short on any one resource and the host is out. The comparison is made against declared reservations, not against what the host is measured to be using. - **Required host attributes.** A workload may require hosts carrying particular labels — an architecture, a disk class, a hardware feature. Hosts without them drop out. - **Hard placement rules.** Required affinity (land near workloads like these) and required anti-affinity (never land beside another of my own copies) are filters, and so is a strict spread rule across failure domains. - **Host markers.** A host can be marked so that it repels workloads unless a workload carries an explicit exemption. This is how hosts being drained, hosts under maintenance and hosts holding specialised hardware are kept clear of ordinary work. - **Host state.** A host whose agent has stopped reporting is not a candidate at all. If the survivor list is empty, nothing is placed: the workload stays **pending**, the scheduler records why each host was rejected, and it retries as the cluster changes. ## Stage two: ranking the survivors Every surviving host is given a number by a set of weighted functions, and the weights are what a platform's placement policy actually is. Common inputs: - how much unreserved room would remain after this placement — weighted one way this spreads load, weighted the other it packs hosts tight and leaves large gaps free elsewhere; - how the placement affects spread of this workload's copies across hosts and failure domains, where the spread rule was expressed as a preference rather than a requirement; - whether the image is already present on the host, which makes the workload start sooner; - soft affinity preferences declared by the workload itself. Platforms differ in what they weigh, and in whether live measured load is an input at all; what does not differ is that scoring only ever ranks hosts that already passed the filter. The highest score wins. Ties are broken arbitrarily, and some schedulers randomise the tie-break so that a burst of identical workloads does not all land on the same host. | stage | question it answers | what failing it does | |---|---|---| | filter | can this host run this workload at all? | removes the host from the candidate list | | score | which of these hosts do we prefer? | the host is simply not chosen this time | ## Why the split is the whole diagnosis 1. A workload the scheduler has examined and left unplaced failed the **filter** stage. Scoring never leaves a workload unplaced — it only chooses among hosts that already fit. So the question to ask about a pending workload is always which condition ruled every host out. 2. A **preference cannot keep a workload off a host**. If it must not land somewhere, that has to be a hard rule; a soft one only lowers a score, and when the preferred hosts are full the workload lands on the host you were trying to avoid. 3. A **hard rule can make a workload unplaceable**, and that is the price of the guarantee rather than a bug. Every requirement you add shrinks the candidate list before scoring ever runs. ## What the decision is not Placement is a one-time decision. The scheduler does not come back and move a running workload because a better host appeared, because the fleet became lopsided, or because the host it chose is now busy; once placed, the workload stays where it is until something ends it. It is also a decision about the **whole unit being placed**: when the unit is a group of containers that must run together, their reservations are summed and the group lands on one host or not at all — there is no partial placement and no splitting of a reservation across hosts.

  • What happens to a workload when the filter stage leaves no candidate hosts?
    Nothing starts. The workload stays pending, and the scheduler records a per-host rejection reason — insufficient unreserved room, a label mismatch, a spread rule, a repel marker with no exemption. It retries as cluster state changes, so freeing room or relaxing a rule makes the same workload placeable without recreating it.
  • Does a higher score mean the workload will perform better on that host?
    No. The score is the scheduler's preference function over inputs it can see at decision time — room left after placement, spread, whether the image is already there. It is not a prediction of runtime behaviour, and it is computed once; nothing re-scores the placement afterwards.

saying these in an interview costs you the question

  • Says the scheduler simply picks the emptiest host, with no filter stage at all
  • Treats a soft preference as a guarantee that a workload will stay off a host
  • Assumes the scheduler re-runs later and moves a workload once a better host appears
  • Thinks a workload no host can take is rejected and deleted rather than left pending
  • Believes a high enough score can rescue a host that the filter stage ruled out
open as a page

A scoring service with a large reservation stays pending while the cluster shows plenty of free capacity — why?

level: middleimportance: must knowfreq 70%

basics

~10 s

Placement is a per-host fit, never a cluster-wide one. That free capacity is a sum of small leftovers spread over many hosts, and the whole reservation has to fit inside one host's unreserved room.

open as a page

A host sits at 20% measured CPU use, yet the scheduler will not place anything more on it — why?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Schedulers place against reservations, not measurements. Everything already on that host has reserved capacity it is not currently using, so the host is fully booked even though it looks idle, and no further reservation fits.

open as a page

Under a strict even-spread rule across three failure domains, why does a workload's fourth replica stay pending?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A strict spread rule is a filter, not a preference. With three failure domains holding one replica each and no unevenness permitted, placing a fourth anywhere would make the counts uneven, so every host is ruled out and the replica waits.

open as a page

On a shared cluster, when is dedicating a group of hosts to one workload class worth the capacity it strands?

level: principalimportance: should knowfreq 32%

basics

~20 s

Dedicated hosts are worth it when the workloads genuinely differ in the hardware they need, the interference they can tolerate, or the separation required of them — not when one team simply wants to be treated as more important.

open as a page