When a cluster scheduler places one new replica, what are the two stages it runs over the candidate hosts?
answer
- two stages, not one
- ruled out, then ranked
- hard rules first, preferences second
- empty candidate list means pending
- score only picks among hosts that fit
basics
~20 sA cluster scheduler first filters: it drops every host that cannot satisfy the workload's reservation, required attributes or hard placement rules. It then scores the survivors against weighted preferences and places the workload on the highest-scoring host.
solid answer
~50 sPlacement is two stages that differ in kind. **Filtering** is a yes/no feasibility pass: for each host, is there enough unreserved room for this workload's reservation, does the host carry the attributes the workload requires, do the hard affinity, anti-affinity and spread rules permit it, and is the host free of a marker that repels work the workload has no exemption for. Hosts that fail drop out. **Scoring** ranks the hosts that survived, using weighted preferences such as how much room is left after placement, how the placement affects spread across failure domains, and whether the image is already on the host. The highest score wins, with ties broken arbitrarily. The split explains the two symptoms you actually see: an unplaced workload failed the filter stage, while a workload on a host you did not expect merely lost the scoring round.
code
pseudocode · 19 linescandidates = every host in the cluster
for each host in candidates:
if unreserved_room(host) < workload.reservation: drop host
if host.labels does not match workload.required_labels: drop host
if host.repels_work and workload has no exemption: drop host
if host.agent is not reporting: drop host
if candidates is empty:
workload stays pending # nothing below this line runs
record a rejection reason per host
retry when cluster state changes
for each host in candidates:
score[host] = w1 * room_left_after_placement(host)
+ w2 * spread_gain(host, workload.copies)
+ w3 * image_already_present(host, workload.image)
place workload on the host with the highest score # ties broken arbitrarilygo deeper
Remember the shape: the scheduler first throws out hosts that cannot work, then ranks what is left and picks one. A workload that never starts failed the first step.
Be able to name the conditions in the filter stage — unreserved room, required host attributes, hard affinity and spread rules, repel markers — and say what scoring typically weighs among the hosts that survive.
Show that you read the per-host rejection reasons before touching anything, and that you know a requirement shrinks the candidate list while a preference only moves a score, so the two produce different incidents.
Frame the scoring weights as policy: packing tight raises utilisation and makes large reservations unplaceable, while spreading wastes room but keeps big shapes fitting. Decide which one your fleet's workload mix should buy.
## One workload, many hosts, one decision A cluster scheduler is handed one unplaced unit of work and a list of hosts, and it answers a single question: which host, if any. It answers in two stages that differ in kind. **Filtering** is boolean and decides feasibility — a host either can take this workload or it cannot. **Scoring** is numeric and decides preference — among the hosts that can, which one do we want. Almost everything else in a placement discussion is a detail of one of those two stages. ## Stage one: ruling hosts out The scheduler walks the candidate list and drops any host that fails a hard condition: - **Unreserved room.** Sum the reservations of everything already placed on the host, subtract that from the host's capacity, and compare what is left against this workload's reservation for every resource it declares. Short on any one resource and the host is out. The comparison is made against declared reservations, not against what the host is measured to be using. - **Required host attributes.** A workload may require hosts carrying particular labels — an architecture, a disk class, a hardware feature. Hosts without them drop out. - **Hard placement rules.** Required affinity (land near workloads like these) and required anti-affinity (never land beside another of my own copies) are filters, and so is a strict spread rule across failure domains. - **Host markers.** A host can be marked so that it repels workloads unless a workload carries an explicit exemption. This is how hosts being drained, hosts under maintenance and hosts holding specialised hardware are kept clear of ordinary work. - **Host state.** A host whose agent has stopped reporting is not a candidate at all. If the survivor list is empty, nothing is placed: the workload stays **pending**, the scheduler records why each host was rejected, and it retries as the cluster changes. ## Stage two: ranking the survivors Every surviving host is given a number by a set of weighted functions, and the weights are what a platform's placement policy actually is. Common inputs: - how much unreserved room would remain after this placement — weighted one way this spreads load, weighted the other it packs hosts tight and leaves large gaps free elsewhere; - how the placement affects spread of this workload's copies across hosts and failure domains, where the spread rule was expressed as a preference rather than a requirement; - whether the image is already present on the host, which makes the workload start sooner; - soft affinity preferences declared by the workload itself. Platforms differ in what they weigh, and in whether live measured load is an input at all; what does not differ is that scoring only ever ranks hosts that already passed the filter. The highest score wins. Ties are broken arbitrarily, and some schedulers randomise the tie-break so that a burst of identical workloads does not all land on the same host. | stage | question it answers | what failing it does | |---|---|---| | filter | can this host run this workload at all? | removes the host from the candidate list | | score | which of these hosts do we prefer? | the host is simply not chosen this time | ## Why the split is the whole diagnosis 1. A workload the scheduler has examined and left unplaced failed the **filter** stage. Scoring never leaves a workload unplaced — it only chooses among hosts that already fit. So the question to ask about a pending workload is always which condition ruled every host out. 2. A **preference cannot keep a workload off a host**. If it must not land somewhere, that has to be a hard rule; a soft one only lowers a score, and when the preferred hosts are full the workload lands on the host you were trying to avoid. 3. A **hard rule can make a workload unplaceable**, and that is the price of the guarantee rather than a bug. Every requirement you add shrinks the candidate list before scoring ever runs. ## What the decision is not Placement is a one-time decision. The scheduler does not come back and move a running workload because a better host appeared, because the fleet became lopsided, or because the host it chose is now busy; once placed, the workload stays where it is until something ends it. It is also a decision about the **whole unit being placed**: when the unit is a group of containers that must run together, their reservations are summed and the group lands on one host or not at all — there is no partial placement and no splitting of a reservation across hosts.
- What happens to a workload when the filter stage leaves no candidate hosts?Nothing starts. The workload stays pending, and the scheduler records a per-host rejection reason — insufficient unreserved room, a label mismatch, a spread rule, a repel marker with no exemption. It retries as cluster state changes, so freeing room or relaxing a rule makes the same workload placeable without recreating it.
- Does a higher score mean the workload will perform better on that host?No. The score is the scheduler's preference function over inputs it can see at decision time — room left after placement, spread, whether the image is already there. It is not a prediction of runtime behaviour, and it is computed once; nothing re-scores the placement afterwards.
saying these in an interview costs you the question
- Says the scheduler simply picks the emptiest host, with no filter stage at all
- Treats a soft preference as a guarantee that a workload will stay off a host
- Assumes the scheduler re-runs later and moves a workload once a better host appears
- Thinks a workload no host can take is rejected and deleted rather than left pending
- Believes a high enough score can rescue a host that the filter stage ruled out