skip to content

A scoring service with a large reservation stays pending while the cluster shows plenty of free capacity — why?

level: middleimportance: must knowfreq 70%

answer

  1. totals hide shapes
  2. the sum is not a host
  3. fit is tested per host
  4. largest single gap, not the total
  5. whole reservation into one host's room

basics

~10 s

Placement is a per-host fit, never a cluster-wide one. That free capacity is a sum of small leftovers spread over many hosts, and the whole reservation has to fit inside one host's unreserved room.

solid answer

~50 s

The cluster figure is a **sum**, and the scheduler's test is a **fit**. For each host it subtracts the reservations already placed there from the host's capacity and asks whether what remains is at least this workload's reservation; a large reservation can fail that test on every host while the leftovers add up to several times its size. Six hosts of 8 units each, 39 units already reserved, leaves 9 units free cluster-wide but at most 2 free on any one host — a workload reserving 3 stays pending. The fixes all change a shape rather than a total: right-size the reservation, split the work into more and smaller copies, free a host by moving a large low-value workload off it, or add a host big enough to hold the shape. Read the per-host rejection reasons first, because a placement rule produces the same symptom for a different cause.

code

pseudocode · 15 lines
pseudocode
workload.reservation = 3 units

total_free = 0
best_fit   = none

for each host in cluster:                       # 6 hosts, capacity 8 each
    free = host.capacity - sum(reservation of each workload placed on host)
    total_free = total_free + free              # 2 + 1.5 + 1 + 2 + 0.5 + 2
    if free >= workload.reservation:
        best_fit = host

# total_free  = 9 units   -> what the cluster view reports
# largest free on one host = 2 units
# 2 < 3, so best_fit stays none on every iteration
# result: workload stays pending with 9 units free

go deeper

for a junior

Remember that a workload runs on one host, so the number that matters is the free room on a single host, not the free capacity of the whole cluster added together.

for a middle

Explain the arithmetic: capacity minus the reservations already placed, per host, per resource, compared against the whole reservation of the unit being placed. Then name the fixes that change a shape.

for a senior

Show the diagnosis order — rejection reasons first, to separate fragmentation from a rule — and pick a fix on evidence rather than adding hosts reflexively when the largest gap is the binding number.

for a principal

Treat it as fleet policy: how tightly to pack, how much per-host headroom to keep for large shapes, and whether the workload mix should be steered towards shapes the fleet can actually absorb.

## A sum is not a fit Every cluster view shows a total: capacity across all hosts, reserved across all hosts, and the difference between them. That difference is real, and it is almost useless for answering 'will this workload be placed'. The scheduler never places a workload on 'the cluster'. It places it on one host, and the test it applies is whether **this host's unreserved room** is at least **this workload's whole reservation**, resource by resource. A total tells you nothing about whether any single host passes that test. This is a packing problem, and the quantity that matters is the **largest single-host gap**, not the sum of gaps. ## The worked case A fleet of six hosts, each with 8 units of a resource, so 48 units of capacity. The workloads already placed reserve 39 units, distributed unevenly: | host | capacity | reserved | unreserved | |---|---|---|---| | 1 | 8 | 6.0 | 2.0 | | 2 | 8 | 6.5 | 1.5 | | 3 | 8 | 7.0 | 1.0 | | 4 | 8 | 6.0 | 2.0 | | 5 | 8 | 7.5 | 0.5 | | 6 | 8 | 6.0 | 2.0 | The dashboard says 9 units free, which is three times what the pending scoring service is asking for. The largest gap on any one host is 2. The service reserves 3. Every host fails the filter, the candidate list is empty, and the workload stays pending — with 9 units of genuinely free capacity in the cluster. Nothing is broken; the fleet is fragmented. Two details make this worse than it first looks. The test runs **per resource**: a host with room for the processor reservation but not the memory reservation still drops out, so the effective gap is the smaller of the two. And it runs on the **whole unit being placed**: when the unit is a group of containers that run together, their reservations are summed, so a group is a larger shape than any of its members. ## What actually fixes it 1. **Right-size the reservation.** Measure what the service really needs and declare that. This is the highest-value fix, because an inflated reservation fragments every host it lands on as well as failing to fit. 2. **Make the shape smaller.** More copies, each reserving less, fit into leftovers you already have — if the workload can be split at all. 3. **Defragment.** Removing or re-placing one large, low-value workload can merge two useless gaps into one usable one. Placement is a one-time decision, so nothing does this for you. 4. **Add capacity of the right shape.** A host must be able to hold the whole reservation to help. Several small hosts can raise the total while leaving the largest gap unchanged, which places nothing. 5. **Keep deliberate headroom.** If the fleet is expected to accept large shapes, leave room for one on a few hosts on purpose rather than packing everything to the last unit. ## Rule out the other cause first Fragmentation is not the only reason every host fails the filter. Each of these produces exactly the same visible symptom — pending, with the cluster looking half empty — and each has a completely different fix: - a required host attribute that no host in the fleet carries, often a typo or a label that was renamed; - a strict spread rule across failure domains that the current copy count cannot satisfy; - a required anti-affinity rule with fewer eligible hosts than copies; - hosts marked to repel work, with no exemption on this workload — draining, maintenance or a dedicated group; - a resource other than the obvious one: the fit test runs per resource, so memory can block a placement while processor room is plentiful. So read the scheduler's per-host rejection reasons before doing anything. 'Insufficient unreserved room' on every host is fragmentation. A mismatch reason is a rule. ## Why totals mislead so reliably Capacity planning is done in totals, so that is the number people carry around. A fleet sized from totals alone will accept small workloads happily for months and then refuse a large one, and the refusal looks like a platform fault because the headline number says there is room. The honest fleet-level question is not 'how much is free' but 'what is the largest shape this fleet can still accept, and how often do we need to accept one'. Two habits keep the answer to that question healthy. Track the **largest single-host gap** beside the cluster total, so the number that actually decides placements is on the same screen as the number everyone quotes; a fleet whose largest gap has been shrinking for a month is one release away from its first pending workload. And treat the **shape mix** as a design input: a fleet that must accept occasional very large reservations either keeps a few hosts deliberately loosely packed or accepts that large shapes wait, and choosing by accident means choosing the second one. Neither habit needs new machinery — both are arithmetic over numbers the scheduler already has.

  • How would you tell fragmentation apart from a placement rule that rules every host out?
    Read the per-host rejection reasons the scheduler records. Fragmentation shows insufficient unreserved room on every host; a rule shows a mismatch — a missing host attribute, a spread rule, a repel marker with no exemption. The symptom is identical and the fixes share nothing.
  • Why can halving each copy's reservation help when adding another host does not?
    Smaller shapes fit the leftovers the fleet already has, and there are many of them. A new host only helps if it can hold the whole reservation by itself; adding small hosts raises the cluster total while leaving the largest single gap exactly where it was.

A stand with nine empty seats, never two of them together, still cannot seat a family of three. The total is real; only the fit decides.

saying these in an interview costs you the question

  • Adds up free capacity across the cluster and concludes the workload must fit somewhere
  • Assumes the scheduler can split one reservation across two hosts
  • Calls it a scheduler bug without reading the per-host rejection reasons
  • Thinks adding any host helps, whatever the shape of the pending reservation
  • Believes the workload will start on its own once average cluster usage drops