A host sits at 20% measured CPU use, yet the scheduler will not place anything more on it — why?
answer
- booked, not measured
- a ledger, not a meter
- the scheduler subtracts what was promised
- idle host, fully reserved host
- right-size the reservation, not the fleet
basics
~20 sSchedulers place against reservations, not measurements. Everything already on that host has reserved capacity it is not currently using, so the host is fully booked even though it looks idle, and no further reservation fits.
solid answer
~50 sA reservation is an entry in a ledger, not a reading from a meter. When the scheduler places a workload it subtracts that workload's declared reservation from the host's capacity, once, at placement time; the host is then booked for that amount whether the process uses it or not. Measured usage is a separate number produced by the host and consumed by dashboards and scaling signals — it plays no part in the fit test, and on many platforms no part in placement at all. So a host at 20% measured use and 100% booked is full for placement purposes. Do not confuse this with the `ceiling`, which is what the runtime enforces on the running process; the reservation is what the scheduler subtracts. The gap between booked and used is stranded capacity, and the fix is to re-measure workloads and bring inflated reservations back down.
go deeper
Learn the distinction: the reservation is what a workload asks to have set aside, and the scheduler subtracts it whether the process uses it or not. Measured usage is a separate reading.
Explain who reads which number and when — the scheduler subtracts the reservation once at placement, the runtime enforces the ceiling continuously, the host produces measurements for humans and scaling signals.
Diagnose the booked-versus-used gap on real hosts, name the stranded capacity it creates, and lead reservations back to measured behaviour including peaks and start-up spikes rather than guessing again.
Own it as an estate-level cost: inflated reservations are hosts you pay for and cannot fill. Decide who re-measures, how often, and what evidence is required to defend a reservation.
## Three numbers people mix up | number | declared or produced | who reads it | when it applies | |---|---|---|---| | reservation | declared in the workload spec | the **scheduler**, which subtracts it from a host's capacity | once, at the placement decision | | ceiling | declared in the workload spec | the **runtime**, which enforces it on the running process | continuously, while the process runs | | measured usage | produced by the host | dashboards, scaling signals, humans | continuously, as a reading | The question 'will another workload fit on this host' is answered with the first row only. A host whose reservations sum to its capacity accepts nothing more, whatever the third row says. That is the whole answer to a 20%-busy host refusing work, and the ceiling row is included only so it is not confused with the reservation — what happens to a process that exceeds its ceiling is a different subject entirely. ## Why schedulers book instead of measuring - **A placement has to survive.** Measured use is a snapshot with no future in it. Place three workloads on a host during a quiet ten minutes and their peaks arrive later, on a host that promised nothing to any of them. - **A reservation is a promise the platform can keep.** Booking is how a platform can tell a workload it will have what it declared, rather than what is left over. - **Decisions are concurrent.** Many placements can be in flight at once. Ledger arithmetic composes; readings race, and two schedulers reading the same idle host would both place on it. - **Platforms differ** in whether current load is weighed at all — some ignore it, some use it to break ties when scoring the hosts that already fit. The feasibility test itself is the ledger everywhere. ## The cost: stranded capacity The gap between booked and used is paid for twice. You own hosts you cannot fill, and you have workloads you cannot place. It has a clear signature, and it is worth putting on a dashboard before it becomes an incident: - booked against capacity, per host — the number the scheduler actually uses; - measured usage against booked, per workload — how much of each promise is real; - the largest unreserved gap in the fleet — what the biggest placeable shape is today; - the count of pending workloads while measured utilisation is low — the signature itself. Reservations inflate for honest reasons. Someone sized from a load test's worst minute, or from a start-up spike, or copied a number from a bigger service, or picked a round figure because nothing forced them not to. Nothing in the platform ever revises that number downwards, so it survives every release the workload ever ships. ## Bringing them back to earth 1. Measure over a representative period, including the daily and weekly peak and any start-up spike, not a quiet afternoon. 2. Set the reservation at what the workload genuinely needs to run correctly, with a margin you can defend, rather than at its theoretical maximum. 3. Re-measure after significant releases — a reservation is a claim that ages, and the workload it described may no longer exist in that form. 4. Make the review periodic and someone's job. Nothing lowers a reservation for you. ## Two related traps A workload that declares **no** reservation contributes nothing to the ledger while still consuming real resources. The scheduler keeps believing the host is freer than it is and keeps stacking work onto it, which is how a host ends up both over-loaded and, on paper, under-booked. And the ledger is only as good as the resources it counts. Hosts are booked for the resources the platform models — processor time, memory, sometimes local disk or a device count. Things it does not model, such as memory bandwidth, cache and the network path, are shared regardless of what the ledger says, so a booked-and-used analysis is a placement tool rather than a complete explanation of interference. ## The sentence to have ready Asked why an idle host refuses work, answer in the order the platform works in: the workload declares a reservation; the scheduler subtracts it from a host's capacity at the moment it chooses that host; the host stays booked for it until the workload ends; measured usage is a separate reading that no part of that arithmetic consults. Then say what you would do about it — find the workloads whose measured usage is a small fraction of what they booked, bring those reservations down to defensible numbers, and watch the largest unreserved gap in the fleet recover. That is a concrete answer with an owner, and it beats adding hosts to a fleet that is already mostly idle.
- Which numbers would you put on a dashboard to catch this before workloads start going pending?Booked against capacity per host, measured usage against booked per workload, and the largest unreserved gap in the fleet. Hosts at 95% booked and 20% used, with pending workloads, is the signature; the per-workload ratio then names which reservations to revisit first.
- Why does a workload that declares no reservation make placement worse for everything else?It uses real resources but adds nothing to the ledger the scheduler subtracts from, so the host keeps looking freer than it is and keeps accepting work. The result is a host that is genuinely over-loaded while still appearing to have unreserved room.
A restaurant holds every booked table for the whole evening. The room can look half empty at eight o'clock and still, correctly, have nowhere to seat you.
saying these in an interview costs you the question
- Says the scheduler consults the host's live usage graph when deciding fit
- Assumes an idle host must have room for one more workload
- Calls the reservation the amount the process is allowed to use
- Believes reservations are recalculated automatically from observed usage
- Reports a host as under-utilised without checking how much of it is booked