A host's workloads reserve 40% of its memory while their declared ceilings sum to 200% of it — what is that overcommit betting on?
answer
- density has an arithmetic
- two declared numbers, one is booked
- the gap between them
- peaks assumed independent
- correlated peaks collect the bill
basics
~20 sIt bets that the workloads will not peak at the same time. Placement only subtracts reservations from capacity, so ceilings may sum far past what the host owns. Density is the payoff; coincident peaks are the unpaid bill.
solid answer
~50 sA workload declares two numbers. The **reservation** is what a cluster scheduler subtracts from a host's free capacity when it decides the workload fits; the **ceiling** is what the runtime enforces on the running process. Nothing forces them to be equal, and the gap between them is overcommit. Placement arithmetic only ever reads reservations, so a host can be full on paper at 40% of its memory while the workloads on it are permitted to demand twice the host. The bet is that peak demand is spiky per workload but smooth in aggregate — that few enough of them peak together for the host to cover it. It pays in utilisation and it loses on correlation, not on averages: a shared clock, a shared failed dependency or a fleet-wide restart makes independent-looking workloads peak in the same minute.
code
pseudocode · 17 lineshost.capacityMemoryUnits = 64
for each workload w placed on host:
w.reservation = 4 # subtracted from capacity when w is placed
w.ceiling = 16 # enforced on w's running process
booked = sum(w.reservation for w placed on host) # 40 of 64
possible = sum(w.ceiling for w placed on host) # 160 of 64
# placement reads only the left-hand column
if booked + candidate.reservation <= host.capacityMemoryUnits:
accept candidate onto host
# the right-hand column is the bet
liveDemand = sum(w.currentUsage for w placed on host)
if liveDemand > host.capacityMemoryUnits:
host is short, and the shortfall is settled by ending or moving a workloadgo deeper
Remember that a workload declares two different resource numbers and that only the smaller one is used to decide which host it goes on. Being able to say which is which already puts you ahead of most first screens.
Explain the arithmetic out loud: reservations are subtracted at placement, ceilings are enforced at runtime, and the sum of ceilings on a busy host routinely exceeds the host. Walk through a worked example with real numbers rather than describing it in words.
Show that you price the gap by correlation. Name the things that make peaks land together — a shared schedule, a failing shared dependency, a fleet-wide restart — and say what headroom you keep so that a coincident peak degrades rather than breaks.
Frame the gap as a fleet economic lever with a failure cost attached. Be ready to say which classes of workload you refuse to overcommit at all, and what evidence would make you narrow the gap across the estate rather than for one team.
## The two numbers, and which one the host counts Every workload placed on a shared host carries two resource figures, and confusing them is the most common error in this whole subject. - The **reservation** is a *booking*. A cluster scheduler subtracts it from a host's free capacity to decide whether the workload fits there at all. It is an accounting entry made before the process exists. - The **ceiling** is a *cap*. The runtime enforces it on the running process: processor time is throttled at the ceiling, and a process that tries to hold more memory than its ceiling is ended. Nothing requires the two to match. The **gap between them is overcommit** (also called oversubscription), and its size is a deliberate density choice, not a misconfiguration. ## The arithmetic of a packed host Take a host with 64 units of memory and ten workloads, each reserving 4 units and each declaring a ceiling of 16. 1. **Placement.** Each reservation of 4 is subtracted as the workload lands. After ten, 40 of 64 units are booked and 24 are free — so the host advertises room for more. 2. **Steady state.** Each workload actually holds about 5 units. Total live demand is roughly 50 of 64. Everything is comfortable, and the host looks well used rather than wasted. 3. **Coincident peak.** Ceilings sum to 160 units on a 64-unit host. Only **four** of the ten workloads need to reach their ceiling at once to consume the entire host, and every one of those four is still inside the ceiling it declared. That third line is the whole subject. No workload misbehaved, no ceiling was breached, and the host is short anyway. ## What the bet actually is Overcommit is a statistical argument. One workload's demand is spiky; the *sum* of many independent spiky demands is far smoother than any one of them, because the peaks land in different minutes and cancel. The wider the gap between reservation and ceiling, the more of that smoothing you are spending in advance. So the gap is priced by one question: **how independent are these peaks?** Where independence holds, a fleet runs at high utilisation and nobody notices. Where it does not, the host discovers it all at once. ## When the bet loses — correlation, not the average Averages never tell you the gap is too wide; only coincidence does. The usual correlators: - **A shared clock.** Reporting runs at midnight, reconciliation at 02:00, expiry sweeps on the hour. Workloads that look unrelated are all triggered by the same calendar. - **A shared dependency.** A caching tier degrades, and every workload behind it simultaneously holds more in memory, retries more, and works harder. - **A fleet-wide event.** A rollout restarts everything at once, and start-up is frequently the most expensive minute in a process's life — caches cold, pools being built, configuration being parsed. - **A retry storm.** One slow component converts every caller's queue into extra concurrent work, in the same second, across the host. | Reservation vs ceiling | Utilisation | Behaviour when peaks coincide | Fits | |---|---|---|---| | Reservation equal to ceiling | Lowest — you pay for every peak all the time | Nothing to lose; the host cannot be oversubscribed | Work whose latency tail is the product | | Moderate gap | Good | Brief overshoot absorbed by host headroom | A mixed general-purpose fleet | | Wide gap | Highest | Simultaneous peaks exceed capacity; something must give | Elastic, restartable, retryable work | ## What the gap does not change Overcommit is a property of the *host's* books. It changes less than people assume about the individual workload: - It does not loosen a workload's own ceiling. A container that reaches its own memory ceiling is ended whether the host is idle or full. - It does not buy placement. Raising a ceiling while leaving the reservation alone changes nothing about where the workload lands, because placement reads reservations only. - It does not make the two resources behave alike. A processor shortage and a memory shortage resolve in opposite ways, which is why sensible fleets set the two gaps separately. - It does not partition anything the ceilings never partitioned in the first place. Two workloads inside every declared number still share one storage device, one network link and one pool of cached file pages. The honest summary for an interview: overcommit is how a fleet converts unused headroom into density, the reservation-to-ceiling gap is the size of the bet, and the bet is settled by correlation rather than by the daily mean.
- What makes two workloads' peaks correlated when they look completely unrelated?Usually a clock or a shared dependency. Anything scheduled on the hour or at midnight peaks with everything else scheduled there. A degraded shared tier makes every workload behind it retry and buffer at the same moment, and a fleet-wide rollout restarts them together, so their most expensive minute is the same minute. Independence is an assumption about the workloads' triggers, not about their code.
- Why not set every reservation equal to its ceiling and be done with it?You can, and for latency-critical work you often should. The cost is that you buy every workload's peak permanently: hosts sit at low average utilisation and the fleet needs far more of them. The gap is exactly the trade between predictable isolation and the number of machines you pay for, which is why it is set per class of workload rather than once.
- A workload is placed on the strength of a small reservation and then runs near its much larger ceiling. Is that misbehaviour?No — it is using what it declared, and the runtime is enforcing the ceiling it was given. The mismatch is a policy problem, not a workload problem: someone chose a reservation that does not describe the workload's normal demand. The fix is to correct the declared numbers or the standard that allowed the gap, not to blame the process for reaching a limit it was permitted to reach.
It is an airline selling more seats than the aircraft has, because on a normal day some passengers never turn up. When everyone does turn up, memory behaves like the seat: somebody is taken off the flight. Processor time behaves differently — everyone boards, and the flight is just slower.
saying these in an interview costs you the question
- Thinking overcommit lets a container exceed its own declared ceiling
- Believing the scheduler adds up ceilings and so cannot overpack a host
- Assuming that if reservations fit inside capacity the host cannot run short
- Treating overcommit as a misconfiguration rather than a deliberate density choice
- Claiming a raised ceiling improves where the scheduler places the workload
- Judging the bet by average utilisation instead of by coincident peaks