skip to content

As a fleet standard, how far below its ceiling should a workload be allowed to reserve, and how would you know that gap is too wide?

level: principalimportance: should knowfreq 38%

answer

  1. no single fleet-wide ratio
  2. price the gap per class
  3. processor and memory governed apart
  4. the average is the wrong input
  5. watch the victim's tail, not utilisation

basics

~20 s

There is no single ratio. Set the gap per class of workload, priced by what that class's failure costs, with separate rules for processor time and memory. The guardrail is the victim class's latency, not utilisation.

solid answer

~50 s

A fleet-wide overcommit ratio is the wrong unit. The gap between reservation and ceiling is a bet, and its right size depends on what losing costs the workload sitting behind it. Classify first: work whose latency tail is the product gets little or no gap; elastic, restartable, retryable work gets a wide one, because a lost peak costs a rerun. Set the processor gap and the memory gap separately, since one loses as latency that self-heals and the other loses as an ended process. Price the gap by **correlation**, not by the daily mean — the number that matters is the aggregate peak on a host at the moment of maximum coincidence. And put the guardrail on the victim: rising utilisation is the policy working, so the signal that it went too far is the interactive class's tail moving with a co-tenant's schedule.

go deeper

for a junior

The takeaway here is that the gap between what a workload reserves and what it is allowed to use is a deliberate decision someone makes, with a cost on both sides — not a default to copy from another team.

for a middle

Be able to argue why one ratio cannot fit every workload, and why the risk of a wide memory gap differs in kind from the risk of a wide processor gap rather than merely in degree.

for a senior

Show that you would measure coincident peak rather than average, and that you would watch the affected class's latency instead of the fleet's utilisation, since utilisation rising is the intended outcome.

for a principal

Own the whole trade: classes and their gap defaults, the incentive that keeps declarations honest, the evidence that would make you tighten the standard across the estate, and the willingness to say one tier is not overcommitted at all.

## Why a single fleet ratio is the wrong answer Asking for *the* overcommit ratio assumes every workload pays the same price when the bet loses. They do not. A rerunnable nightly job that is displaced costs a rerun; a checkout path that is displaced costs revenue and trust. The gap between reservation and ceiling is therefore a **policy per class of workload**, expressed as a default with an exception process, rather than a constant. The honest opening answer in an interview is: *it depends on what the class's failure costs, and I would say how I decide it rather than give you a number.* ## Classify the fleet, then set a gap per class | Class | Gap policy | Why | |---|---|---| | Latency tail is the product | Reservation at or very near the ceiling | There is no acceptable degradation; you are buying predictability, and the extra hosts are the price | | General interactive services | Modest gap, watched | Some absorption is fine; the tail is the guardrail | | Elastic, retryable, restartable work | Wide gap | Losing a peak costs a rerun, which is the cheapest failure in the fleet | | Development and pre-production | Widest gap | Nobody's users are behind it, and utilisation is the only thing being optimised | The second question that follows immediately is *which hosts these classes share*, and the cleanest answer for the top row is often that it shares hosts only with its own class. A standard that permits any class anywhere has already lost the argument. ## Govern the two resources separately This is the part most fleet standards get wrong by writing one number. - **Processor time** can be oversubscribed aggressively. Losing the bet costs latency on a curve you can watch, and the host recovers on its own when demand falls. - **Memory** should be oversubscribed narrowly if at all. Losing the bet costs a process that is ended or a workload that is displaced, chosen by pressure rather than by you — and the workload that pays may be one that stayed inside everything it declared. A standard that says 'reserve at least 40% of your ceiling' with no distinction between the two is expressing a single risk appetite for two failure modes that are not comparable. ## Price the gap by correlation, not by the average The average is a comfortable number and it is nearly useless here. What determines whether a host survives is the **sum of live demand at the moment of maximum coincidence**. The inputs worth measuring: 1. Per host, the distribution of aggregate demand against capacity — specifically its upper tail, not its mean. 2. How much of the fleet's demand is clock-driven. Anything on the hour, at midnight, or at the end of a period is correlated by construction. 3. What a fleet-wide event does. A rollout restarts everything, and start-up is often a workload's most expensive minute, so the deploy window is a coincident peak you create yourself. 4. The share of hosts where observed aggregate peak has already crossed capacity, even briefly. ## The guardrail goes on the victim, not on the host The trap in measuring an overcommit policy is that the obvious metric moves the wrong way. **Rising utilisation is the policy working**, so it can never tell you the gap is too wide. Signals that actually indicate the bet is losing: - The interactive class's tail latency moves with a co-tenant's schedule rather than with its own traffic. - Workloads are being ended for memory shortage while inside their own declared ceilings — the clearest possible evidence that the host's books, not the workload, are wrong. - The count of hosts whose aggregate live demand crossed capacity is growing month over month. - Teams are quietly raising ceilings or adding padding to reservations, which is the fleet telling you through behaviour that the standard does not match their experience. Batch work finishing later on busy nights is *not* one of these signals. Absorbing the squeeze is the job that class was given. ## Why the gap widens on its own A standard with no incentive attached decays in one predictable direction, because the two levers point the same way for every team: - Teams raise **ceilings** to avoid being ended under load, since a higher ceiling has no visible cost to them. - Teams lower **reservations** to be scheduled more easily and, where cost follows the reservation, to lower their bill. Both widen the gap, and neither team is behaving badly — they are each optimising what they are measured on. So a durable standard makes the gap **visible and priced**: charge against the higher of reservation and observed peak, require a stated justification for a gap beyond the class default, and publish per-team gap alongside per-team spend. Otherwise the fleet drifts towards maximum overcommit one reasonable local decision at a time. ## The answer that lands Name the classes, set a gap per class with separate processor and memory rules, price it by coincident peak rather than by average, put the guardrail on the tail of the class that cannot absorb a squeeze, attach a cost to the gap so it does not widen by itself — and reserve the right to say that for one tier the correct gap is zero and the fleet simply buys more hosts.

  • If you could put the guardrail on one metric only, which would it be?
    The tail latency of the class that cannot absorb a squeeze, sliced by host. Utilisation tells you the bet is on and never that it is losing; a workload-ended-for-memory count tells you only after it has cost someone. Tail latency per host catches the loss while it is still degradation, and slicing by host shows whether the cause is the workload or its neighbours.
  • How do you stop teams from simply inflating their ceilings under this standard?
    Attach cost to the gap rather than to the reservation alone. If billing follows the reservation, every team is rewarded for reserving less than it uses; if it follows the higher of reservation and observed peak, the incentive flips towards honest declarations. Pair that with a class default that needs a written justification to exceed, and publish each team's gap next to its spend.
  • Is there a case for a gap of zero across a whole tier?
    Yes, and a lead should be willing to say so. Where the latency tail is the product and the cost of one displaced workload exceeds the cost of the hosts saved, reserving at the ceiling and accepting low utilisation is the correct engineering answer. The mistake is treating maximum density as an unconditional goal rather than a trade with a price.

saying these in an interview costs you the question

  • Proposing one overcommit ratio and applying it to every workload class
  • Judging the policy safe because average utilisation sits below capacity
  • Treating a wider gap as free on the grounds that unused memory is wasted anyway
  • Applying the same gap to processor time and to memory
  • Reading rising host utilisation as evidence the policy is not too aggressive
  • Assuming teams will declare honest numbers with no cost attached to the gap