When would you set a latency-sensitive replica's reservation equal to its ceiling rather than below it?
answer
- the gap is the bet
- predictability against density
- equal means neighbours stop mattering
- still killed by your own ceiling
- latency commitments cannot be conditional
basics
~20 sSet them equal when the replica's performance must not depend on what else lands beside it: the capacity it is placed against becomes the capacity it may use. You pay for the peak continuously, which is worth it for latency commitments and not for elastic background work.
solid answer
~50 sThe distance between the two numbers is a bet. With the reservation **below** the ceiling, a replica is placed cheaply and may burst into whatever the node has spare — so a fleet fits more work, but each replica's speed becomes partly a property of its neighbours, and its throttling appears and disappears with their load. With the reservation **equal** to the ceiling, the capacity claimed at placement is exactly the capacity the runtime will allow, so the replica behaves the same on a quiet node and a busy one. That predictability is what you buy, and the price is paying for the peak all the time. Choose equality for workloads with a latency commitment, for replicas whose burst level is really their normal level, and for anything you need to reproduce in isolation. Leave a gap for batch and background work that can finish later without anyone noticing.
go deeper
Understand what the gap means before arguing about it: capacity above the reservation is allowed but not promised, so whether a replica gets it depends on what else is running beside it.
Explain both directions of the trade — equality buys behaviour that does not depend on neighbours, a gap buys density — and be able to say which column a given workload belongs in and why.
Argue it against a real workload, including the limits: equality does not stop a replica breaching its own ceiling, and it does not divide the resources these two numbers never governed. Back the choice with a measured peak, not a round number.
Set the rule for the estate. Decide where equality is mandatory, what gap is permitted elsewhere, and what evidence justifies a ceiling — and publish the reason, because the first template you ship becomes every team's default for years.
## What the gap actually means A reservation is the capacity a scheduler claims on a node for a replica; a ceiling is what the runtime enforces on that replica while it runs. Whenever the ceiling is higher than the reservation, the difference is capacity the replica is **allowed to use but has not been promised**. Whether it gets that capacity depends entirely on whether the node has it spare at that moment — which depends on the other workloads placed there. So the size of the gap is not a tuning detail. It decides whether a replica's performance is a property of the replica or a property of its neighbourhood. ## The case for equality Setting reservation equal to ceiling collapses the gap. Everything the replica may use has been claimed for it at placement, so: - **Its behaviour is reproducible.** The same workload on the same version behaves the same whether the node is idle or full, which makes load tests, capacity models and incident post-mortems believable. - **Its latency stops being a neighbour's decision.** The replica cannot be starved of capacity it was promised, so the throttling it experiences is only ever against its own quota. - **It is easier to reason about under pressure.** Platforms differ in how they rank workloads when a node genuinely runs short, but a workload using only what it claimed is the easiest case to argue for in any of those schemes. The cost is real and continuous: the peak is paid for at all times, the node's claimed capacity is consumed whether the replica is busy or idle, and fewer replicas fit on the same fleet. ## The case for a gap Leaving the reservation below the ceiling is the density play. More replicas are placed per node, idle headroom is recycled by whoever needs it, and a workload with genuinely rare spikes gets to absorb them without being charged for them continuously. What you give up is determinism: the same replica is fast at 3 a.m. and stuttering at peak, the symptom moves when co-tenants move, and reproducing a latency complaint requires reproducing the whole node's composition. For work with no latency commitment, that is a good trade. ## Choosing between them | Workload | Sensible setting | Reasoning | |---|---|---| | Request path with a latency objective | Reservation equal to ceiling | The commitment cannot be conditional on neighbours | | Replica whose burst is its normal level | Reservation equal to ceiling | The gap would be used continuously, so it was never headroom | | Service being load-tested for a capacity model | Reservation equal to ceiling | The measurement is worthless if it is neighbour-dependent | | Periodic batch or report generation | Reservation below ceiling | Finishing later is acceptable; density is worth more | | Development and test environments | Reservation below ceiling | Predictability is not the goal and cost is | | Genuinely rare, short spikes | Reservation below ceiling | Paying for the spike continuously is poor value | ## The two things equality does not buy This is where candidates overreach, so state both limits explicitly. 1. **It does not protect a workload from itself.** A replica that crosses its own memory ceiling is still ended, and one that exhausts its own CPU quota is still throttled. Equality removes the neighbours from the equation, not the ceiling. 2. **It does not partition everything.** The reservation and ceiling pair governs CPU time and memory. Other things a workload depends on are shared with everything else on the host and are not divided by these numbers at all, so a co-tenant can still leave a mark on a latency tail. How much that matters is a contention question in its own right. ## Making it a standard rather than a per-team habit Because the first workload template a platform ships is copied indefinitely, this choice is usually made once and inherited forever. A workable default is to require equality for anything with a published latency commitment, permit a bounded gap elsewhere, and require that the ceiling be justified by a measured peak in both cases. State the rule and the reason, because a team that sees only the numbers will assume the smaller reservation is simply a cheaper version of the same thing — and it is not; it is a different guarantee.
- Does equal reservation and ceiling mean the replica can never be throttled or ended?No. It removes the neighbours from the equation, not the enforcement. The replica is still throttled when it exhausts its own per-period CPU quota, and still ended when it crosses its own memory ceiling. What changes is that both now depend only on its own behaviour, which is exactly why the failure becomes reproducible.
- What is the fleet-level cost of making equality the default everywhere?Every workload's peak is claimed continuously, so a node's claimable capacity is consumed by headroom that is mostly idle and far fewer replicas fit per node. On a large estate that is a direct and permanent hardware bill. The usual compromise is to require equality where a latency commitment exists and to allow a bounded gap for everything else.
- How would you decide the ceiling itself once you have chosen to make the reservation equal to it?By measuring the real peak rather than the average, and including the start-up high-water mark, because with equality that single number is now both what you pay for and what enforces the wall. Add margin for a larger or slower start, then revisit it with measurements rather than leaving the first guess in place forever.
saying these in an interview costs you the question
- Claims equality prevents the workload being ended at all
- Thinks a lower reservation is just a cheaper version of the same guarantee
- Believes equality partitions every resource a workload uses
- Says a reservation below the ceiling is always wasteful
- Assumes bursting above the reservation is guaranteed to be available