skip to content

What is the difference between a resource reservation the scheduler places against and the ceiling enforced at runtime?

level: middleimportance: must knowfreq 68%

answer

  1. two numbers, two different readers
  2. one decides where, one caps there
  3. reservation is claimed, not capped
  4. never placed versus throttled or killed
  5. the gap is what creates the choice

basics

~20 s

A reservation is a placement input: the scheduler subtracts it from a node's free capacity when deciding where a workload fits. A ceiling is enforced on the running process by the kernel's accounting mechanism. Different actors, different moments, different failures.

solid answer

~50 s

They are two numbers per resource with two different jobs. The **reservation** is what the scheduler reasons about: it subtracts each replica's reservation from a candidate node's free capacity and places the workload only where the sum still fits. It is a claim on capacity, not a cap — a running container may use more than it reserved when the node has room. The **ceiling** is what the runtime enforces once the process is alive: passing the CPU ceiling throttles it, passing the memory ceiling ends it. So an unplaceable workload is almost always a reservation problem, and a throttled or terminated one is almost always a ceiling problem. Setting the two to the same value collapses the gap and makes the capacity a replica is placed against identical to the capacity it may actually use.

code

pseudocode · 21 lines
pseudocode
declare workload "order-api":
    replicas = 4
    for each replica:
        cpu_reservation    = 0.5 cores   # placement input: subtracted from a node's free capacity
        cpu_ceiling        = 1.0 cores   # runtime cap: over it, the process is throttled
        memory_reservation = 512 MB      # placement input
        memory_ceiling     = 768 MB      # runtime cap: over it, a process is ended

# scheduler, placing one replica on a candidate node:
free = node.allocatable - sum(reservation of every workload already placed)
if free.cpu >= 0.5 and free.memory >= 512 MB:
    place replica here
    free = free - (0.5 cores, 512 MB)     # only the reservation is subtracted
else:
    try the next node

# runtime, every period, on the placed replica:
if cpu_time_used_this_period > cpu_ceiling_share:
    stop scheduling the process until the next period    # throttle, still alive
if accounted_memory > memory_ceiling and nothing reclaimable:
    end a process inside the boundary                    # kill, no catchable signal

go deeper

for a junior

Remember that the two numbers are read by different things: one decides which node a workload lands on, the other caps what the process may use once it is there. Getting that pairing the right way round is most of the answer.

for a middle

Explain the arithmetic — a node's free capacity is its allocatable total minus the reservations already placed on it, not its current measured usage — and say what each number's failure looks like: never placed, versus throttled or terminated.

for a senior

Triage with it. Show that you reach for the reservation when a workload never starts, for the ceiling when a running replica is slow or dies, and for the gap between them when a replica is only slow on busy nodes.

for a principal

Own the default. The distance you standardise between reservation and ceiling is a fleet-level bet on density against predictability, and teams will copy whatever the first template sets. Decide it deliberately rather than letting it be inherited.

## Two numbers, two actors, two moments Almost every container platform lets you declare two numbers per resource for a workload, and the single most common confusion in this material is treating them as a soft limit and a hard limit on the same axis. They are not on the same axis at all. They are read by different components at different times. - The **reservation** is read by the **scheduler**, **before anything runs**. It is arithmetic on capacity: this node has so much unclaimed CPU and memory, this replica claims this much, does it still fit? - The **ceiling** is read by the **runtime and the kernel's accounting mechanism**, **while the process runs**. It is enforcement: this container may use this much and no more. Everything else follows from that split. ## What the reservation actually does When a scheduler evaluates a node, it does not look at how busy the node currently is — or at least, that is not the number it must respect. It looks at how much has already been **claimed** by the workloads placed there. Each placement subtracts that workload's reservation from the node's remaining allocatable capacity. A node with plenty of idle CPU can still be closed to new work because its existing reservations already add up to its capacity. That has three practical consequences: 1. A reservation that no node can satisfy means the workload **never starts at all**. There is no container to inspect, no logs, no exit code — only a workload waiting for a place to go. 2. A reservation is **not a cap**. Nothing stops a container from using more than it reserved, up to its ceiling, when the node happens to have the room. 3. A reservation is also **not free**. It is capacity nobody else can claim, whether or not the workload ever uses it, which is exactly why teams are tempted to set it low. ## What the ceiling actually does The ceiling is enforced continuously on the live process and does not care what the scheduler thought. Passing the CPU ceiling means the process is throttled: it waits for the next period's quota and stays alive. Passing the memory ceiling means a process inside the boundary is ended by the kernel. A ceiling therefore explains symptoms you can only observe after the workload is running. ## The gap between them The interesting design choice is the **distance** between the two numbers. | | Reservation | Ceiling | |---|---|---| | Read by | The scheduler, at placement | The runtime, continuously | | Meaning | Capacity claimed on a node | Maximum the process may use | | Too high | The workload is never placed | Wasted headroom; one replica can take more than its share | | Too low | Placed onto a node that cannot really feed it | Throttling, or termination on memory | | Failure timing | Before the container exists | While the container runs | When the reservation is set **below** the ceiling, the workload can burst into whatever the node has spare. That raises how many replicas fit on a fleet, and it makes each replica's speed partly a function of what else landed beside it. When the reservation **equals** the ceiling, the capacity a replica is placed against is exactly the capacity it can use, and its behaviour stops depending on its neighbours. The cost of that predictability is that the peak is paid for continuously. Deciding which side to sit on for a given workload is a separate judgment call; what matters here is knowing that the gap is what creates the choice, and that the sum of a node's ceilings routinely exceeds what the node physically has. ## Diagnosing with the distinction The fastest use of this model is triage by symptom: - Workload is waiting, no container ever created → look at the **reservation** against real node capacity. - Container runs but requests are slow under load, nothing errors → look at the **CPU ceiling**. - Container exits abnormally and restarts, logs stop mid-line → look at the **memory ceiling**. - Container is fine alone and slow when the node is busy → look at the **gap** between the reservation and the ceiling. ## Where platforms differ Platforms genuinely vary in how much the scheduler weighs beyond reservations — some treat declared reservations as the only placement input, others also score nodes on current measured load or on spreading a workload across failure domains. They also vary in what happens when you declare only one of the two numbers. Because that default differs, state your own numbers explicitly rather than relying on one platform's convention travelling with you.

  • A node shows 60% idle CPU, yet the scheduler will not place another replica on it. How is that possible?
    Placement is arithmetic on reservations, not on measured usage. The workloads already on that node have claimed its capacity through their reservations, and claimed capacity stays claimed whether or not it is being used. The idle time is real but unclaimable, which is precisely the cost of setting reservations above what workloads actually consume.
  • If a reservation is not enforced on the running process, what makes it worth declaring at all?
    It is the only thing that prevents a node being handed more work than it can feed. Without it the scheduler has no basis for saying no, so replicas pile onto whichever node looks available and then contend for the same CPU and memory. The reservation is a promise about density, redeemed at placement time rather than at runtime.
  • What happens if you declare a memory ceiling larger than any single node has?
    The ceiling itself is never the thing that blocks placement — only the reservation is — so the outcome depends on which number is unrealistic. A reservation larger than any node's allocatable capacity leaves the workload permanently unplaced. An enormous ceiling with a modest reservation places fine and simply means the ceiling will never be the thing that stops the container; the node runs out first.

saying these in an interview costs you the question

  • Calls the reservation a soft limit and the ceiling a hard one
  • Says the scheduler subtracts the ceiling from a node's free capacity
  • Thinks a container cannot exceed its reservation while running
  • Believes an unplaced workload is being throttled by its ceiling
  • Assumes node placement is based on current measured usage