skip to content

Your scope's aggregate reservation budget is full: a new copy is refused while the four running keep serving — why?

level: middleimportance: should knowfreq 50%

answer

  1. evaluated at creation, not continuously
  2. sums what was declared, not used
  3. all dimensions or none
  4. admitted workloads are never revisited
  5. the refusal lands on the creating controller

basics

~20 s

A scope budget is evaluated when an object is created or changed, against the sum of the reservations already declared in that scope. A request that would push the sum past the ceiling is refused, and workloads already admitted are never revisited, shrunk or killed.

solid answer

~40 s

A scope budget is a ceiling on the aggregate of what the workloads in that scope have declared — typically the sum of their reservations, often object counts and requested storage too. It is evaluated at the moment something is created or changed: the platform adds the incoming request to the scope's current total and, if any dimension would go over, refuses the whole request even where the other dimensions fit. It is never applied backwards. Nothing already admitted is re-examined, so four healthy copies stay four healthy copies and the fifth simply never appears. That asymmetry is what makes the failure awkward to diagnose: no container crashed, so the workload looks fine, and the refusal is recorded on whatever controller was trying to create the next copy.

code

yaml · 18 lines
yaml
scope: reports

budget:
  totalCpuCoresReserved: 8
  totalMemoryGiBReserved: 16

currentlyReserved:
  totalCpuCoresReserved: 7
  totalMemoryGiBReserved: 14

incomingRequest:
  cpuCoresReserved: 1
  memoryGiBReserved: 4

# cpu:    7 + 1  = 8   -> within the ceiling of 8
# memory: 14 + 4 = 18  -> over the ceiling of 16
# result: the whole request is refused on the memory dimension;
#         the workloads already admitted are untouched

go deeper

for a junior

Recall the shape: a scope has a ceiling on the total its workloads may reserve, and once that total is reached the platform stops letting new workloads be created there.

for a middle

Explain the timing and the arithmetic — the check happens when an object is created or changed, it adds the request to the scope's current total, and one dimension going over refuses the whole request.

for a senior

Demonstrate the diagnosis: nothing crashed, so read the controller that was trying to create the copy and the scope's declared total per dimension, not the containers that are serving happily.

for a principal

The judgment is where to set the ceiling and who is allowed to move it. Too tight and teams are blocked by an accounting number rather than by real capacity; too loose and the budget stops being a signal at all.

## What the budget is a ceiling on A scope budget caps the **aggregate** of what the workloads inside one named scope have declared. In practice it covers several dimensions at once: - the **sum of declared CPU reservations** and the **sum of declared memory reservations** across every workload in the scope; - often the **sum of declared ceilings** as well, which is a different and usually larger number; - commonly **object counts** — how many workload objects, how many claims for durable storage, how many external entry points; - often the **total requested durable storage**. Two words carry the whole idea. **Aggregate**: the budget knows nothing about any individual workload, only about the scope's running total. **Declared**: it adds up what specs asked for, which is not what the processes are using. ## When it is evaluated The check runs at write time. Whenever something in the scope is created or changed, the platform computes what the scope's total *would become* if the write were accepted, compares each dimension against that dimension's ceiling, and refuses the write if any one of them would go over. Two consequences follow immediately: 1. **It is all-or-nothing across dimensions.** A request that fits comfortably on CPU and misses on memory is refused entirely. Nothing is trimmed to fit, and no dimension is admitted on its own. 2. **It is not a background sweep.** Nothing walks the scope afterwards looking for workloads to reclaim. A ceiling lowered below the scope's current total does not evict anything — it simply refuses everything new until the total comes down by other means. ## Why nothing already running is touched | A full budget does | A full budget does not do | |---|---| | Refuse creation of the next workload object | Kill, evict or restart anything already admitted | | Refuse an edit that would raise the scope's total | Shrink an admitted workload's declared reservation | | Stop an autoscaler adding copies past the total | Change how the running copies are placed or served | | Stop a replacement that needs headroom for an extra copy | Free any capacity on any host | Admission is a one-time decision. Once a workload's reservation has been counted into the scope's total, the accounting treats the matter as settled. Reclaiming it would mean taking capacity away from something that is currently serving traffic, and an accounting rule does not do that on its own. ## Reading the symptom That asymmetry is why this failure gets diagnosed badly. Every copy that exists is healthy, every check it has passes, and the dashboards are green — the thing that failed is an object that was never created, and absent objects raise no alerts. The triage that works: 1. **Compare the scope's current total against its ceiling, dimension by dimension.** One dimension is over and the others usually are not, which is exactly why the request looked reasonable to whoever wrote it. 2. **Look at the controller that owns the workload, not at the containers.** The refusal is recorded where the creation was attempted: a repeated failure to create, carrying the reason and the dimension. 3. **Ask what most recently raised the total.** A new workload added to the scope, a reservation edited upward, or an autoscaler that grew a different workload can exhaust a budget that has been comfortable for months. The copy that got refused is often not the change that caused the problem. ## Reservations are not usage The most useful thing to know about a scope budget is that it can read one hundred percent full while the hosts underneath are half idle. It sums what specs declared; the processes may be using a fraction of that. This is not a flaw in the accounting. A reservation is what the scheduler subtracts from a host's capacity when it decides placement, so over-declared reservations really do consume placeable capacity even while the process sits idle — the budget is telling the truth about capacity that has been promised away. What it does mean is that the fix is often to correct the declarations rather than to raise the ceiling, and that a scope whose declarations are honest gets far more workloads out of the same number. Platforms differ in the details — whether the sum of ceilings is capped alongside the sum of reservations, whether object counts live in the same budget, whether a scope may be left unbudgeted at all — but the shape above is common to all of them: an aggregate, over declarations, enforced at write time, never applied backwards.

  • Where does the refusal actually surface, if no container ever crashes?
    On whatever was trying to create the object. The controller for that workload records a repeated failure to create, with the dimension that was exceeded, and stops at the copies it already has. The running copies stay healthy, so health dashboards look normal — you find it by reading the creating controller's recent failures and the scope's declared total against its ceiling.
  • What do scope budgets usually cap besides CPU and memory reservations?
    Object counts — how many workload objects, how many claims for durable storage, how many external entry points — and the total durable storage requested. Count caps matter more than they look: they are what stops a misbehaving loop creating objects forever and filling the cluster's state store, which no CPU ceiling would catch.

It behaves like a credit limit rather than a meter: the limit is tested when you try to spend, and reaching it declines the next purchase without reversing the ones already settled.

saying these in an interview costs you the question

  • Says the platform kills running workloads to make room for a new one
  • Thinks the budget tracks measured usage rather than declared reservations
  • Expects a request that fits on CPU to be admitted with memory trimmed
  • Reads healthy running copies as proof the deploy succeeded
  • Assumes a full budget shows up as a crashing container