skip to content

How do you work backwards from a hard five-hundred-megabyte container ceiling to the maximum heap you configure for a worker?

level: seniorimportance: should knowfreq 50%

answer

  1. the heap number is a remainder
  2. measure warm, at peak concurrency
  3. subtract, then check the ratio
  4. budget pools from their bound
  5. live set after collection, not peak

basics

~20 s

Treat the heap maximum as a remainder, not a choice. Measure every non-heap line at peak concurrency — stacks, runtime metadata, buffers outside the managed heap, allocator overhead — add an explicit margin, subtract the total from the ceiling, then check the remainder against the live set.

solid answer

~50 s

Budget the lines you cannot configure first and give the heap what survives. For a payment-notification worker under a 500 MB ceiling that might read: runtime metadata and mapped code 60 MB, 40 threads at 1 MB of stack each 40 MB, pooled buffers outside the managed heap 30 MB, native allocator caches and overhead 30 MB, and a 10 percent margin of 50 MB — 210 MB in total. That leaves 290 MB, so you configure a 280 MB maximum heap and keep the slack. The final step is the sanity check: measure the live set after a collection, and if 280 MB is not a comfortable multiple of it — here 140 MB, a factor of two — the fix is a smaller live set or fewer threads, not a bigger number in the configuration.

code

pseudocode · 23 lines
pseudocode
CEILING         = 500 MB        // enforced outside the process

metadata        = 60 MB         // measured warm, after all paths ran
stacks          = 40 threads * 1 MB     // pool maximum, not steady state
offheap_buffers = 30 MB         // pool's configured bound
allocator_cost  = 30 MB         // process total minus what we can name
margin          = 10% * CEILING // = 50 MB, held back deliberately

non_heap = metadata + stacks + offheap_buffers + allocator_cost + margin
         = 60 + 40 + 30 + 30 + 50
         = 210 MB

remainder = CEILING - non_heap = 290 MB
max_heap  = round_down(remainder) = 280 MB

// sanity check against demand, not against the setting
live_set  = 140 MB              // measured right after a collection
headroom  = max_heap / live_set = 2.0

if headroom < 1.5 then
    shrink the live set, or cut concurrency, or raise the ceiling
else
    ship 280 MB and record the assumed thread count and live set

go deeper

for a junior

Know that the number you configure for the heap has to leave room for everything else the process holds, and that the container limit covers all of it.

for a middle

Be able to list the budget lines and do the subtraction, and explain why the post-collection live set rather than peak heap usage is the measure of demand.

for a senior

Show the operational discipline: warm measurements at peak concurrency, pools budgeted from their bounds, an explicit margin, and the headroom ratio checked at the end.

for a principal

Make the budget an artefact. Record its assumed thread count and live set beside the deployment configuration so a concurrency change is reviewed as a claim on the margin.

## Why the heap number is computed last Most engineers reach for the heap setting first because it is the one number that is easy to change. That is exactly why it should be decided last. Every other line in the budget is a measurement; the heap setting is the remainder those measurements leave behind. Deciding it first means guessing at the remainder and then discovering the guess was wrong in production, at the ceiling, where the failure mode is a process that vanishes. So the procedure is subtraction, in this order. 1. **Fix the ceiling.** It is given, and it is charged against the whole process. 2. **Measure the non-heap lines at peak concurrency**, not at idle and not at average. 3. **Add an explicit margin** for measurement error and short spikes. 4. **Subtract**, and give the heap what is left. 5. **Validate the remainder against the live set**, and if it fails, change the workload rather than the setting. ## A worked budget for a notification worker The service consumes payment events, renders a message and hands it to a delivery channel. It runs 40 threads and holds a modest set of long-lived structures. Its container ceiling is a hard 500 MB. | Line | Figure | How it was obtained | |---|---|---| | Runtime metadata and mapped program code | 60 MB | measured after the service had exercised every code path, not at startup | | Thread call stacks | 40 MB | 40 threads at 1 MB each, taken from the thread-pool maximum rather than the steady state | | Buffers outside the managed heap | 30 MB | the pool's configured upper bound, not its current occupancy | | Native allocator caches and overhead | 30 MB | process total minus everything accounted for, observed at peak | | Margin | 50 MB | 10 percent of the ceiling, held back deliberately | | **Non-heap subtotal** | **210 MB** | | | **Remainder for the heap** | **290 MB** | 500 minus 210 | | **Configured maximum heap** | **280 MB** | rounded down, keeping 10 MB of slack | The measured live set after a collection is 140 MB, so the configured heap is **2.0 times the live set**. That is the number the budget actually turns on. ## Why the live set, and not peak heap usage, is the input Peak heap usage includes garbage that had not yet been reclaimed when the sample was taken. It therefore measures *when the collector last ran* as much as it measures what the program needs. The live set measured immediately after a collection is the closest available estimate of genuine demand, and it is the figure the headroom ratio should be computed against. - A heap sized at roughly 1.1 times the live set leaves almost nowhere to allocate between collections; the worker either collects continuously or fails to satisfy an allocation. - A heap of two to three times the live set is a common working range for a service that must fit a ceiling; more headroom buys easier collection, and the value of that trade for a whole fleet belongs to a separate discussion. - If the remainder cannot reach a sane multiple of the live set, the budget has told you something useful: this workload does not fit this ceiling at this concurrency. ## The lines people forget - **Stacks are budgeted from the pool maximum**, because the pool will reach it exactly when the service is under the load that also fills everything else. - **Buffer pools are budgeted from their bound**, not their current size, for the same reason — and a pool with no bound cannot be budgeted at all, which is itself the finding. - **Metadata is measured warm.** A worker that has served only health checks has not loaded the code paths that matter. - **The allocator's own overhead is real** and usually shows up as the gap between the sum of everything you can name and the process total. If that gap is large and growing, investigate it rather than absorbing it into the margin. ## Recording the result A budget that lives in one engineer's head is re-derived badly six months later. Write it down next to the deployment configuration as a table of lines and their sources, including the concurrency it assumes. The two facts that make it re-checkable are the **assumed thread count** and the **assumed live set**, because a change in either invalidates the heap number without touching it. That is also what makes the budget reviewable: a change that raises the thread pool from 40 to 80 is visibly a 40 MB claim on a 10 MB slack, and the review catches it before the ceiling does.

  • The thread pool maximum is raised from 40 to 80. What happens to this budget?
    Stacks go from 40 MB to 80 MB, a 40 MB claim against 10 MB of slack and a 50 MB margin, so the margin is nearly consumed before any other line moves. Per-request buffers usually rise with concurrency too, so in practice the heap maximum must come down or the ceiling must go up.
  • What do you do when the remainder is smaller than the live set itself?
    Stop configuring and change the workload. A heap that cannot hold the live set with room to collect will fail whatever number you write. The levers are bounding caches, streaming large payloads instead of buffering them, cutting concurrency per process, or raising the ceiling.
  • Why budget buffer pools from their configured bound rather than their observed size?
    A pool grows to the high-water mark of concurrency and then holds it, and it reaches that mark under exactly the load that fills every other line. Budgeting the observed size under-counts the peak, and a pool with no bound cannot be budgeted at all — which is a finding, not a rounding error.

saying these in an interview costs you the question

  • Sets the maximum heap equal to the container ceiling
  • Budgets thread stacks from the steady-state thread count rather than the pool maximum
  • Uses peak heap usage instead of the post-collection live set as demand
  • Leaves no explicit margin for measurement error and spikes
  • Measures the footprint at startup before code paths are warm
  • Treats an unbounded buffer pool as a small constant in the budget