Your measured memory budget for a worker exceeds its container ceiling by twenty percent — which lever do you pull, and on what evidence?
answer
- read the budget before choosing
- fixed per process versus variable
- only one lever lowers demand
- more replicas repeat fixed cost
- name what each lever spends
basics
~20 sPick the lever from the shape of the budget, not from habit. Only shrinking the live set or the per-request working memory reduces real demand; raising the ceiling and re-splitting the work relocate it, and re-splitting multiplies fixed per-process costs.
solid answer
~40 sFour levers exist and they are not equivalent. Bounding caches and streaming large payloads instead of buffering them reduces the demand itself, and is the only lever that makes the workload cheaper everywhere. Cutting concurrency per process lowers stacks and per-request buffers, but the throughput has to reappear somewhere. Splitting into more replicas relocates the variable cost and repeats the fixed cost — metadata and mapped code are paid once per process, so more processes cost more in total. Raising the ceiling is legitimate but is a capacity decision that reduces how many processes fit on a host. The evidence that decides it is the fixed-versus-variable split in the measured budget: a budget dominated by fixed per-process cost argues for fewer, larger processes, and one dominated by per-request cost argues for the opposite.
go deeper
Know the levers exist: hold less data, run fewer threads, run more processes, or ask for a bigger limit. Each one has a cost somewhere.
Explain why the levers are not interchangeable, especially that fixed per-process costs are paid again by every extra replica.
Bring the evidence: a warm budget at peak concurrency, the post-collection live set against the logical working set, and the headroom the remainder leaves.
Decide and defend: choose from the fixed-versus-variable shape, name the cost of the lever you took and the one you rejected, and leave a guard that lets the next person re-derive the number.
## Read the budget before choosing the lever A twenty percent overrun is not one problem, it is at least two different problems wearing the same symptom, and the measured budget tells you which. Split every line into **fixed per process** and **variable with concurrency or data**. | Line | Fixed per process | Variable | |---|---|---| | Runtime metadata and mapped program code | yes | no | | Native allocator base structures | largely | grows with thread count | | Long-lived caches and registries in the live set | yes | grows with data retained | | Call stacks | no | one per thread | | Per-request working buffers, inside and outside the heap | no | concurrency times payload size | | Collection headroom | no | scales with the live set you chose | The ratio between the two columns is the single most decisive fact, because it determines whether splitting the work helps or hurts. ## The four levers, and what each one really does 1. **Reduce demand.** Bound the caches, evict on a size limit rather than on hope, stream large payloads instead of materialising them whole, and hold identifiers instead of objects. This is the only lever that lowers the number on every replica at once, and it usually also lowers collection cost as a side effect. It is the slowest to implement and the most durable. 2. **Cut concurrency in this process.** Fewer threads means fewer stacks and fewer simultaneous working buffers. Cheap and immediate, but the throughput has to be recovered somewhere, so it is rarely a solution on its own. 3. **Split into more, smaller processes.** Each replica carries the full fixed column again. If metadata and base structures are 90 MB of a 500 MB budget, going from two processes to four spends an extra 180 MB across the fleet to buy nothing. If the budget is dominated by per-request buffers, splitting works well. 4. **Raise the ceiling.** Honest and sometimes correct, but it is a capacity decision, not a configuration one: it reduces how many processes fit on a host, and it should be argued in those terms rather than slipped in as a settings change. ## Evidence that distinguishes them - **The fixed-versus-variable ratio**, measured on a warm process at peak concurrency. High fixed share argues for fewer, larger processes and for lever 1 or 4; high variable share argues for lever 2 or 3. - **The live set after a collection**, compared with what the service logically needs to hold. A live set far above the logical working set points at retention you can remove, which is lever 1. - **The headroom ratio the budget leaves.** If the remainder for the heap is barely above the live set, the overrun is structural and configuration cannot fix it. - **The concurrency the throughput target actually requires.** Frequently fewer threads serve the same rate, because the bottleneck is elsewhere; that makes lever 2 free rather than a trade. ## Argue it as a trade, with the cost named Each lever spends something, and a defensible answer names what. - Lever 1 spends engineering time and may lower a cache hit rate, which moves cost to whatever the cache was protecting. - Lever 2 spends throughput per process, and therefore replica count. - Lever 3 spends fixed overhead per replica and adds coordination, more connections and more moving parts to operate. - Lever 4 spends host density, which is a real budget owned by someone else. A smaller heap also leaves less room to collect in, which costs collection work; how much that trade is worth for a fleet is a separate discussion and not one to settle inside a sizing exercise. ## The decision, stated the way it should be reviewed Make the recommendation in three parts, and it becomes reviewable rather than personal: 1. **The measured budget**, line by line, split fixed against variable, taken warm and at peak concurrency. 2. **The lever chosen and the one rejected**, each with the cost it spends — for example, bound the notification-template cache to 40 MB and cut the thread pool from 40 to 30, rejecting a larger ceiling because host density is already the binding constraint. 3. **The guard that proves it held**: an alarm on whole-process footprint against the ceiling, not on heap occupancy, plus the assumed thread count and live set recorded beside the deployment configuration so the next change to either is reviewed as a claim on the same budget. The mark of a principal answer here is not picking a favourite lever. It is choosing from the budget's shape, naming what the choice costs, and leaving behind a number that the next person can re-derive.
- Why can splitting a worker into twice as many replicas increase total memory use across the fleet?Fixed per-process lines — runtime metadata, mapped program code, the allocator's base structures, any per-process long-lived registry — are paid again by every replica. When those dominate the budget, doubling the replica count doubles that entire column while the variable work is merely redistributed.
- What single measurement most often makes this decision obvious?The split of a warm, peak-concurrency budget into fixed per-process and variable lines. A high fixed share argues for fewer, larger processes and for reducing demand; a high variable share means concurrency and payload size are the cost, and cutting or redistributing them works.
- Which guard shows the chosen lever actually held?An alarm on whole-process footprint as a fraction of the container ceiling, together with the assumed thread count and post-collection live set recorded beside the deployment configuration. Heap occupancy against the configured maximum is the wrong signal, because it improves whenever the maximum is raised.
saying these in an interview costs you the question
- Reaches for a larger ceiling first without measuring the budget
- Assumes splitting into more replicas always lowers memory per unit of work
- Ignores that metadata and base structures are paid once per process
- Cuts concurrency without saying where the throughput is recovered
- Presents a lever without naming the cost it spends
- Sizes from average concurrency rather than the peak the pools reach