How would you choose memory and CPU limits for a latency-sensitive containerized service, given that exceeding the memory limit kills the process while exceeding the CPU quota only delays it?
answer
- memory = cliff (SIGKILL), CPU = slope (throttle)
- size memory from peak anon + 25-50% headroom
- runtime cap below container limit -> diagnosable OOM
- quota above burst, not average; pools sized to quota
- alert on oom_kill and nr_throttled; density vs blast radius
basics
~20 sTreat the two asymmetrically. Memory is a cliff: size it from observed peak plus headroom and cap the runtime below it so failures are diagnosable. CPU is a slope: quota only delays work, so set it above burst demand, size thread pools to it, and watch throttling counters instead of shaving it to average usage.
solid answer
~60 sThe failure modes differ, so the sizing method should too. **Memory is a cliff.** Hitting memory.max ends in a SIGKILL with no application-level diagnostics. Size from the *peak* non-reclaimable footprint (memory.peak, and the anon component of memory.stat), not the average, add headroom for spikes, and cap the runtime's own allocation below the container limit - a JVM at roughly 70% MaxRAMPercentage fails with an OutOfMemoryError and a heap dump rather than being killed opaquely. Remember off-heap: metaspace, thread stacks, buffers, native libraries. **CPU is a slope.** Exceeding the quota throttles, degrading latency proportionally. For latency-sensitive services set quota above burst demand rather than average, or use weights so the service can absorb spikes while still being protected from noisy neighbours, and size thread and GC pools from the quota so a burst does not drain the period in milliseconds. Then close the loop: load-test to find the real peak and the throttling knee, alert on the oom_kill counter and on nr_throttled, and revisit after major dependency or traffic changes. The remaining tradeoff is density versus blast radius - tight limits pack more per node and make correlated failure more likely.
code
bash · 5 lines# peak non-reclaimable footprint and current breakdown
cat /sys/fs/cgroup/system.slice/docker-<id>.scope/memory.peak
grep -E '^(anon|file|slab|sock) ' /sys/fs/cgroup/system.slice/docker-<id>.scope/memory.stat
# throttling knee under load
cat /sys/fs/cgroup/system.slice/docker-<id>.scope/cpu.statgo deeper
Know that exceeding memory kills the process while exceeding CPU quota only slows it, and that limits should exceed observed peak usage.
Derive numbers from measured peak and burst behaviour, and keep the runtime's heap setting below the container limit.
Run the measurement loop end to end - load test, cgroup counters, alerting - and justify quota versus weight per workload class.
Own the policy: which tiers get hard limits, headroom standards, density versus blast radius, error-budget-driven slack, and how the decision is revisited as the system evolves.
## Start from the failure modes Memory and CPU limits are both cgroup settings, but they fail in opposite ways. Crossing memory.max triggers reclaim and then a SIGKILL: an abrupt, undiagnosable termination, dropped in-flight requests, and a restart. Crossing the CPU quota causes throttling: the workload keeps running, just slower, and recovers on its own when demand falls. One is a cliff, the other is a slope. Sizing policy should reflect that asymmetry rather than applying one formula to both. ## Sizing memory Measure the real footprint under realistic load, including warmup, peak traffic, and the tail-risk operations that only happen occasionally - a large export, a cache rebuild, a retry storm inflating in-flight buffers. The number that matters is the non-reclaimable part: the anon component from memory.stat and the high-water mark from memory.peak. Page cache in memory.current is reclaimable and should not drive the limit. Add headroom, commonly 25-50% over observed peak, because the cost of being wrong is asymmetric: a slightly generous limit wastes some node capacity, a slightly tight one kills the process during exactly the traffic spike you cared about. Then make the failure diagnosable. Cap the runtime's own allocation *below* the container limit - a JVM with MaxRAMPercentage in the 65-75% range, Node with max-old-space-size, Go with GOMEMLIMIT - so the process hits an application-level memory error, logs it, dumps state, and exits in a recognizable way instead of vanishing under SIGKILL. Account for everything the container limit covers: heap plus metaspace, thread stacks (stack size times thread count is not negligible at high concurrency), code cache, direct and mapped buffers, native allocator overhead and any helper processes sharing the cgroup. ## Sizing CPU The key insight is that throttling is driven by burstiness against a fixed 100 ms period, not by average utilization. A limit equal to mean usage guarantees throttling on every spike. For latency-sensitive services, either set the quota comfortably above the burst rate, or drop the hard quota in favour of a weight so the service can borrow idle capacity while still being protected under contention. Where the platform supports it, burst tolerance (cpu.max.burst) lets unused quota absorb short spikes. Whatever the quota, align concurrency to it: thread pools, GC parallelism, connection pools and parallel-stream defaults sized from host core count will drain a small quota instantly and turn a burst into a multi-tens-of-milliseconds stall. This is often a bigger win than raising the limit. ## Choosing a posture - **Hard limits everywhere** maximize predictability and protect neighbours, at the cost of utilization and the risk of self-inflicted throttling and kills. - **Weights plus generous memory limits** maximize utilization and latency, but a misbehaving workload can degrade its neighbours' CPU, and a node-level memory shortfall becomes possible. A common compromise: hard memory limits always (memory is not shareable in a graceful way, and the node-level OOM killer is worse than a container-level one), CPU quotas generous or replaced by weights for latency-critical tiers, strict quotas for batch and untrusted work. ## Density versus blast radius Limits are also a packing decision. Tight limits fit more work per node and increase the chance that a normal spike causes a kill; generous limits waste capacity but degrade gracefully. Ask what the service's error budget can absorb: a user-facing API justifies slack that a nightly batch job does not. ## Close the loop Sizing is not a one-time calculation. Load-test to find the peak footprint and the throttling knee. Alert on the memory.events oom_kill counter (which also catches kills of non-PID-1 children that leave a container silently degraded), on memory.current approaching memory.max, and on the nr_throttled to nr_periods ratio. Re-derive limits after dependency upgrades, traffic growth or concurrency changes, and record why each number was chosen so the next engineer does not treat it as arbitrary. In orchestrated environments these numbers become the platform's requests and limits, but the reasoning above is what produces them.
- Why cap the runtime's heap below the container memory limit instead of letting it use everything?Because the two failures are not equally useful. If the runtime hits its own cap it raises a memory error it can log, dump and exit on, and monitoring sees a clear signal; if the cgroup limit is hit first, the kernel sends an uncatchable SIGKILL and you get an exit code and nothing else. The gap also has to cover non-heap memory - metaspace, thread stacks, code cache, direct buffers, native allocations - which the heap setting does not control.
- When would you deliberately set no CPU limit at all?For latency-critical services on capacity you control, where throttling costs more than the isolation is worth. You rely on weights plus scheduling to prevent starvation, keep hard memory limits, and accept less predictable per-container capacity in exchange for absorbing bursts. It is a poor choice for untrusted or batch workloads, where an unbounded consumer can degrade every neighbour.
- How do you validate the numbers rather than guessing?Run a load test that reproduces realistic burst patterns and the rare heavy operations, then read the cgroup counters: memory.peak and the anon component for the footprint, nr_throttled and throttled_usec against latency percentiles for the CPU knee. Set the limit from peak plus headroom, re-run, and confirm no throttling at target load and no OOM under the worst case; then keep the alerts that would tell you the assumptions have drifted.
saying these in an interview costs you the question
- Sizing memory limits from average usage rather than peak non-reclaimable usage
- Setting the container memory limit equal to the configured heap size
- Treating CPU limits as equally dangerous as memory limits, or vice versa
- Assuming a limit at average utilization is safe for a bursty workload
- Setting limits once and never revisiting them after traffic or dependency changes