skip to content

You own a shared, multi-tenant Kubernetes cluster. How do you decide what CPU and memory requests and limits to set for a service, and what is your policy on overcommitting CPU versus memory?

level: principalimportance: should knowfreq 44%

answer

  1. CPU compressible → overcommit; memory incompressible → do not
  2. memory request == limit
  3. CPU request from p50-p90, not peak
  4. CFS 100 ms burst throttling argues against CPU limits
  5. VPA recommendation mode + maxLimitRequestRatio guardrail

basics

~20 s

Size requests from observed usage percentiles plus headroom, not guesses. Overcommit CPU deliberately — it is compressible and degrades gracefully. Do not overcommit memory: set memory request equal to limit so failures stay contained and attributable rather than node-wide.

solid answer

~50 s

I start from measurement, not intuition: p50-p90 CPU and peak working-set memory over a representative window, from a load test or production traffic. **Memory:** request equals limit. Memory is incompressible, so overcommitting it converts one service's growth into node-level OOM and eviction that hits innocent neighbours. Equality makes the failure a clean, attributable OOMKill and gives the Pod good eviction ranking. The number comes from peak working set plus runtime overhead (JVM non-heap, Go GC headroom) plus margin. **CPU:** request from steady-state usage so the scheduler packs honestly, and treat the limit as a per-tier policy choice. Latency-sensitive services often run with no CPU limit — shares still enforce fairness under contention, and no limit means no CFS throttling on bursts. Batch and untrusted tenants keep limits so their bursts cannot ruin a node. Then I close the loop: VPA in recommendation mode or usage-versus-request dashboards to catch drift, LimitRange defaults so nothing is unbounded, and quota plus PriorityClass to bound each tenant.

code

yaml · 7 lines
yaml
resources:
  requests:
    cpu: "500m"      # p90 of observed usage
    memory: "1Gi"    # peak working set + runtime overhead
  limits:
    memory: "1Gi"    # equal to request: no memory overcommit
    # no cpu limit: rely on shares for fairness, avoid CFS throttling

go deeper

for a junior

Say you would measure actual usage first and start conservative, and that memory running out kills the container while CPU running out only slows it.

for a middle

Give concrete inputs — usage percentiles, startup peaks, runtime overhead — and justify memory request equal to limit.

for a senior

Argue the CPU-limit trade-off using the CFS throttling mechanism, and describe the guardrails (LimitRange, quota, alerts) that keep sizing honest over time.

for a principal

Frame it as cluster economics and tenancy policy: packing efficiency versus blast radius, per-tier policies, node-shape alignment, and the organisational loop that reviews and enforces the numbers.

## Why this is a judgement call, not a formula Requests decide bin packing, which decides node count, which decides cost. Limits decide blast radius and tail latency. Every number is a trade between utilisation and safety, and the right answer differs per workload tier. A principled answer therefore names the inputs, the asymmetry between CPU and memory, and the feedback loop — not a magic ratio. ## Getting the inputs Start with measurement over a window that includes a real peak: a business-hours peak, a batch window, a deploy, a cache-cold restart. Useful signals are CPU-usage rate percentiles and working-set memory maxima. Startup is often the true peak for both — JIT warm-up and cache priming can double steady-state CPU, and a too-tight CPU request slows startup enough to trip startup probes. New services with no history get an estimate from a load test plus a deliberately generous first setting, tightened later. ## The CPU/memory asymmetry This is the core of the answer. CPU is compressible: taking it away slows a service down. Memory is incompressible: taking it away kills a process. So the two resources deserve opposite overcommit policies. **Memory should not be overcommitted.** Set request equal to limit. The scheduler then reserves exactly what the container may use, so a node can never be promised more memory than it has. The consequences are excellent: a leaking service kills only itself, with `OOMKilled` naming the culprit; neighbours are untouched; the Pod sits in the safest eviction tier. The cost is packing efficiency — you pay for memory that is only occasionally used. That is almost always the right trade, because the alternative failure is a node-level OOM or eviction storm that takes out unrelated workloads and is miserable to diagnose. **CPU should be overcommitted, deliberately.** Almost nothing uses its peak CPU continuously, and CPU shortage manifests as slowness, not death. Requesting p50-p90 usage rather than peak lets the scheduler pack tightly, and shares ensure that when the node is contended each container still gets at least its requested proportion. This is where real cost savings live. ## The CPU limit debate Whether to set a CPU limit at all is a genuine fork with defensible answers on both sides, and an interviewer usually wants reasoning rather than recital. *For no CPU limit on latency-tier services:* CFS quota is enforced in 100 ms windows, so multi-threaded services get throttled during bursts even when average usage is far below the limit — real tail-latency damage for no capacity benefit. Requests already provide proportional fairness under contention. Removing limits typically improves p99 immediately. *For keeping limits:* they cap blast radius from a runaway loop, they make performance reproducible across differently-sized and differently-loaded nodes, they are necessary for untrusted or multi-tenant workloads, and Guaranteed QoS — required for exclusive CPU pinning — needs them. A workable policy: no CPU limits for first-party latency-sensitive services on nodes you control, generous limits (several times the request) for everything else, hard limits for batch and untrusted tenants, and always a `maxLimitRequestRatio` so nobody declares a 50m request with an 8-core limit. ## Guardrails and the feedback loop Numbers set once rot. The loop that keeps them honest: - **LimitRange defaults** so no container is ever accidentally BestEffort or unbounded. - **ResourceQuota per namespace** to bound a tenant's total claim, plus **PriorityClass** so a low-tier team's growth cannot preempt critical services. - **Vertical Pod Autoscaler in recommendation mode** as an advisor; full auto-update mode is risky for stateful or latency-critical services because it restarts Pods to apply changes. In-place Pod resize reduces that pain on newer clusters but is not yet a universal default. - **Dashboards on request-versus-usage** in both directions. Massively over-requested services are the cluster's biggest hidden cost; under-requested ones cause the incidents. - **Node shape matters:** requests that do not tile neatly into node allocatable leave stranded capacity, so sizing requests and choosing node types are part of the same decision. ## What good sounds like "Memory request equals limit, derived from peak working set plus runtime overhead. CPU request from p90 of observed usage; limit generous or absent for the latency tier, strict for batch. Enforced by LimitRange and quota per namespace, reviewed against VPA recommendations quarterly, with alerts on throttled-period ratio and OOMKill rate to catch drift."

  • What is the risk of removing CPU limits across the whole cluster?
    A single runaway container can consume every spare core on its node, so unrequested burst capacity becomes first-come-first-served. Requests still guarantee each container its proportional share under contention, so well-sized neighbours degrade rather than starve — but services that under-request are exposed, performance becomes node-dependent and less reproducible, and untrusted or multi-tenant workloads lose an important containment boundary.
  • How do you handle a service whose memory usage is genuinely bursty, so request-equals-limit wastes a lot?
    First check whether the burst is real working set or reclaimable page cache, since the latter does not need to be reserved. If it is real, either size for the peak and accept the cost, or split the bursty path into a separate Job or worker Deployment that can be scheduled and scaled independently. Deliberately overcommitting memory is a last resort, and then only on nodes reserved for workloads that tolerate eviction.

saying these in an interview costs you the question

  • Quoting a fixed universal ratio such as 'limits are always twice requests' with no workload evidence.
  • Treating CPU and memory symmetrically — overcommitting memory the same way you overcommit CPU.
  • Setting requests from peak usage everywhere, which strands capacity and inflates node count.
  • Assuming removing CPU limits is always safe, with no mention of untrusted workloads or reproducibility.
  • Recommending VPA in auto mode for latency-critical or stateful services without acknowledging the Pod restarts it causes.

context