How do you expose a container's own CPU and memory requests and limits to the process running inside it in Kubernetes, and what value is injected when a limit is not set?
answer
- resourceFieldRef: limits.cpu / limits.memory / requests.*
- divisor 1m for CPU millicores, 1Mi for memory; rounds UP
- containerName required in the volume form
- unset limit → node allocatable, not empty
- modern JVM/Go read cgroups directly
basics
~20 sUse resourceFieldRef in env or in a downwardAPI volume, selecting limits.cpu, limits.memory, requests.cpu or requests.memory, with a divisor to choose units. If the limit is unset, the injected value falls back to the node's allocatable capacity for that resource — which can badly oversize a runtime.
solid answer
~50 sThe downward API's second selector is `resourceFieldRef`: ```yaml env: - name: MEMORY_LIMIT valueFrom: resourceFieldRef: containerName: app # required in the volume form resource: limits.memory divisor: 1Mi ``` Selectable resources are `limits.cpu`, `requests.cpu`, `limits.memory`, `requests.memory` and `limits.ephemeral-storage`/`requests.ephemeral-storage`. The `divisor` sets the unit: `1` (default) gives cores/bytes, `1m` gives CPU millicores, `1Mi`/`1Gi` give memory in those units. Values are rounded **up** to the next whole unit. The trap: if the container has no limit for that resource, the downward API does not inject empty — it injects the node's **allocatable** amount. A Pod with no memory limit on a 64 GiB node gets `MEMORY_LIMIT=68719476736`, and a runtime sized from that will be OOM-killed or evicted the moment it grows. So always set the limit you intend to read, or validate the value. Modern JVMs and Go read cgroup limits directly, so this is mostly for runtimes and scripts that do not.
code
yaml · 37 linesapiVersion: v1
kind: Pod
metadata:
name: sizing-demo
spec:
containers:
- name: app
image: example/app:1.0
resources:
requests: {cpu: "500m", memory: "512Mi"}
limits: {cpu: "2", memory: "1Gi"}
env:
- name: CPU_LIMIT_MILLICORES
valueFrom:
resourceFieldRef:
resource: limits.cpu
divisor: 1m
- name: MEM_LIMIT_MI
valueFrom:
resourceFieldRef:
resource: limits.memory
divisor: 1Mi
- name: GOMAXPROCS
valueFrom:
resourceFieldRef:
resource: limits.cpu # divisor 1 → cores, rounded up
volumeMounts:
- {name: podinfo, mountPath: /etc/podinfo, readOnly: true}
volumes:
- name: podinfo
downwardAPI:
items:
- path: mem_limit_bytes
resourceFieldRef:
containerName: app
resource: limits.memory
divisor: "1"go deeper
Show the resourceFieldRef snippet and name the four selectors; know that divisor picks the unit.
Add the round-up rule, the containerName asymmetry between env and volume forms, and the fallback to node allocatable when a limit is unset.
Lead with the failure story — no limit set, big node, runtime sized from allocatable, OOM kill or eviction — and enforce explicit limits via policy; note that modern runtimes read cgroups directly.
Treat it as part of a resource-governance policy: mandatory limits, admission rules for containers that self-size, node-pool homogeneity assumptions, and where runtime cgroup awareness makes injection unnecessary.
## The selector Alongside `fieldRef` (Pod metadata/spec/status), the downward API offers `resourceFieldRef`, which reads the *container's* compute resources out of the Pod spec. It works in both delivery forms: ```yaml # environment variable env: - name: CPU_LIMIT_MILLICORES valueFrom: resourceFieldRef: resource: limits.cpu divisor: 1m # volume file volumes: - name: podinfo downwardAPI: items: - path: mem_limit resourceFieldRef: containerName: app resource: limits.memory divisor: 1Mi ``` `containerName` is optional in the env form (it defaults to the container the `env` entry belongs to) and **required** in the volume form, because a volume is Pod-scoped and could be mounted by several containers — the manifest must say whose numbers to write. That asymmetry is a favourite interview detail. ## Divisors and rounding The raw value is a quantity: CPU in cores, memory and ephemeral storage in bytes. `divisor` divides it before injection, and the result is **rounded up** to the next integer. - CPU with `divisor: 1` → cores; `500m` becomes `1` after rounding up, which is usually wrong. - CPU with `divisor: 1m` → millicores; `500m` becomes `500`. Use this for CPU. - Memory with `divisor: 1` → bytes; with `1Mi` → mebibytes; with `1Gi` → gibibytes (again rounded up, so 1.2 GiB becomes 2). Rounding up matters for sizing decisions: computing a JVM heap as a percentage of a rounded-up gibibyte value can exceed the real limit. ## The unset-limit fallback This is the part candidates miss. If the selected resource is not specified on the container, the downward API does not fail and does not emit an empty string — it substitutes the **node's allocatable** amount for that resource (node capacity minus reservations for the kubelet, system daemons and eviction thresholds). Consequences: - A container with no memory limit on a 64 GiB node reads a 60-ish GiB "limit". If a startup script sets the heap to 75 % of that, the process will be OOM-killed by the kernel the moment it actually grows past its cgroup — or, with no limit at all, will trigger node-level eviction and take neighbours down with it. - The value is **node-dependent**, so the same Deployment behaves differently across a heterogeneous node pool. This produces the classic "works on the small nodes, dies on the big ones" bug — the *bigger* node is the dangerous one. Mitigations: always set explicit limits on containers whose sizing you derive this way; prefer `requests.*` when you want the guaranteed floor rather than the ceiling; add a validating policy (Gatekeeper/Kyverno) that rejects Pods with a `resourceFieldRef` to `limits.*` but no such limit; or sanity-check the value in the entrypoint. ## Why anyone does this Historically, language runtimes were unaware of cgroups: a JVM inside a 512 MiB container would size its heap from the *node's* RAM and get OOM-killed. Teams therefore injected `limits.memory` and computed `-Xmx` from it. The same applied to thread pools and `GOMAXPROCS`, which read the machine's CPU count rather than the container's quota. Today most runtimes read the cgroup themselves: modern JVMs enable container support by default and honour `-XX:MaxRAMPercentage` against the cgroup limit; Go's `GOMAXPROCS` still defaults to the visible CPU count and is commonly set from the downward API or by `automaxprocs`. So `resourceFieldRef` remains useful for: - **`GOMAXPROCS`** from `limits.cpu` with `divisor: 1` (rounded up) or a computed value. - **Runtimes and interpreters without cgroup awareness**, and shell entrypoints that size worker counts, buffer pools or connection pools. - **Self-reporting.** Emitting the container's own limits as metrics alongside its usage, so a dashboard can show headroom without joining against the API. - **Sidecars** that must size themselves relative to the main container — the volume form with an explicit `containerName` is exactly for this. ## Interaction with other features With in-place Pod resize (`resources` updates on a running Pod, beta in recent versions), a container's limits can change without a restart. Environment variables injected via `resourceFieldRef` will *not* follow that change; a downwardAPI volume file is republished. Version-sensitive detail worth flagging rather than asserting confidently. `resourceFieldRef` covers only compute resources on the container. It cannot expose extended resources such as `nvidia.com/gpu` counts, QoS class, or the node's total capacity as a distinct field — for those you are back to the API server or to the device plugin's own environment injection. ## How to answer Show the snippet, name the four common resource selectors, explain the divisor and round-up rule, then lead with the fallback-to-node-allocatable trap and its heterogeneous-node consequence. Finish by noting that modern JVM/Go runtimes read cgroups directly, so reach for this when the runtime does not.
- A container has no memory limit but injects limits.memory via resourceFieldRef. What value does it see?The node's allocatable memory — capacity minus kubelet/system reservations and eviction thresholds — not zero or an empty string. Any sizing computed from it will be wildly too large, and it changes with the node the Pod lands on, so the same Deployment misbehaves differently across a heterogeneous node pool.
- Why is `divisor: 1` a poor choice for limits.cpu?CPU quantities are in cores, and the injected value is rounded up to a whole number, so a `500m` limit is reported as `1`. Use `divisor: 1m` to get 500 millicores. The rounding-up rule is deliberate (it never under-reports) but it makes sub-core limits meaningless at core granularity.
- Do modern JVMs still need limits.memory injected this way?Usually not: since JDK 10 (backported to 8u191) the JVM reads the container's cgroup memory limit by default and `-XX:MaxRAMPercentage` sizes the heap against it. Injecting the limit is still useful for non-cgroup-aware runtimes, for shell entrypoints that size worker or pool counts, and for exporting a container's own limits as metrics.
saying these in an interview costs you the question
- Assuming an unset limit yields an empty value rather than the node's allocatable capacity
- Using divisor 1 for CPU and reporting 500m as 1 core
- Forgetting containerName is mandatory in the downwardAPI volume form
- Believing resourceFieldRef can expose GPUs or other extended resources
- Expecting an env-injected limit to change after an in-place Pod resize