skip to content

In Kubernetes, why must a JVM or Go runtime size itself from its container's cgroup limits rather than the node's totals, and how does each do it?

level: middleimportance: should knowfreq 52%

answer

  1. cgroup caps, /proc still shows node
  2. thread count versus CPU quota
  3. default heap percentage base
  4. limit, not request, drives CPU count
  5. downward API rounds up

basics

~20 s

Container CPU and memory limits are enforced by cgroups, but a runtime that reads the node's core count or RAM oversizes threads and heap. Current JVMs read cgroup limits by default; Go 1.25+ caps GOMAXPROCS at the CPU limit.

solid answer

~40 s

A container limit is a cgroup setting, not a smaller machine: `/proc/cpuinfo` and `/proc/meminfo` still describe the node. A runtime that sizes GC threads, worker pools or heap from those numbers plans for 32 cores inside a 1.5-CPU quota and gets throttled or OOM-killed. **The JVM** has `-XX:+UseContainerSupport` on by default in current releases: it reads the cgroup memory limit as the base for heap percentages such as `MaxRAMPercentage` (default 25%) and derives the processor count from the CPU limit; `-XX:ActiveProcessorCount` overrides it. With no CPU limit set, it sees the node's CPUs. **Go** from 1.25 sets the default `GOMAXPROCS` from the cgroup CPU limit; older Go uses the visible CPU count, so teams set `GOMAXPROCS` explicitly, for example from the downward API's `limits.cpu`.

code

yaml · 20 lines
yaml
apiVersion: v1
kind: Pod
metadata:
  name: checkout-go
spec:
  containers:
    - name: checkout
      image: registry.example.com/booking/checkout-go:2.7.1
      env:
        - name: GOMAXPROCS
          valueFrom:
            resourceFieldRef:
              resource: limits.cpu
              divisor: "1"
      resources:
        requests:
          cpu: 750m
        limits:
          cpu: 1500m
          memory: 3Gi

go deeper

for a junior

Remember that a container limit is enforced by cgroups but the process may still see the whole node, and that the JVM and Go have their own ways to read the real limit.

for a middle

Explain which runtime settings read the cgroup (UseContainerSupport, MaxRAMPercentage base, GOMAXPROCS default) and what each does when no CPU limit is set.

for a senior

Show how you prove container-awareness before release: run with production limits, print detected CPUs and heap, and pin explicit values where the runtime version cannot be trusted.

for a principal

Weigh a platform default of no CPU limits against runtimes that derive thread counts from limits, and decide whether base images or templates should set processor counts explicitly.

## The problem: a limit is not a smaller machine When a Kubernetes container declares `resources.limits`, the kubelet asks the container runtime to create a **cgroup** with those bounds. For memory, the kernel caps the cgroup's usage and invokes the OOM killer when it is exceeded. For CPU, the kernel's CFS **quota** lets the cgroup use a given amount of CPU time per scheduling period, then pauses it until the next period. What the cgroup does **not** do is change what the process sees when it asks about the hardware. `nproc`, `/proc/cpuinfo` and `/proc/meminfo` generally still report the node. On a 32-core, 128 GiB node, a checkout container limited to `1500m` CPU and `3Gi` memory can still be told "32 CPUs, 128 GiB". ## What goes wrong when the runtime believes the node - **Too many threads**: GC worker threads, a JIT compiler pool, a `ForkJoinPool` or a Go scheduler with 32 active slots all compete for 1.5 CPUs' worth of quota. They burn the quota early in each period and the whole container is paused until the next one, which shows up as tail latency, for example a 310 ms p99 spike on the ticket-booking checkout during a sale. - **Too much memory**: a heap sized as a fraction of 128 GiB is far larger than 3 GiB. The process grows until the kernel OOM-kills it, often without the runtime ever reporting an out-of-memory error of its own. Diagnosing throttling in a live cluster and tuning heap flags are separate topics; the readiness point is that the runtime must be **container-aware before it ships**. ## How the JVM handles it | Concern | Current JVM behaviour | Override | |---|---|---| | Detect container | `-XX:+UseContainerSupport`, on by default on Linux | `-XX:-UseContainerSupport` to disable | | Memory base | cgroup memory limit | explicit `-Xmx` | | Default max heap | `MaxRAMPercentage`, default 25% of that base | `-XX:MaxRAMPercentage=...` | | CPU count | derived from the cgroup CPU limit (quota) | `-XX:ActiveProcessorCount=N` | Key points: 1. Container support has existed for years, but **cgroup v2** detection arrived later than v1 detection, so a very old JDK on a cgroup v2 node can silently fall back to host values. Check the JDK version before the migration. 2. Current JDKs derive the processor count from the CPU **limit**, not from the CPU **request**. A pod with no CPU limit therefore sees every CPU on the node, which is a common surprise on clusters that deliberately omit CPU limits. 3. You can confirm what the JVM decided by running `java -XshowSettings:system -version` inside the container. ## How Go handles it `GOMAXPROCS` is the number of OS threads that may execute Go code at the same time. - **Go 1.25 and later** set the default from the cgroup CPU limit when that limit is lower than the CPU count, and re-check it periodically, so a changed limit is picked up. It does not look at the CPU request. - **Older Go** uses the number of CPUs the process may run on, which in a container is typically the node's count. Teams either link a library that reads the cgroup quota at startup or set the `GOMAXPROCS` environment variable explicitly. ## Setting the value from the pod spec The **downward API** can inject a container's own limit as an environment variable through `resourceFieldRef`. For `limits.cpu` with the default divisor of `1`, Kubernetes rounds **up** to a whole number, so `1500m` becomes `2`. That value is a sensible `GOMAXPROCS` for older Go, or an `-XX:ActiveProcessorCount` input for a JVM when you want the choice explicit and visible in the manifest. ## A readiness check you can run - start the image locally with the same CPU and memory limits it will get in the cluster; - print what the runtime detected (JVM system settings, or `runtime.GOMAXPROCS(0)` in Go); - compare with the limits in the manifest, and fail the pipeline if they disagree. This turns "the runtime reads its cgroup limits" from an assumption into a tested property of the image. ## Common mistakes - **Trusting the base image blindly**: an old runtime baked into a base image can ignore cgroup v2 limits even though the application code is modern. - **Setting a heap equal to the limit**: the heap is only part of process memory; thread stacks, metaspace and native buffers also count toward the cgroup limit. - **Removing CPU limits without revisiting pools**: when a platform drops CPU limits, runtimes that derived thread counts from the quota suddenly size themselves to the whole node.

  • A JVM pod has a CPU request of 2 but no CPU limit. How many processors does a current JVM report?
    It reports the CPUs the process can run on, typically the node's full count, because current JDKs derive the processor count from the CFS quota and a pod without a CPU limit has no quota. The request only sets CPU weight. If you want pools sized to about two CPUs, set `-XX:ActiveProcessorCount=2` or add a limit.
  • What does the downward API inject for GOMAXPROCS when the container sets no CPU limit?
    When no limit is set, `resourceFieldRef` on `limits.cpu` falls back to the node's allocatable CPU, so the injected value is roughly the node's core count, which is exactly the oversizing you were trying to avoid. The pattern only helps when the container has a CPU limit, or when you reference `requests.cpu` deliberately.

saying these in an interview costs you the question

  • Believes a CPU limit makes nproc and /proc/cpuinfo show fewer cores
  • Says every JVM version reads cgroup v2 limits correctly
  • Thinks the JVM sizes its processor count from the CPU request
  • Assumes Go always sets GOMAXPROCS from the container's CPU limit, whatever its version
  • Claims a runtime that ignores the memory limit gets a heap OutOfMemoryError before any OOM kill