A JVM service runs in a container capped at 512 MB and keeps getting killed by the kernel, even though its heap usage looks small and it never throws `java.lang.OutOfMemoryError`. How does a modern JVM discover container CPU and memory limits, and how would you size it so this stops happening?
answer
- cgroup limits RSS; -Xmx limits heap only
- UseContainerSupport default since JDK 10 / 8u191
- default heap ~25% of container limit
- non-heap: metaspace, 1MB stacks, direct buffers, NMT
- cgroup v2 needs JDK 15+ / backports
basics
~20 sModern JVMs read cgroup limits (container support is default-on since JDK 10, backported to 8u191) and default the heap to about 25% of the container limit. The kill comes from total process memory: heap plus metaspace, thread stacks, code cache, direct buffers and native allocations. Size with -XX:MaxRAMPercentage and leave real non-heap headroom.
solid answer
~50 sThe cgroup limit covers the whole process's resident memory, not just the Java heap. A small heap plus metaspace, per-thread stacks (~1 MB each), JIT code cache, GC structures, direct or mapped byte buffers and native library allocations can easily exceed 512 MB — and the kernel sends SIGKILL, so the JVM never gets to throw `OutOfMemoryError`. Since JDK 10 (backported to 8u191) `-XX:+UseContainerSupport` is on by default: the JVM reads `memory.max` and `cpu.max` from the cgroup instead of host values, derives its default max heap as a percentage (about 25%) of the container limit, and computes active processor count from the CPU quota. Cgroup v2 detection needs a newer JDK (15+, backported into later 11/17 builds); older JDKs on a v2 host silently see host memory. Fix by budgeting: set `-XX:MaxRAMPercentage=70` (or an explicit `-Xmx`), bound `-XX:MaxMetaspaceSize` and `-XX:MaxDirectMemorySize`, cap thread pools, verify with Native Memory Tracking, then raise the container limit if real demand requires it.
code
bash · 2 linesdocker run --rm --memory=512m --cpus=0.5 eclipse-temurin:21-jre \
java -XX:+PrintFlagsFinal -version | grep -E 'MaxHeapSize|ActiveProcessorCount|UseContainerSupport'go deeper
Know the heap is only part of the process, that the container limit covers everything, and that modern JVMs read the container limit rather than host RAM.
Enumerate the non-heap regions, explain the default heap being a percentage of the container limit, and set MaxRAMPercentage plus metaspace and direct-memory bounds deliberately.
Drive the diagnosis with Native Memory Tracking and RSS-versus-limit evidence, know the JDK and cgroup-version pitfalls, and prefer a catchable in-JVM error with a heap dump over a kernel kill.
Standardise a memory-budget policy across services and base images — percentage sizing, bounded non-heap regions, mandated headroom — and weigh the fleet-wide cost of that headroom against the blast radius of silent kills.
## Why the kill happens with no OutOfMemoryError Two different budgets are in play. `-Xmx` bounds the Java heap; the container's cgroup limit bounds the **resident set size of the whole process**. When the heap is exhausted the JVM throws `java.lang.OutOfMemoryError`, which you see in the logs. When the cgroup limit is exceeded the kernel sends SIGKILL — uncatchable, no log line, exit code 137 with `OOMKilled: true`. A silent death alongside a healthy-looking heap is the signature of the second case. Everything outside the heap still counts against the limit: - **Metaspace**: class metadata, unbounded by default; frameworks that generate classes grow it steadily. - **Thread stacks**: roughly 1 MB each on 64-bit, so 300 threads is ~300 MB — often the single biggest surprise. - **JIT code cache** and compiler arenas. - **GC overhead**: card tables, remembered sets, region metadata — a few percent of heap, more for region-based collectors. - **Direct and mapped byte buffers**: NIO, pooled network buffers, memory-mapped files. - **Native allocations**: database drivers, compression, TLS, and glibc malloc arenas, which can retain a surprising amount of freed-but-unreturned memory (`MALLOC_ARENA_MAX=2` is a common mitigation). ## How the JVM sees the container Historically the JVM read `/proc/meminfo` and the host CPU count, so a container limited to 512 MB on a 64 GB host would default its max heap to about 16 GB and die almost immediately. `-XX:+UseContainerSupport`, added in JDK 10 and backported to 8u191, made the JVM read the cgroup instead: `memory.max` for the memory ceiling and `cpu.max` (quota over period) for the processor count. It is enabled by default. Two caveats matter. First, the default heap is a **fraction** of the container limit — historically `MaxRAMFraction=4`, now expressed as `MaxRAMPercentage`, defaulting to roughly 25% for containers of meaningful size. That conservatism leaves room for the non-heap regions above, but it also means a 512 MB container defaults to about a 128 MB heap, which is often why an app instead OOMs *inside* Java on small containers. Second, **cgroup v2** detection required a newer JVM (JDK 15, later backported into 11.0.16+ and 17.x); an older JDK on a v2-only host falls back to host values and mis-sizes everything. CPU awareness follows the same path: active processor count is derived from the CPU quota, and that number drives GC thread counts, JIT compiler threads, the common ForkJoinPool, and any library that calls `Runtime.availableProcessors()`. `--cpus=0.5` therefore yields one reported processor, changing GC and pool sizing rather than merely slowing execution. This is not JVM-specific — Go and .NET have equivalent container-awareness stories, and older runtime versions in all three ecosystems read host values. ## Sizing method 1. **Measure the real total.** Run under representative load and watch RSS against the limit, plus Native Memory Tracking (`-XX:NativeMemoryTracking=summary`, then `jcmd <pid> VM.native_memory summary`) to attribute non-heap usage. 2. **Budget explicitly.** Prefer `-XX:MaxRAMPercentage` (say 60–75% for a service with modest native usage) over a raw `-Xmx`, so the same image behaves sensibly when the limit changes. Bound `-XX:MaxMetaspaceSize` and `-XX:MaxDirectMemorySize` so runaway non-heap growth surfaces as a catchable Java error instead of a kernel kill. 3. **Cap concurrency.** Thread pools and connection pools have a memory cost; unbounded pools defeat any heap budget. 4. **Leave headroom.** The container limit should exceed measured peak RSS by a real margin — typically 20–50% — because the limit is a cliff, not a slope. 5. **Prefer failing inside the JVM.** An `OutOfMemoryError` with a heap dump (`-XX:+HeapDumpOnOutOfMemoryError`, written to a mounted volume) is diagnosable; a SIGKILL is not. Tight internal bounds plus a slightly larger container limit buy exactly that. ## Sanity checks `java -XX:+PrintFlagsFinal -version` inside the container prints the effective `MaxHeapSize` and `ActiveProcessorCount` as the JVM computed them — the fastest way to prove whether container detection worked. If `MaxHeapSize` looks like a quarter of the *host's* RAM rather than the container's, you are on an old JDK, on cgroup v2 without support, or someone disabled container support explicitly.
- Heap, metaspace and thread stacks are all accounted for, yet RSS still creeps past the limit. Where else do you look?Native allocations outside JVM accounting: direct and mapped byte buffers, database, TLS and compression libraries, and glibc malloc arenas that retain freed memory. Enable Native Memory Tracking to attribute what the JVM knows about, compare that total to RSS, and treat the gap as native. Bounding `-XX:MaxDirectMemorySize` and setting `MALLOC_ARENA_MAX=2` are the usual mitigations.
- Why does `--cpus=0.5` change application behaviour beyond making it slower?The JVM derives its active processor count from the CPU quota, and that number sizes GC threads, JIT compiler threads and the common ForkJoinPool, and is what `Runtime.availableProcessors()` returns to libraries. At 0.5 CPU the JVM reports a single processor, so parallel GC and every pool sized from that value shrink — a configuration change, not just a throughput change.
- Would raising the container memory limit alone fix this?It buys headroom and may stop the kills, but it bounds nothing: unbounded metaspace, thread counts or direct-buffer usage will grow into whatever you give them. The durable fix is explicit budgets per region plus a limit sized above measured peak, so a breach surfaces as a catchable Java error with a heap dump rather than an untraceable SIGKILL.
-Xmx is the size of your suitcase; the container limit is the airline's total baggage allowance, including the carry-on, laptop and coat you forgot to weigh.
saying these in an interview costs you the question
- Assuming `-Xmx` bounds total container memory usage
- Expecting an `OutOfMemoryError` when the kernel OOM-kills the process
- Not knowing container detection is default-on since JDK 10 / 8u191, or that cgroup v2 needs a newer JDK
- Setting `-XX:MaxRAMPercentage=100`, leaving no room for stacks, metaspace and native memory
- Ignoring thread count as a memory cost of roughly 1 MB of stack each