How do you diagnose an OutOfMemoryError in production? Distinguish the common OOM kinds and the tools you'd use.
answer
- Read the message: heap space / GC overhead / Metaspace / native thread / direct memory
- -XX:+HeapDumpOnOutOfMemoryError → .hprof → Eclipse MAT
- Dominator tree + retained size + path to GC roots
- GC logs: rising floor = leak; sawtooth = healthy
- Metaspace OOM = classloader leak; native thread = unbounded threads
basics
~20 sRead the OutOfMemoryError message — it tells you which kind: heap space, Metaspace, GC overhead limit, or unable to create native thread. Capture a heap dump (-XX:+HeapDumpOnOutOfMemoryError), open it in a tool like Eclipse MAT, and find what's holding the most memory.
solid answer
~50 sFirst, classify the OOM from its message, because each points at a different cause. 'Java heap space' means live objects exceed -Xmx — usually a leak (objects still reachable from a root) or genuine under-sizing. 'GC overhead limit exceeded' means GC is running constantly while reclaiming almost nothing — same root cause, caught earlier. 'Metaspace' means class metadata exhausted, typically a classloader leak (repeated redeploys, dynamic proxies). 'Unable to create new native thread' or 'Direct buffer memory' are native/off-heap, not the Java heap. For heap issues I enable -XX:+HeapDumpOnOutOfMemoryError to capture a dump at the moment of failure, then analyze it in Eclipse MAT or VisualVM: look at the dominator tree and 'leak suspects' to find which GC root retains the largest subgraph. I correlate with GC logs (-Xlog:gc) to see whether the heap grows monotonically (leak) or just spikes. Add monitoring (JFR, metrics) so it's caught before it crashes.
go deeper
Knows OutOfMemoryError means memory ran out, that you can enable a heap dump and open it in a tool, and that bumping -Xmx isn't always the fix.
Distinguishes heap vs Metaspace OOM, captures a heap dump, and uses MAT/VisualVM to find big objects; reads basic GC log growth.
Classifies all the common OOM kinds from the message, uses dominator tree + retained size + path-to-roots to pinpoint the retaining root, correlates with GC logs, and proposes the right fix.
Designs the observability up front (JFR, metrics, auto heap dump), reasons about container memory limits, off-heap/direct memory and classloader-leak patterns, and drives systemic fixes (bounded caches, pool sizing) plus capacity/SLO trade-offs.
## Start by reading the message — OOM is not one thing `OutOfMemoryError` (an `Error`, not an `Exception` — you generally don't catch it) is thrown for several *different* exhaustion conditions. The text after the colon tells you which, and each implies a different investigation: | Message | What ran out | Typical cause | |---|---|---| | `Java heap space` | The **heap** (object area), capped by `-Xmx` | Memory **leak** (objects still reachable from a GC root) or genuine under-sizing / load spike | | `GC overhead limit exceeded` | Effectively heap — JVM spends **>98% time in GC reclaiming <2%** | Same as above, detected earlier as a thrash signal | | `Metaspace` | **Metaspace** (class metadata, native) | **Classloader leak**: repeated hot-redeploys, dynamic proxy/CGLIB class generation, too many classloaders | | `Unable to create new native thread` | **OS threads / native memory**, not the heap | Thread leak / unbounded thread creation; OS ulimit reached | | `Direct buffer memory` | **Off-heap direct ByteBuffers** (`-XX:MaxDirectMemorySize`) | NIO/Netty buffers not released | | `Requested array size exceeds VM limit` | A single huge array allocation | Bad size computation | **Key distinction:** `OutOfMemoryError` is about exhausting a memory pool; `StackOverflowError` is a *different* error about a single thread's stack (deep recursion). Don't conflate them. ## Why heap OOM usually means a leak Recall reachability: the GC keeps every object **reachable from a GC root**. A **leak** in Java is an object that's still reachable (e.g. accumulating in a `static` collection, an unbounded cache, un-removed listeners, or `ThreadLocal`s on pooled threads) but no longer needed. The heap grows monotonically until it hits `-Xmx` → OOM. So diagnosing heap OOM = **finding which root retains a growing subgraph.** ## The toolchain ### 1. Capture a heap dump at the failure Set, in production, **before** it happens: ``` -XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/var/dumps ``` This writes an **`.hprof` snapshot of the whole heap** the moment the OOM is thrown — the most valuable artifact. You can also capture on demand with `jmap -dump:live,format=b,file=heap.hprof <pid>` or via `jcmd <pid> GC.heap_dump`. ### 2. Analyze the dump - **Eclipse MAT (Memory Analyzer Tool)** — the standard. Its **Leak Suspects** report and **Dominator Tree** show which objects **retain** (keep alive) the most memory, and the **path to GC roots** explains *why* they can't be collected. - **VisualVM / JDK Mission Control** — lighter, good for live inspection. Key concepts: **shallow size** (the object itself) vs **retained size** (everything that would be freed if it were collected). You hunt the object with huge *retained* size and trace its GC-root path. ### 3. Correlate with GC logs Enable `-Xlog:gc*:file=gc.log` (JDK 9+; older: `-XX:+PrintGCDetails`). A **monotonically rising used-heap-after-GC** baseline = leak; **sawtooth that returns to a stable floor** = healthy, just needs more heap or fewer allocations. ### 4. Continuous / low-overhead profiling **Java Flight Recorder (JFR)** + Mission Control gives always-on allocation and GC data (`-XX:StartFlightRecording`). Great for catching slow leaks before the crash, and for finding allocation hot spots. ### 5. For Metaspace / thread / direct-memory OOMs - **Metaspace**: watch loaded-class count (`jcmd ... VM.metaspace`, or metrics); a climbing class count across redeploys = classloader leak — find references pinning old classloaders. - **Native threads**: count threads (`jstack`, metrics); fix unbounded executors / missing pools, check OS `ulimit`. - **Direct memory**: track `BufferPool` MXBean; ensure buffers are freed/pooled. ## A pragmatic workflow 1. **Classify** from the message. 2. If heap: ensure `-XX:+HeapDumpOnOutOfMemoryError` was on; grab the `.hprof`. 3. Open in **MAT** → Leak Suspects → dominator tree → path to GC roots → identify the retaining root (often a static or a cache). 4. Confirm with **GC logs** (growth shape) and code review of that root. 5. Fix (bound the cache, remove/weaken the reference, size correctly), add a **metric/alert** and JFR so it's caught earlier next time. ## Tie-in to this topic Everything here rests on the platform internals: the **memory areas** (which pool the message names), **reachability + GC roots** (why a leak survives), and **reference types** (a `WeakReference`/bounded cache is often the *fix*). OOM diagnosis is the applied capstone of the JVM-memory concepts.
- What does 'GC overhead limit exceeded' tell you that plain 'heap space' doesn't?It's an early-warning variant: the JVM is spending over ~98% of time in GC while reclaiming under ~2% of the heap, so it gives up before fully filling memory. It strongly signals a leak or severe under-sizing rather than a one-off spike — the heap is effectively full of live (reachable) objects.
- In a heap dump, what's the difference between shallow and retained size, and which matters for leaks?Shallow size is the memory of the object itself; retained size is all memory that would be freed if that object were collected (its exclusively-owned subgraph). For leaks you look at retained size in the dominator tree — the object with huge retained size is what's holding the leak alive.
saying these in an interview costs you the question
- Treating all OutOfMemoryError the same / just bumping -Xmx without diagnosing
- Confusing OutOfMemoryError with StackOverflowError
- Catching and swallowing OutOfMemoryError to 'keep running'
- Looking only at shallow size instead of retained size / dominator tree
- Assuming Metaspace OOM is a heap problem