skip to content

How do you diagnose an OutOfMemoryError in production? Distinguish the common OOM kinds and the tools you'd use.

level: seniorimportance: should knowfreq 60%

answer

  1. Read the message: heap space / GC overhead / Metaspace / native thread / direct memory
  2. -XX:+HeapDumpOnOutOfMemoryError → .hprof → Eclipse MAT
  3. Dominator tree + retained size + path to GC roots
  4. GC logs: rising floor = leak; sawtooth = healthy
  5. Metaspace OOM = classloader leak; native thread = unbounded threads

basics

~20 s

Read the OutOfMemoryError message — it tells you which kind: heap space, Metaspace, GC overhead limit, or unable to create native thread. Capture a heap dump (-XX:+HeapDumpOnOutOfMemoryError), open it in a tool like Eclipse MAT, and find what's holding the most memory.

solid answer

~50 s

First, classify the OOM from its message, because each points at a different cause. 'Java heap space' means live objects exceed -Xmx — usually a leak (objects still reachable from a root) or genuine under-sizing. 'GC overhead limit exceeded' means GC is running constantly while reclaiming almost nothing — same root cause, caught earlier. 'Metaspace' means class metadata exhausted, typically a classloader leak (repeated redeploys, dynamic proxies). 'Unable to create new native thread' or 'Direct buffer memory' are native/off-heap, not the Java heap. For heap issues I enable -XX:+HeapDumpOnOutOfMemoryError to capture a dump at the moment of failure, then analyze it in Eclipse MAT or VisualVM: look at the dominator tree and 'leak suspects' to find which GC root retains the largest subgraph. I correlate with GC logs (-Xlog:gc) to see whether the heap grows monotonically (leak) or just spikes. Add monitoring (JFR, metrics) so it's caught before it crashes.

go deeper

for a junior

Knows OutOfMemoryError means memory ran out, that you can enable a heap dump and open it in a tool, and that bumping -Xmx isn't always the fix.

for a middle

Distinguishes heap vs Metaspace OOM, captures a heap dump, and uses MAT/VisualVM to find big objects; reads basic GC log growth.

for a senior

Classifies all the common OOM kinds from the message, uses dominator tree + retained size + path-to-roots to pinpoint the retaining root, correlates with GC logs, and proposes the right fix.

for a principal

Designs the observability up front (JFR, metrics, auto heap dump), reasons about container memory limits, off-heap/direct memory and classloader-leak patterns, and drives systemic fixes (bounded caches, pool sizing) plus capacity/SLO trade-offs.

## Start by reading the message — OOM is not one thing `OutOfMemoryError` (an `Error`, not an `Exception` — you generally don't catch it) is thrown for several *different* exhaustion conditions. The text after the colon tells you which, and each implies a different investigation: | Message | What ran out | Typical cause | |---|---|---| | `Java heap space` | The **heap** (object area), capped by `-Xmx` | Memory **leak** (objects still reachable from a GC root) or genuine under-sizing / load spike | | `GC overhead limit exceeded` | Effectively heap — JVM spends **>98% time in GC reclaiming <2%** | Same as above, detected earlier as a thrash signal | | `Metaspace` | **Metaspace** (class metadata, native) | **Classloader leak**: repeated hot-redeploys, dynamic proxy/CGLIB class generation, too many classloaders | | `Unable to create new native thread` | **OS threads / native memory**, not the heap | Thread leak / unbounded thread creation; OS ulimit reached | | `Direct buffer memory` | **Off-heap direct ByteBuffers** (`-XX:MaxDirectMemorySize`) | NIO/Netty buffers not released | | `Requested array size exceeds VM limit` | A single huge array allocation | Bad size computation | **Key distinction:** `OutOfMemoryError` is about exhausting a memory pool; `StackOverflowError` is a *different* error about a single thread's stack (deep recursion). Don't conflate them. ## Why heap OOM usually means a leak Recall reachability: the GC keeps every object **reachable from a GC root**. A **leak** in Java is an object that's still reachable (e.g. accumulating in a `static` collection, an unbounded cache, un-removed listeners, or `ThreadLocal`s on pooled threads) but no longer needed. The heap grows monotonically until it hits `-Xmx` → OOM. So diagnosing heap OOM = **finding which root retains a growing subgraph.** ## The toolchain ### 1. Capture a heap dump at the failure Set, in production, **before** it happens: ``` -XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/var/dumps ``` This writes an **`.hprof` snapshot of the whole heap** the moment the OOM is thrown — the most valuable artifact. You can also capture on demand with `jmap -dump:live,format=b,file=heap.hprof <pid>` or via `jcmd <pid> GC.heap_dump`. ### 2. Analyze the dump - **Eclipse MAT (Memory Analyzer Tool)** — the standard. Its **Leak Suspects** report and **Dominator Tree** show which objects **retain** (keep alive) the most memory, and the **path to GC roots** explains *why* they can't be collected. - **VisualVM / JDK Mission Control** — lighter, good for live inspection. Key concepts: **shallow size** (the object itself) vs **retained size** (everything that would be freed if it were collected). You hunt the object with huge *retained* size and trace its GC-root path. ### 3. Correlate with GC logs Enable `-Xlog:gc*:file=gc.log` (JDK 9+; older: `-XX:+PrintGCDetails`). A **monotonically rising used-heap-after-GC** baseline = leak; **sawtooth that returns to a stable floor** = healthy, just needs more heap or fewer allocations. ### 4. Continuous / low-overhead profiling **Java Flight Recorder (JFR)** + Mission Control gives always-on allocation and GC data (`-XX:StartFlightRecording`). Great for catching slow leaks before the crash, and for finding allocation hot spots. ### 5. For Metaspace / thread / direct-memory OOMs - **Metaspace**: watch loaded-class count (`jcmd ... VM.metaspace`, or metrics); a climbing class count across redeploys = classloader leak — find references pinning old classloaders. - **Native threads**: count threads (`jstack`, metrics); fix unbounded executors / missing pools, check OS `ulimit`. - **Direct memory**: track `BufferPool` MXBean; ensure buffers are freed/pooled. ## A pragmatic workflow 1. **Classify** from the message. 2. If heap: ensure `-XX:+HeapDumpOnOutOfMemoryError` was on; grab the `.hprof`. 3. Open in **MAT** → Leak Suspects → dominator tree → path to GC roots → identify the retaining root (often a static or a cache). 4. Confirm with **GC logs** (growth shape) and code review of that root. 5. Fix (bound the cache, remove/weaken the reference, size correctly), add a **metric/alert** and JFR so it's caught earlier next time. ## Tie-in to this topic Everything here rests on the platform internals: the **memory areas** (which pool the message names), **reachability + GC roots** (why a leak survives), and **reference types** (a `WeakReference`/bounded cache is often the *fix*). OOM diagnosis is the applied capstone of the JVM-memory concepts.

  • What does 'GC overhead limit exceeded' tell you that plain 'heap space' doesn't?
    It's an early-warning variant: the JVM is spending over ~98% of time in GC while reclaiming under ~2% of the heap, so it gives up before fully filling memory. It strongly signals a leak or severe under-sizing rather than a one-off spike — the heap is effectively full of live (reachable) objects.
  • In a heap dump, what's the difference between shallow and retained size, and which matters for leaks?
    Shallow size is the memory of the object itself; retained size is all memory that would be freed if that object were collected (its exclusively-owned subgraph). For leaks you look at retained size in the dominator tree — the object with huge retained size is what's holding the leak alive.

saying these in an interview costs you the question

  • Treating all OutOfMemoryError the same / just bumping -Xmx without diagnosing
  • Confusing OutOfMemoryError with StackOverflowError
  • Catching and swallowing OutOfMemoryError to 'keep running'
  • Looking only at shallow size instead of retained size / dominator tree
  • Assuming Metaspace OOM is a heap problem

context