skip to content

Given an OutOfMemoryError in production, how do you read the message to identify the variant and choose the right first remediation for each?

level: middleimportance: must knowfreq 68%

answer

  1. Read the text after 'OutOfMemoryError:' first
  2. heap space + GC overhead → heap dump, leak vs size
  3. Metaspace → loaded-class count, ClassLoader leak, MaxMetaspaceSize (not -Xmx)
  4. native thread / array size / direct buffer = other resources
  5. Diagnose before touching any flag

basics

~30 s

Read the text right after 'OutOfMemoryError:'. 'Java heap space' means object memory is full — take a heap dump and check for a leak vs. a too-small heap. 'Metaspace' means too many classes are loaded — look for a classloader leak or raise MaxMetaspaceSize. 'GC overhead limit exceeded' means the collector is thrashing — treat it like a heap problem. The message names the region; that tells you where to look.

solid answer

~50 s

The string after `OutOfMemoryError:` is the primary diagnostic — it names the exhausted resource, which routes your investigation. **'Java heap space'**: the object heap is full. First step — enable `-XX:+HeapDumpOnOutOfMemoryError`, capture a dump, analyze dominator tree / retained size; a growing structure = leak (fix reachability), flat-but-full = under-size (raise `-Xmx`), one big operation = stream/paginate. **'Metaspace'**: class metadata exhausted; watch loaded-class counts — a classloader leak (climbs with redeploys) means a loader can't be unloaded (find path-to-GC-root on it), otherwise raise `-XX:MaxMetaspaceSize`. Note: raising `-Xmx` does nothing here. **'GC overhead limit exceeded'**: the throughput collector is thrashing (~98% GC, <2% reclaimed) — same root cause and fix path as heap space, caught earlier. Other messages exist too ('unable to create native thread', 'Requested array size exceeds VM limit', 'Direct buffer memory'), each pointing at a different resource. The discipline: identify the variant from the message, then diagnose leak vs. sizing before touching any flag.

code

java · 16 lines
java
// Capturing the variant + a dump is the first move, not guessing.
// Run with:
//   -XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/var/dumps
//   -Xlog:gc*:file=gc.log
//
// The thrown error's message routes you:
catch (OutOfMemoryError e) {
    String m = e.getMessage();            // e.g. "Java heap space"
    if (m != null && m.contains("Metaspace")) {
        // class-metadata / classloader-leak path: inspect loaded-class counts
    } else if (m != null && m.contains("heap space")) {
        // object-heap path: analyze the .hprof dominator tree
    }
    // ... then trigger a controlled shutdown so a supervisor restarts us
    throw e;
}

go deeper

for a junior

Knows to read the message after 'OutOfMemoryError:' and that 'heap space' vs 'Metaspace' are different problems needing different attention.

for a middle

Maps each common variant to its cause family and correct first step (heap dump for heap/GC-overhead; loaded-class counts and MaxMetaspaceSize for Metaspace), and knows -Xmx isn't a universal fix.

for a senior

Runs a disciplined triage: variant → dump/loaded-class analysis → leak vs. size vs. spike, distinguishes the lesser-known variants (native thread, direct buffer, array size), and resists flag-first fixes.

for a principal

Bakes this into runbooks and platform defaults: always-on heap dumps, GC logging, container-aware sizing, class-loading and thread observability, and alerting that routes each variant to the right remediation owner.

## The message is the map `OutOfMemoryError` always carries a **detail message** naming *which* resource ran out. Reading it is the single highest-leverage diagnostic step, because the fix for each variant is different — and applying the wrong one (e.g. raising `-Xmx` for a Metaspace problem) wastes time and masks the real issue. ## Variant 1 — `OutOfMemoryError: Java heap space` **Meaning:** the object heap (where `new` objects live, sized by `-Xmx`) is full of reachable objects and GC can't reclaim enough. **Causes:** leak (objects unintentionally still reachable — unbounded caches, un-deregistered listeners, ThreadLocals in pools); under-sizing (working set legitimately exceeds `-Xmx`); spike (loading a huge dataset/file at once). **First step:** capture a heap dump (`-XX:+HeapDumpOnOutOfMemoryError`, or on demand via `jmap`/`jcmd`), open in **Eclipse MAT / VisualVM**, look at **retained size** and the **dominator tree**, and compare dumps over time. Growing structure → fix reachability. Flat-but-full → raise `-Xmx`. Spike → stream/paginate. ## Variant 2 — `OutOfMemoryError: Metaspace` **Meaning:** the native-memory region holding **class metadata** (definitions of loaded classes; replaced PermGen in Java 8) is exhausted. **Causes:** too many classes loaded/generated (proxies, bytecode generation, scripting) or, the dangerous one, a **classloader leak** — a `ClassLoader` can't be unloaded because it (or one of its classes/instances/static fields) is still reachable, so its classes stay in Metaspace. Classic in app servers across hot redeploys. **First step:** watch **loaded vs. unloaded class counts** (`jstat -class`, JFR, `-verbose:class`). Rising loaded count tied to redeploys → loader leak: take a heap dump and run **path-to-GC-root** on the stuck `ClassLoader` to find the pinning reference (ThreadLocal, registered JDBC driver, JMX bean, spawned thread). If it's genuinely just many classes, raise **`-XX:MaxMetaspaceSize`**. *Raising `-Xmx` does nothing here* — different region. ## Variant 3 — `OutOfMemoryError: GC overhead limit exceeded` **Meaning:** the throughput (Parallel) collector detects it's spending **~98% of time in GC while reclaiming <2% of the heap** — i.e. thrashing — and fails fast. **Causes & first step:** same as 'Java heap space' (it's the same heap exhaustion caught earlier). Heap dump → leak vs. under-size → fix reachability or raise `-Xmx`. Disabling the check with `-XX:-UseGCOverheadLimit` only delays a harder crash. ## Other messages you may meet (don't confuse with the big three) - **`unable to create new native thread`**: not heap — the OS/process limits on threads (or native memory for thread stacks) are hit. Fix by reducing thread count (pool!) or raising OS limits, not `-Xmx`. (In fact a *smaller* per-thread `-Xss` or fewer threads helps.) - **`Requested array size exceeds VM limit`**: code tried to allocate an array larger than the JVM allows (near `Integer.MAX_VALUE`). A logic bug — bound the size. - **`Direct buffer memory`**: off-heap NIO `DirectByteBuffer` space exhausted; tune `-XX:MaxDirectMemorySize` or fix unreleased direct buffers. These share the `OutOfMemoryError` type but are *not* heap-space problems; the message is what distinguishes them. ## The decision flow ``` Read text after 'OutOfMemoryError:' ├── 'Java heap space' → heap dump → leak? under-size? spike? → fix reachability / -Xmx / stream ├── 'GC overhead limit exceeded' → same as heap space (caught earlier; don't just disable the check) ├── 'Metaspace' → loaded-class count → loader leak (path-to-GC-root) / else -XX:MaxMetaspaceSize ; NOT -Xmx ├── 'unable to create new native thread' → thread/OS limits, pool threads, native memory ├── 'Requested array size exceeds VM limit' → array-size logic bug └── 'Direct buffer memory' → off-heap NIO buffers / -XX:MaxDirectMemorySize ``` ## Key takeaways - The detail message names the exhausted resource — read it first. - Heap space & GC overhead → heap dump, then leak vs. size vs. stream. - Metaspace → loaded-class counts + find the stuck ClassLoader; tune MaxMetaspaceSize, not -Xmx. - 'native thread' / 'array size' / 'direct buffer' are different resources entirely. - Diagnose before changing any flag.

  • You see 'OutOfMemoryError: unable to create new native thread'. Why is raising -Xmx the wrong fix?
    Because it's not a heap problem at all — it's hitting OS/process thread limits or native memory for thread stacks. A bigger heap doesn't add native memory; ironically a huge heap can leave less native memory for stacks. Fix by pooling/limiting threads or raising OS limits.
  • Two services both crash with 'Java heap space' — one's heap climbs steadily, one's is high but flat. Same fix?
    No. The climbing one is a leak — find and fix the growing structure (a bigger heap only delays it). The flat-but-high one is under-sized — raising -Xmx (and aligning the container limit) is the legitimate fix.

saying these in an interview costs you the question

  • Reaching for -Xmx before reading the variant message.
  • Using -Xmx to fix a Metaspace or 'native thread' OOME (wrong resource).
  • Treating all OutOfMemoryError messages as the same heap problem.
  • Skipping the heap dump and guessing at the cause.
  • Assuming the OOME stack trace pinpoints the leak.

context