skip to content

A long-running Java service shows steadily climbing heap usage and eventually OOMs. How do you confirm it's a leak and find the cause?

level: seniorimportance: should knowfreq 55%

answer

  1. Leak signature: post-GC (live) heap floor trends up, not sawtooth-to-baseline
  2. Corroborate with GC logs/JFR: old gen up, collections longer, reclaim less
  3. Capture: -XX:+HeapDumpOnOutOfMemoryError or jmap -dump:live
  4. Analyze in MAT: retained size + dominator tree + path to GC roots
  5. Root path names the cause; fix, then confirm the floor flattens

basics

~20 s

Watch the heap after garbage collection: if the live (post-GC) size keeps trending up over time instead of returning to a baseline, that's a leak, not just busy memory. Then take a heap dump, find the objects with the largest retained size, and trace what GC root is keeping them alive — that points at the offending reference.

solid answer

~50 s

First distinguish a leak from normal churn: plot the heap's live size right after full GCs. Healthy services sawtooth back to a stable baseline; a leak shows the post-GC floor creeping upward over hours or days until OOM. Confirm with GC logs or JFR — rising old-gen occupancy and increasingly frequent/long collections reclaiming little. To find the cause, capture a heap dump (jmap, or automatically with -XX:+HeapDumpOnOutOfMemoryError), then open it in Eclipse MAT or VisualVM. Look at the dominator tree and retained size to find which objects hold the most memory, then inspect the GC-root path to the leak suspect — that path names the long-lived holder (a static map, a listener list, a Thread's ThreadLocalMap). MAT's 'Leak Suspects' report often points straight at it. Reproduce, fix the retaining reference, and confirm the post-GC floor flattens.

go deeper

for a junior

Knows heap can grow and that a heap dump plus a tool like VisualVM helps find big objects; may not yet distinguish leak vs. busy.

for a middle

Distinguishes a leak (rising post-GC floor) from churn, can trigger a heap dump and open it, and finds large objects.

for a senior

Confirms via post-GC trend + GC logs, uses retained size, dominator tree, and path-to-GC-roots to name the holder, maps it to a code pattern, fixes, and verifies the floor flattens.

for a principal

Builds this into operations: standing heap/GC dashboards and alerts on live-heap growth, automatic OOM heap dumps shipped to storage, runbooks for triage, and load tests that catch retention regressions before release.

## Step 0: Is it actually a leak? High heap usage alone isn't a leak — a busy service legitimately uses memory. The defining signature of a **leak** is **monotonic growth of the *live* heap** — the amount still occupied *after* a full garbage collection. - A **healthy** heap **sawtooths**: it fills with garbage, a GC runs, and it drops back to a roughly **constant baseline** (the live set). Over time the baseline is flat. - A **leaking** heap sawtooths too, but the **post-GC floor creeps upward** run after run. The GC keeps running but can't get back to the old baseline because more and more objects are permanently retained. Eventually the floor reaches the ceiling → **`OutOfMemoryError: Java heap space`**. So: **measure the post-GC live size over time.** Upward floor = leak; flat floor = just sized tight or busy. ## Step 1: Corroborate with GC behavior Enable **GC logging** (`-Xlog:gc*`) or record a **Java Flight Recorder (JFR)** session. Tell-tale signs: - **Old generation** occupancy after each old/full GC trending up. - Collections becoming **more frequent and longer**, reclaiming **less** each time. - Approaching the dreaded **'GC overhead limit exceeded'** (JVM spends ~98% of time in GC reclaiming <2%). `jstat -gcutil <pid> <interval>` gives a quick live view of generation occupancy and GC counts without extra tooling. ## Step 2: Capture a heap dump A **heap dump** is a snapshot of every live object and reference at a moment in time. Get one by: - **On OOM automatically:** run with `-XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/path` so the JVM writes a dump exactly when it fails — the most representative moment. - **On demand:** `jmap -dump:live,format=b,file=heap.hprof <pid>` (the `live` option triggers a GC first, so the dump shows only retained objects — ideal for leak hunting). Dump when the heap is already large; an early dump may not show the leak clearly. ## Step 3: Analyze the dump Open the `.hprof` in **Eclipse MAT** (Memory Analyzer Tool) or **VisualVM**. Two concepts are central: - **Retained size:** the total memory that would be freed if this object were collected — i.e., everything it *exclusively* keeps alive. This, not shallow size, tells you who the real memory hogs are. - **Dominator tree:** a tree where each node 'dominates' (exclusively retains) its children. The big nodes near the top are your leak candidates. Workflow: 1. Sort by **retained size** / inspect the **dominator tree** → find the object(s) hoarding the heap (e.g. one `HashMap` retaining 1.2 GB across 4M entries). 2. Select the suspect → **'Path to GC Roots'** (in MAT, 'merge shortest paths to GC roots', excluding weak/soft refs). This shows *why* it's alive — the chain from a root to the object. That chain **names the leak**: `class X static field SESSIONS → HashMap → Node[] → Session...`, or `Thread → ThreadLocalMap → Entry → value`. 3. MAT's **'Leak Suspects' report** automates much of this and often points straight at the culprit class. ## Step 4: Map the root path back to code The GC-root path translates directly to a code smell: - root path through a **static field** → an unbounded static collection/cache. - root path through a **listener list** in a long-lived source → an unremoved listener. - root path through **`Thread` → `ThreadLocalMap`** → a ThreadLocal not `remove()`d in a pooled thread. - root path through an **unclosed stream/connection** holding buffers. ## Step 5: Fix and verify Drop the retaining reference (evict/bound the cache, deregister the listener, `ThreadLocal.remove()`, close the resource). Then **re-run under load and re-measure the post-GC floor** — a real fix makes it flatten. Comparing two heap dumps over time (a **'dominator delta'**) confirms the suspect's growth has stopped. ## Tooling cheat-sheet - **Live monitoring:** VisualVM, JFR, `jstat`, JMX heap metrics, async-profiler (alloc profiling). - **Dump capture:** `-XX:+HeapDumpOnOutOfMemoryError`, `jmap -dump:live`. - **Dump analysis:** Eclipse MAT (dominator tree, retained size, Leak Suspects, path-to-GC-roots), VisualVM. ## The principle Leak hunting is two questions: **(1) Is the live set growing?** (post-GC floor trend) and **(2) Who is holding it?** (largest retained size → path to GC roots). Answer both and the offending reference is named.

  • What is 'retained size' and why use it over shallow size?
    Shallow size is just the object's own memory; retained size is everything that would be freed if the object were collected — all the objects it exclusively keeps alive. Leaks are about deep retention, so the dominator with the largest retained size is the real culprit, even if its own shallow size is tiny.
  • Why does 'GC overhead limit exceeded' often accompany a leak?
    As the live set grows toward the heap ceiling, the GC runs more often and reclaims less. When it spends ~98% of time collecting and recovers <2% of the heap, the JVM throws 'GC overhead limit exceeded' — a near-exhaustion signal that frequently indicates an underlying leak rather than mere undersizing.

saying these in an interview costs you the question

  • Calling high heap usage a leak without checking the post-GC live trend
  • Reading shallow size instead of retained size to find the hog
  • Taking a heap dump too early when the leak hasn't accumulated
  • Forgetting to use 'live' so the dump is full of soon-to-be-collected garbage
  • Declaring it fixed without re-measuring the post-GC floor under load

context