skip to content

Walk through how you would diagnose a slow memory leak in a long-running production Java service, from first symptom to root cause.

level: seniorimportance: should knowfreq 46%

answer

  1. Confirm: post-GC baseline creeps up (sawtooth floor rising) → leak, not load
  2. Capture: -XX:+HeapDumpOnOutOfMemoryError and/or two dumps to diff
  3. Analyze: MAT Leak Suspects + dominator tree sorted by retained size
  4. Root cause: path to GC roots → static collection / listener / ThreadLocal / cache
  5. Fix then verify: flat baseline under load

basics

~20 s

Confirm the leak by watching memory climb over time and never fully recover after GC. Enable an automatic heap dump on OutOfMemoryError, or take dumps at intervals and compare. Open the dump in a tool like Eclipse MAT, sort by retained size, find the object retaining the most, and trace its path to GC roots to see what code keeps it alive.

solid answer

~50 s

First, confirm it is actually a leak: watch the heap trend (GC logs, jstat, or JFR) — a leak shows old-generation occupancy climbing across full GCs and never returning to baseline, eventually heading toward OutOfMemoryError. Second, capture evidence with minimal disruption: set -XX:+HeapDumpOnOutOfMemoryError so the failing instance dumps itself, and/or take two live dumps hours apart so you can diff growth. Third, analyze: open the .hprof in Eclipse MAT, run the Leak Suspects report, and sort the dominator tree by retained size — the culprit usually has a small shallow but huge retained size. Fourth, find the cause: 'path to GC roots' reveals the reference chain (often a static collection, an unbounded cache, an un-deregistered listener, or a ThreadLocal in a pooled thread). Fifth, fix and verify: bound the cache / deregister / clean up the ThreadLocal, then re-run under load and confirm the heap returns to baseline after GC. Throughout, prefer low-overhead tools (JFR) on production and do heavy MAT analysis offline.

code

java · 11 lines
java
// A classic leak: an unbounded static cache reachable from a GC root.
public final class SessionCache {
    // static field => reachable from a GC root for the JVM's lifetime
    private static final Map<String, Session> CACHE = new HashMap<>();

    public static void put(String id, Session s) {
        CACHE.put(id, s); // never evicted -> retained size grows forever
    }
    // In a dump: CACHE has tiny shallow size but enormous retained size,
    // and 'path to GC roots' points straight at this static field.
}

go deeper

for a junior

Knows the rough loop: see memory climb, take a heap dump, open it in a tool, find the big object. May not yet distinguish leak vs. undersized heap.

for a middle

Confirms via GC trend, captures a dump with the OOM flag, uses MAT to sort by retained size, and names common leak patterns.

for a senior

Runs the full disciplined workflow including dump-diffing, path-to-GC-roots root-causing, low-overhead production handling, and verifying the fix under load.

for a principal

Establishes the org's leak-response runbook and guardrails (auto-dump config, dump retention/redaction, JFR always-on), prevents leaks via review patterns (bounded caches, ThreadLocal hygiene), and mentors on trend-vs-snapshot reasoning.

## What a Java memory leak actually is Java has a garbage collector, so it cannot leak the way C does (forgetting to `free`). A **Java leak** is *unintended reachability*: an object you are finished with is still reachable from a **GC root** (a thread stack, a static field, etc.), so the collector is obligated to keep it. Over time these accumulate, retained heap grows, GC works harder for less, and eventually the JVM throws `OutOfMemoryError: Java heap space`. A disciplined diagnosis has five phases. The goal is to move from a vague 'it runs out of memory' to a specific line of code keeping objects alive. ## Phase 1 — Confirm it is a leak (not just a big workload) Not every OOM is a leak; sometimes the heap is simply undersized for the load. Distinguish them by the **trend after garbage collection**: - Turn on **GC logging** (`-Xlog:gc*` on modern JVMs) or watch live with `jstat -gcutil <pid> 1000`, or record with **JFR**. - A healthy service has a **sawtooth**: occupancy rises, a GC drops it back near a stable baseline, repeat. The baseline is flat over hours. - A **leak** shows the **post-GC baseline creeping upward** — even **full GCs** can no longer return old-gen to its earlier low. Extrapolated, the line runs into the heap ceiling. That rising floor is the signature of a leak. ## Phase 2 — Capture evidence with minimal disruption A heap dump is the object-level evidence, but capturing one pauses the app and the file is heap-sized, so be deliberate: - **Best, hands-off:** start the service with `-XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/var/dumps`. The instance that finally OOMs writes a dump at the exact moment of failure — the most informative snapshot, and one you usually can't reproduce manually. - **Diff approach:** take **two live dumps** several hours apart (`jcmd <pid> GC.heap_dump` or `jmap -dump:live`). Comparing them shows which object populations *grew*, which is often clearer than one snapshot because it isolates the *accumulating* objects from the merely-large-but-stable ones. - Do this on **one** instance behind the load balancer if possible, to absorb the pause; or take the node out of rotation first. ## Phase 3 — Analyze the dump Move the `.hprof` off the production box and open it in **Eclipse Memory Analyzer Tool (MAT)** (handles large dumps; VisualVM works for smaller ones): 1. Run MAT's **Leak Suspects** report — it automatically flags dominators holding a disproportionate share of the heap. 2. Open the **dominator tree** and **sort by retained size** descending. **Retained size** = the memory freed if that object died (itself plus all objects only it keeps alive). The leak culprit typically has a **tiny shallow size but a huge retained size** (e.g. one `HashMap` retaining gigabytes). 3. If you took two dumps, compare **histograms** to see which classes' instance counts grew between snapshots. ## Phase 4 — Find the root cause (path to GC roots) Finding the big object isn't enough — you must learn *why it is still reachable*. Run **'path to GC roots'** (in MAT, exclude weak/soft references so you see the strong chain holding it). This prints the reference chain from a root to the object, e.g. `class AppRegistry (static) -> HashMap -> Node[] -> SessionData`. That chain points at the offending code. The usual culprits map to known **leak patterns**: - **Unbounded static collection / cache** that is only ever added to. - **Listeners/callbacks** registered but never deregistered. - **ThreadLocal** values not removed on pooled (reused) threads, so they live as long as the thread pool. - **Unclosed resources** holding buffers, or a custom classloader pinned by a stray reference (a metaspace leak). ## Phase 5 — Fix and verify Apply the targeted fix: bound the cache (size or time eviction, or a `WeakHashMap`/`SoftReference` if appropriate), deregister the listener, call `ThreadLocal.remove()` in a `finally`, or close the resource with try-with-resources. **Then verify**: re-run under representative load with GC logging on and confirm the **post-GC baseline is flat again** over a sustained period. A fix you don't verify under load is a guess. ## Cross-cutting principles - **Low overhead on production, heavy analysis offline.** Monitor live with JFR/jstat (cheap); do MAT analysis on a copied dump on a workstation. - **Trend beats snapshot for confirming** a leak; **snapshot (dump) beats trend for locating** it. Use both. - **Reproduce in staging if you can** — it lets you iterate on the fix without risking production.

  • Why are two heap dumps taken hours apart often more useful than a single dump?
    A single dump shows everything large, mixing stable big structures with the slowly-accumulating leak. Diffing two dumps isolates the object populations that *grew* between them, which points directly at the leaking objects rather than at legitimately large but stable ones.
  • How does the GC-trend graph distinguish a leak from an undersized heap?
    An undersized heap has a stable post-GC baseline — it just runs close to the ceiling under load and may OOM at peaks, but GC returns it to the same floor. A leak shows the post-GC floor itself creeping upward over time, even after full GCs, trending into the ceiling regardless of load.

saying these in an interview costs you the question

  • Jumping straight to a heap dump without confirming the post-GC baseline is actually rising (could be undersized heap, not a leak).
  • Finding the biggest object but never running 'path to GC roots', so you miss why it is retained.
  • Doing heavy MAT analysis on the production box instead of copying the dump off and analyzing offline.
  • Shipping a fix without re-running under load to confirm the baseline flattens.
  • Assuming every OutOfMemoryError is a leak; some are simply too much live data for the configured heap.

context