How do you use a VisualVM heap dump to hunt a memory leak, and what does the live monitoring view show beforehand?
answer
- Healthy heap = sawtooth with flat baseline; leak = rising post-GC floor
- Leak = objects still reachable but no longer needed
- Sort by RETAINED size, not shallow size
- Path to GC root = what keeps it alive = the smoking gun
- Suspects: static collections, unbounded caches, unremoved listeners/ThreadLocals
basics
~20 sFirst the live memory graph hints at a leak: used heap keeps climbing across GCs and never settles. Then you capture a heap dump — a snapshot of every object — and look for which classes hold the most memory and what chain of references keeps them alive so the garbage collector can't free them.
solid answer
~50 sBefore dumping, VisualVM's (or jconsole's) memory tab tells the story: in a healthy app the heap sawtooths — rising as objects allocate, dropping after each garbage collection. A leak shows the post-GC floor creeping ever upward, until you approach OutOfMemoryError. To find it, take a heap dump: a snapshot of every live object and its references. In the dump you sort classes by retained size (memory that would be freed if that object went away), find the unexpectedly large set, then examine the GC roots / 'paths to root' — the reference chain keeping the objects reachable. The leak is usually that chain: a static collection, a cache without eviction, listeners never unregistered, or ThreadLocals on a pooled thread. Note a dump is heavy: capturing it can pause the app and the file can be gigabytes, so it's offline analysis, not live monitoring.
go deeper
Recognizes a leak as memory that keeps growing and that a heap dump shows the objects in memory.
Reads the sawtooth-with-rising-baseline signal and can open a dump to find the biggest classes by size.
Uses retained size and path-to-GC-root to pinpoint the keeping-alive reference and names the common leak patterns; understands dump cost as offline analysis.
Designs leak-detection into ops (HeapDumpOnOutOfMemoryError, MAT for huge dumps, JFR allocation profiling), prevents classloader/cache leaks architecturally, and balances capture cost against production impact.
## What a memory leak is in Java Java has a **garbage collector (GC)** that automatically frees objects no longer **reachable** — i.e. no live reference chain leads to them from a **GC root** (a thread stack, a static field, etc.). So Java can't 'leak' in the C sense of forgetting to free. A **Java memory leak** is subtler: objects you no longer *need* are still **reachable**, so the GC is obligated to keep them. They accumulate, the heap fills, and eventually you hit **OutOfMemoryError (OOM)**. ## Step 1 — spot it in live monitoring Open the **Monitor / Memory** tab (VisualVM or jconsole). Watch the **used heap** over time: - **Healthy:** a **sawtooth** — usage climbs as objects are allocated, then drops sharply at each GC back toward a stable baseline. The baseline (the post-GC low) stays roughly flat. - **Leaking:** the **post-GC floor trends upward** run after run — each collection reclaims less, the baseline creeps higher, GCs get more frequent and longer, until usage hugs the max and OOM looms. This is the cheap, live signal that says 'investigate'. It tells you *that* you leak, not *what*. ## Step 2 — capture a heap dump A **heap dump** is a complete snapshot of the heap: every live object, its fields, and its references, written to a file (`.hprof`). In VisualVM, right-click the process → **Heap Dump** (or it auto-grabs one on OOM if you set `-XX:+HeapDumpOnOutOfMemoryError`). Important properties: - It's **heavy**: the JVM typically **stops-the-world** while writing, and the file can be **gigabytes** (≈ live heap size). This is **offline analysis**, the opposite end of the spectrum from lightweight live monitoring. - For leaks, dumping *after* the heap has grown large makes the offender obvious. ## Step 3 — read the dump Key concepts when browsing: - **Shallow size:** memory of an object itself (its own fields). - **Retained size:** memory that would be **freed if this object were collected** — i.e. it plus everything only it keeps alive. **Retained size is what you sort by** to find the real hog; a small object can retain a huge subtree (e.g. a `HashMap` retaining millions of entries). - **Dominator:** an object through which *all* paths to a set of objects pass — it 'owns' that retained memory. - **GC root / path to root (path to GC root):** the reference chain from a GC root down to the suspect object. **This is the smoking gun** — it shows *what* is still pointing at the objects and thus preventing collection. Workflow: open the dump → look at **Classes** sorted by instance count / retained size → spot the type with surprisingly many or huge instances → pick an instance → **show 'path to GC root' (nearest roots)** → read the chain. The culprit is whatever at the top of that chain shouldn't still hold the reference. ## Common leak patterns the path reveals - **Static collection / singleton** that only ever `add`s (an ever-growing `static List`/`Map`). - **Cache without eviction or bounds** (use a bounded/`WeakHashMap`-based or `LinkedHashMap`-LRU cache, or a real cache lib). - **Listener / callback never unregistered** — the publisher keeps the subscriber alive. - **ThreadLocal not removed** on a **pooled** thread (thread-pool threads are long-lived, so a forgotten `ThreadLocal` value lives forever). - **ClassLoader leak** — common in app servers on redeploy; a lingering reference pins an entire old classloader and all its classes. ## Step 4 — confirm and fix Compare two dumps over time (before/after load) to confirm the suspect set *grows*. Fix the reference (bound the cache, unregister the listener, `remove()` the ThreadLocal), then re-run and confirm the live sawtooth returns to a flat baseline. ## Live vs offline — the recurring contrast Live monitoring (sawtooth watching) is cheap, continuous, and tells you *something* is wrong. Heap-dump analysis is expensive and point-in-time but tells you *exactly what and why*. Use the first to detect, the second to diagnose. (For lower-overhead production capture, JFR's allocation events and Eclipse MAT for big dumps are common companions.)
- Why is retained size more useful than shallow size for finding a leak?Shallow size is only the object's own fields; a leak hog is often a small container (e.g. a Map) retaining a vast subtree. Retained size measures everything that would be freed if it went away, so it surfaces the true owner of the memory.
- How can a ThreadLocal cause a leak even though it's meant for per-thread state?On a thread pool, threads are long-lived and reused. A value set in a ThreadLocal and never removed stays referenced by the living thread indefinitely, so it (and its retained graph) is never collected. Always remove() in a finally block on pooled threads.
saying these in an interview costs you the question
- Sorting by shallow size instead of retained size when hunting the hog
- Thinking the rising sawtooth peaks (not the baseline) prove a leak
- Believing Java can't leak because it has GC
- Ignoring that capturing a heap dump can pause the app and is gigabytes