How would you diagnose and confirm a classloader leak in a running JVM?
answer
- Confirm metaspace ratchet, not heap-space OOM
- -Xlog:class+unload — silence after redeploy = pinned loader
- jmap -dump:live (forces GC first)
- MAT: count stale classloaders, path-to-GC-roots excluding weak refs
- Set -XX:MaxMetaspaceSize to fail fast
basics
~20 sWatch metaspace climb with each redeploy and never fall, ending in OutOfMemoryError: Metaspace. Then take a heap dump, find the classloaders that should be dead, and use the tool's 'path to GC roots' to see which reference is keeping each one alive.
solid answer
~40 sFirst confirm the symptom: monitor metaspace usage (jstat -gcmetacapacity/-gcmeta, JFR, or a metrics dashboard) and look for a saw-tooth that ratchets upward across redeploys, ending in 'OutOfMemoryError: Metaspace' rather than heap space. Enable '-XX:+HeapDumpOnOutOfMemoryError' or take a live dump with 'jmap -dump:live'. Open it in Eclipse MAT (or VisualVM), and look for multiple instances of the same web-app classloader that should have been collected — duplicate copies of your application classes are the giveaway. For each stale classloader, run 'path to GC roots' excluding weak/soft references; the resulting strong path identifies the exact lingering reference (a ThreadLocal, a static cache, a registered driver/listener, a running thread). You can also enable class-unloading logging ('-Xlog:class+unload') to confirm classes aren't being unloaded on redeploy. The fix is to sever that one reference.
go deeper
Knows to take a heap dump and look at metaspace metrics; may need guidance on interpreting paths.
Confirms the metaspace ratchet symptom, takes a live heap dump, and uses MAT path-to-GC-roots to find the reference.
Adds class-unload logging, distinguishes design-growth from true leaks, and validates the fix by re-running the redeploy loop.
Builds production observability (JFR/metaspace alerts), sets sane caps, and bakes leak-detection into the deploy pipeline.
## Step 1 — Confirm the symptom (metaspace, not heap) The defining signature of a classloader leak is **metaspace that grows and never recovers**, especially across redeploys. - **Monitor metaspace**: `jstat -gc <pid>` (look at the `MC`/`MU` metaspace capacity/used columns), or `jstat -gcmetacapacity`. Modern setups read it from JMX/Micrometer or **Java Flight Recorder (JFR)**. - **The fingerprint**: usage climbs in steps after each redeploy and never falls back. Eventually: `java.lang.OutOfMemoryError: Metaspace`. If instead you see 'Java heap space', it's a different (data) leak. - **Size the cap deliberately**: set `-XX:MaxMetaspaceSize` so the leak fails fast and visibly rather than slowly consuming native memory; without a cap, metaspace can grow until the OS kills the process. ## Step 2 — Watch class unloading directly - **`-Xlog:class+unload`** (JDK 9+; older: `-XX:+TraceClassUnloading`) logs each class as it is unloaded. After a redeploy you *expect* a burst of unloads of the old version's classes. **Silence is the smoking gun** — the old classes aren't being unloaded, i.e. the old loader is pinned. ## Step 3 — Take a heap dump - **On OOM automatically**: `-XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/path`. - **On demand (live objects only)**: `jmap -dump:live,format=b,file=heap.hprof <pid>`. The `live` flag forces a full GC first, so anything remaining is genuinely reachable — perfect for proving a leak. ## Step 4 — Analyze the dump Using **Eclipse MAT** (the standard tool) or **VisualVM**: 1. **Count classloaders**: look for **multiple live instances** of your app/web-app classloader class (e.g. `WebappClassLoader`, `ParallelWebappClassLoader`). After N redeploys you may see N+1 of them — each pinned generation is one leak. 2. **Spot duplicate classes**: the same application class loaded many times (once per stale loader) confirms the buildup. 3. **Path to GC roots**: select a stale classloader and run **'Merge Shortest Paths to GC Roots' → exclude weak/soft/phantom references**. The remaining **strong** path *is* the leak. Typical endings: a `Thread` (→ ThreadLocalMap → entry), a static field of a shared library, a `DriverManager` driver entry, a registered MBean/listener/shutdown hook. 4. MAT's **'Leak Suspects'** report and the **dominator tree** help quantify how much each loader retains. ## Step 5 — Cross-check live - **JFR** can record metaspace and old-object-sample events over time, useful in production where you can't easily pause for a dump. - **async-profiler** / allocation profilers help if the leak co-occurs with allocation churn, though metaspace leaks are primarily about retained class metadata, not allocation rate. ## Step 6 — Fix and verify Sever the single strong reference the path-to-GC-roots revealed (remove the ThreadLocal, stop the thread, deregister the driver/listener, null the static cache). Then **re-run the redeploy loop** and confirm with `-Xlog:class+unload` that the old classes now unload and metaspace returns to baseline. ## Common false leads - A genuinely *growing* class set (e.g. heavy runtime code generation, lots of dynamic proxies/lambdas, scripting) can fill metaspace **without** a classloader leak — distinguish 'too many classes by design' from 'old loaders not unloading' by checking whether usage *recovers* after redeploy. - A too-small `MaxMetaspaceSize` can OOM under legitimate load; verify the ratchet-across-redeploys pattern before blaming a leak.
- Why use jmap -dump:live rather than a plain dump?The live flag triggers a full GC before dumping, so only genuinely reachable objects remain. That proves a suspected-dead classloader is actually still reachable (a real leak) rather than just uncollected garbage, and reduces noise in the dump.
- How can you tell a metaspace OOM is a leak versus just too many classes?Check whether metaspace recovers after a redeploy/undeploy. A leak shows a monotonic ratchet across redeploys (old loaders never unload); a by-design large class count is high but stable and recovers when the app is removed.
saying these in an interview costs you the question
- Jumping to a heap-space diagnosis when the OOM message says Metaspace
- Not excluding weak/soft references in path-to-GC-roots (false roots)
- Confusing 'many classes loaded by design' with old loaders not unloading
- Taking a dump without -live, so unreachable garbage muddies the analysis