Compare GC.heap_dump, GC.class_histogram, and GC.run in jcmd. When would you reach for each, and what are the production cautions?
answer
- histogram = cheap class/instance/bytes table
- heap_dump = full HPROF file → MAT/VisualVM, find leak path
- GC.run = request full GC (== System.gc()), may be disabled
- Dump size ~ live heap; stop-the-world; check disk
- Two-step: histogram first, then targeted dump
basics
~20 sGC.class_histogram lists how many instances and bytes each class uses — a quick memory overview. GC.heap_dump writes a full snapshot file you analyze in a tool to find leaks. GC.run requests a garbage collection. All can pause the app, so use them carefully in production.
solid answer
~50 sThese three jcmd GC commands answer different memory questions. 'GC.class_histogram' prints a ranked table of every class with its live instance count and total bytes — a fast, cheap way to spot which type is dominating the heap. 'GC.heap_dump <file>' writes a complete HPROF snapshot of all objects and references, which you open in Eclipse MAT or VisualVM to chase a leak via dominator trees and GC roots. 'GC.run' politely requests a full garbage collection (equivalent to System.gc()), mainly to see what's retained after collection. Cautions: a full heap dump triggers a stop-the-world pause and writes a file as large as the live heap (potentially gigabytes) — it can stall a latency-sensitive service. The histogram is lighter but still causes a brief pause. GC.run can blow your latency budget and may be ignored if -XX:+DisableExplicitGC is set. In production, prefer histogram first, dump deliberately (off-peak, with disk space checked), and lean on JFR for continuous allocation profiling.
go deeper
Knows histogram shows counts/bytes per class, heap_dump writes a file, and GC.run triggers collection.
Picks the right command per question and knows a heap dump is big and pausing; can open an HPROF in a tool.
Runs the histogram-then-targeted-dump workflow, reasons about pause/disk cost, knows DisableExplicitGC, and reads dominator trees / GC-root paths to prove a leak.
Sets org-wide practice: HeapDumpOnOutOfMemoryError defaults, JFR-first allocation profiling, capacity for dump storage, and runbooks that avoid stalling peak traffic.
## The shared background: heap and GC The **heap** is the region of JVM memory where all Java objects live. The **garbage collector (GC)** automatically reclaims objects that are no longer reachable from **GC roots** (live stack variables, static fields, etc.). A **memory leak** in Java isn't forgotten `free()` calls — it's objects that *stay reachable* by accident (e.g. a growing static map) so the GC can never collect them. These three jcmd commands give you progressively heavier views into that heap. ## GC.class_histogram — the cheap overview ``` jcmd <pid> GC.class_histogram ``` Produces a table sorted by retained bytes: ``` num #instances #bytes class name 1: 4,210,113 168,404,520 [B (byte arrays) 2: 980,442 23,530,608 java.lang.String 3: 310,005 14,880,240 com.acme.Order ``` This instantly answers *"what kind of object is eating my heap?"*. If `com.acme.Order` is millions of instances when you expect thousands, you have your lead. It's relatively light because it walks the heap to tally classes but produces no file. It does force a safepoint (brief pause) and, by default, may trigger a GC so it reports *live* objects. ## GC.heap_dump — the full forensic snapshot ``` jcmd <pid> GC.heap_dump /var/tmp/app.hprof # live objects after a GC jcmd <pid> GC.heap_dump -all /var/tmp/app.hprof # ALL objects, no GC first ``` This writes an **HPROF** file containing every object, its fields, and every reference between objects. You then load it into a heap analyzer — **Eclipse MAT** (Memory Analyzer Tool) or **VisualVM** — and use: - the **dominator tree** (what would be freed if X died), - **GC-root paths** (why an object is still reachable — the leak path), - **retained size** (memory kept alive *because of* an object). This is how you actually *prove* a leak: find the object whose retained size grows, then trace its path to a GC root. The cost is real: it is **stop-the-world** for the duration and the file is roughly the size of the live heap — **gigabytes** on a big service. It can stall request handling and fill the disk. ## GC.run — request a collection ``` jcmd <pid> GC.run ``` Asks the JVM to perform a full GC — the diagnostic equivalent of `System.gc()`. Used to answer *"is this memory actually garbage, or genuinely retained?"*: run it, then take a histogram; whatever survives is truly reachable. Caveats: it's a deliberate **stop-the-world** pause that can violate latency SLOs, and if the JVM was launched with `-XX:+DisableExplicitGC` the request is **ignored**. It does not fix leaks — leaked objects are reachable, so GC won't reclaim them. ## Choosing between them | Question | Reach for | |---|---| | "Which class dominates the heap right now?" | `GC.class_histogram` (cheap, fast) | | "Why is this object retained / where's the leak path?" | `GC.heap_dump` + MAT/VisualVM | | "Is this memory garbage or genuinely live?" | `GC.run` then a histogram | ## Production cautions (the senior bit) - **Pauses:** all three reach a safepoint; the heap dump pause scales with heap size. Don't dump a 40 GB heap during peak traffic. - **Disk:** an HPROF can be as large as the live heap — check free space and a fast target path first; a failed/partial dump under disk pressure makes things worse. - **Two-step leak workflow:** histogram to identify the suspect type, *then* a targeted dump for the reference paths — avoids dumping blindly. - **Prefer continuous, low-overhead tooling:** JFR (`JFR.start`) records *allocation* profiles continuously at ~1% overhead, often letting you find the allocating call site without ever taking a giant dump. - **Automate the worst case:** `-XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=...` captures the dump exactly when an OOM happens, which is usually more useful than catching one by hand.
- You suspect a leak but don't want to dump a 30 GB heap on a live node. What's a safer sequence?Take a GC.class_histogram first (light) to identify the dominating type. If you must dump, do it off-peak, confirm disk space, write to fast local storage, or fail over the node first. Better still, use JFR allocation profiling continuously to find the allocating call site without a full dump.
- Why might GC.heap_dump still show objects that you think are garbage?By default GC.heap_dump dumps live objects after a GC; with -all it skips the GC and includes unreachable objects too. If even after a GC they remain, they are genuinely reachable — that's the leak, not garbage.
saying these in an interview costs you the question
- Thinking GC.run fixes a leak (leaked objects are reachable)
- Taking a multi-GB heap dump at peak with no disk check
- Assuming GC.run always runs (DisableExplicitGC ignores it)
- Treating the histogram as a full reference graph — it has no paths, only totals