When debugging a memory issue, when would you reach for `jmap -histo` versus capturing a full heap dump? What can each tell you that the other cannot?
answer
- histo = WHAT (flat counts/bytes), cheap
- dump = WHY (object graph), expensive
- Two histograms over time = leak trend
- MAT gives dominator tree + path to GC roots
- Retained size names the real owner
basics
~20 sUse -histo for a quick, cheap look at which classes have the most instances and bytes. Capture a full heap dump when you need to know why objects are kept alive — the reference chains and who is holding them — which only an analyzer working on a dump can show.
solid answer
~50 sA histogram (`jmap -histo:live <pid>`) is fast and light: it lists every class with instance counts and total bytes, so you immediately see *what* dominates memory — say, ten million `char[]` or a giant cache class. It is great for a first triage and for comparing two snapshots over time to spot a growing class. But it is **flat**: it cannot tell you *who retains* those objects or the reference paths keeping them alive, so it rarely identifies the leak's owner. A full heap dump (`-dump:live,format=b,file=...`) captures the entire object graph. Loaded into Eclipse MAT or VisualVM you get the **dominator tree** (which object keeps the most memory alive), **paths to GC roots** (the exact reference chain pinning a suspected object), and **retained size** (memory that would be freed if an object went away). The cost: a full GC plus a dump pause and a multi-gigabyte file. So: histogram to triage cheaply, dump when you must trace ownership and root cause.
go deeper
Knows -histo is the quick look and a dump is the deep look, even if not the precise reasons.
Can describe what each produces and that the dump is opened in MAT/VisualVM for deeper analysis.
Articulates the WHAT-vs-WHY split, retained vs shallow size, dominator tree and path-to-GC-roots, and a triage-then-escalate workflow with cost awareness.
Sets org-wide memory-diagnostic practice: histogram-diff monitoring, automatic OOM dumps, pause budgets/SLO impact, and analysis tooling/automation at scale.
## The two questions a memory investigation asks Memory debugging usually splits into **"what is filling the heap?"** and **"why is it being kept alive?"** jmap offers a cheap tool for the first and the raw material for the second. ## `jmap -histo` — the class histogram The **heap** is where Java objects live; the **garbage collector (GC)** frees objects no longer **reachable** (followable from a **GC root** such as a thread stack or static field). `jmap -histo <pid>` walks the heap and prints, per class: the number of instances and the total bytes they occupy, sorted largest first. Add `:live` (`-histo:live`) to run a GC first and count only reachable objects. What it is good for: - **Fast triage.** It is comparatively cheap and quick, so it is the first thing to run when memory looks high. - **Trend detection.** Take a histogram now and another 10 minutes later; a class whose count climbs steadily is a leak candidate. - **Confirming a suspicion.** "Are we really holding millions of `BigDecimal`?" — a histogram answers yes/no instantly. What it **cannot** do: - It is **flat**. It tells you `String[]` occupies 2 GB, but not *which* objects reference those arrays, nor the chain of references keeping them alive. You cannot find the *owner* of a leak from a histogram alone, because the real culprit is often a small object (one cache, one static map) that *retains* a huge subtree. ## The full heap dump — the object graph `jmap -dump:live,format=b,file=heap.hprof <pid>` writes the entire object graph to an **.hprof** file. Opened in **Eclipse MAT (Memory Analyzer Tool)** or **VisualVM**, it exposes structure a histogram cannot: - **Retained size:** for an object, the total memory that would be freed if it were collected — i.e., everything it *exclusively* keeps alive. A class can have small *shallow* size but huge *retained* size; that is the leak signature. - **Dominator tree:** ranks objects by retained size, so the single object responsible for the most memory floats to the top — usually the leak's owner. - **Path to GC roots:** for any suspected object, the exact reference chain back to a root that prevents collection. This is what actually *names the bug* — e.g., "this `byte[]` is held by an entry in a static `HashMap` in `CacheManager`." - **Duplicate strings, collection fill ratios, OQL queries** and more. The cost is real: `live` triggers a **stop-the-world full GC**, the dump itself pauses the app at a **safepoint**, and the file is roughly the size of the live heap (gigabytes). ## Decision rule 1. **Start with a histogram** — cheap, fast, often enough to confirm the suspect class and trend. 2. **Escalate to a full dump** when you must answer *why retained / who holds it* — the histogram has taken you as far as flat counts can. 3. In production, prefer comparing **two histograms** over taking a giant dump if pause budget is tight; reserve the full dump for when ownership analysis is unavoidable, and consider `-XX:+HeapDumpOnOutOfMemoryError` to capture one automatically at the moment of failure. ## Modern front end `jcmd <pid> GC.class_histogram` and `jcmd <pid> GC.heap_dump <file>` are the recommended equivalents on current JDKs.
- Why can a class with small shallow size still be the cause of a leak?Because its *retained* size is huge — a small object (e.g., a static map) can exclusively keep alive a large subtree of other objects. The histogram shows the big subtree's classes, but only the dump's dominator tree/retained size points at the small owner.
- How do you use a histogram to detect a leak without ever taking a full dump?Capture `-histo:live` at intervals and diff them; a class whose instance count grows monotonically over time is leaking, even though the histogram never shows the reference chain.
saying these in an interview costs you the question
- Believing a histogram can identify the leaking owner — it shows the retained subtree's classes, not who retains them
- Always taking a multi-GB dump first when a cheap histogram (or two) would triage faster
- Confusing shallow size (an object alone) with retained size (everything it keeps alive) — the latter finds owners