When analyzing a heap dump, what is the difference between shallow size and retained size, and how does a dominator tree help you find a leak?
answer
- Shallow = the object's own fields only
- Retained = everything freed if it dies (itself + what only it keeps alive)
- Leak signal: tiny shallow, huge retained
- Dominator: A dominates B if every GC-root path to B goes through A
- Sort dominator tree by retained size, then 'path to GC roots'
basics
~20 sShallow size is the memory one object uses by itself. Retained size is all the memory that would be freed if that object were removed — the object plus everything only it keeps alive. A dominator tree shows which objects hold the most memory, so you look at the biggest retained sizes to find the leak.
solid answer
~50 sShallow size is the memory occupied by a single object's own fields, ignoring what it references. Retained size is the total memory that becomes reclaimable if that object is garbage-collected — itself plus the entire set of objects reachable only through it. A leak culprit usually has a huge retained size while its shallow size is tiny (e.g. one HashMap whose shallow size is bytes but retained size is gigabytes). The dominator tree captures this: object B is dominated by A if every path from a GC root to B passes through A, so A's retained set is exactly the subtree it dominates. Tools like Eclipse MAT build this tree and sort by retained size, letting you jump straight to the few objects that dominate most of the heap, then inspect the 'paths to GC roots' to learn why they are still reachable.
go deeper
Knows shallow size is the object alone and retained size is bigger because it includes what the object holds onto; knows a tool sorts by retained size.
Can explain retained size as 'memory freed if this object dies' and use MAT/VisualVM to sort the dominator tree and open paths to GC roots.
Fluently defines domination (every GC-root path passes through A), reads a dominator tree to isolate a culprit, and traces the reference chain to the root cause of a leak.
Builds the team's leak-analysis playbook, reasons about why counts mislead vs. retained size, interprets weak/soft-reference exclusions, and correlates dump findings with GC logs and live profiling to confirm root cause.
## The problem heap analysis solves When the heap is full, you need to find *which objects* are hogging memory and *why they are still alive*. Raw object counts mislead you — a million tiny strings might matter less than one giant cache. The concepts of **shallow size**, **retained size**, and the **dominator tree** exist to answer this precisely. ## Reachability and GC roots (the foundation) The garbage collector keeps an object alive only if it is **reachable**: there is a chain of references from a **GC root** to it. GC roots are the starting points the GC always considers live — e.g. local variables on thread stacks, static fields of loaded classes, and JNI references. An object becomes collectible the instant no GC root can reach it. A **memory leak** in Java is therefore *unintended reachability*: an object you are done with is still reachable from some root and so never freed. ## Shallow size The **shallow size** of an object is the memory its own instance occupies: the object header plus its own fields. A reference field counts only as the pointer (a few bytes), **not** the object it points to. So an `ArrayList` with 100,000 elements has a small shallow size — its fields are just a length and a pointer to a backing array. ## Retained size The **retained size** of object X is the total amount of memory that would be freed if X were garbage-collected — that is, X's shallow size **plus** the shallow sizes of every object that is kept alive *only* through X (the objects that would also become unreachable once X is gone). This is the number you care about for leaks: it tells you how much memory removing this one object would actually reclaim. The gap between the two is the key signal. A leaking object often has a tiny shallow size but an enormous retained size — for example, a single static `Map` field (shallow size: a few bytes) that retains gigabytes of entries that should have been removed. ## The dominator tree To compute retained sizes efficiently and present them, analyzers build a **dominator tree**. Borrowed from graph theory: in the object graph (rooted at the GC roots), object A **dominates** object B if *every* path from a GC root to B goes through A. Intuitively, A is a mandatory gatekeeper for B — remove A and B becomes unreachable. The dominator tree links each object to its **immediate dominator** (the closest such gatekeeper). Key property: **the retained set of A is exactly the subtree it dominates.** So once the tree is built, the retained size of any node is just the total size of its subtree. Tools sort the top-level dominators by retained size, and the few biggest entries are your suspects. ## The analysis workflow (e.g. in Eclipse MAT) 1. Open the `.hprof` dump in **Eclipse Memory Analyzer Tool (MAT)** or VisualVM. 2. Look at the **dominator tree**, sorted by **retained size** descending. The objects retaining the most heap are at the top. 3. For a suspect with a huge retained size, run **'path to GC roots'** (in MAT, usually excluding weak/soft references) to see *why* it is still reachable — this exposes the reference chain holding the leak (e.g. `static ServiceRegistry -> HashMap -> ... -> your object`). 4. MAT's **Leak Suspects report** automates much of this: it flags dominators that account for an unusually large share of the heap. ## A concrete leak shape Classic example: a `static final Map<Key, Value> cache` you only ever `put` into and never evict. Each entry's `Key` is reachable from the static map, which is reachable from a GC root (the class). The map's shallow size is trivial, but its retained size grows without bound. In the dominator tree it appears as a single node retaining most of the heap, and 'path to GC roots' shows the static field — pointing you straight at the unbounded cache. ## Why not just count objects? Object *counts* (the histogram, e.g. `jmap -histo`) tell you there are 5 million `char[]` instances but not who is keeping them alive. Retained size + dominator tree connect *what* is large to *who* owns it, which is what you need to actually fix the leak.
- Why can two objects' retained sizes overlap, and how does the dominator tree avoid double-counting?If an object is reachable through two independent owners, neither alone retains it (removing one leaves the other path), so it belongs to neither's retained set — it is dominated only by a common ancestor higher up. The dominator tree assigns each object to its single immediate dominator, so retained sizes partition the heap without double-counting.
- When taking a dump to hunt a leak, why prefer a 'live' dump (jmap -dump:live)?It forces a full GC first, so the dump excludes already-unreachable garbage. That removes noise so the dominator tree reflects only genuinely retained objects — the ones a leak is actually keeping alive.
saying these in an interview costs you the question
- Confusing shallow and retained size — quoting shallow size as the memory a collection 'uses'.
- Hunting leaks by object count (histogram) alone, which shows what is large but not who retains it.
- Thinking the object with the largest shallow size is the leak; the culprit usually has a small shallow but huge retained size.
- Not running 'path to GC roots', so you find the big object but never learn why it is still reachable.