You have two heap snapshots of a search-indexing service taken an hour apart; how do you diff them to find what is growing?
answer
- an instant cannot show a trend
- same phase, same settings, both warm
- aggregates, never object addresses
- type count up, holder count flat
- confirm on one root path
basics
~20 sCompare aggregates, not individual objects: per-type instance counts and shallow bytes, then retained size along the same root paths. The growing type plus a holder count that stayed flat points at one container accumulating entries.
solid answer
~40 sA single snapshot shows a big structure but cannot say whether it is growing. Two snapshots an hour apart can, provided you diff them the right way. Compare **aggregates**: instances and total shallow bytes per type, then the retained size of the same dominator paths in both files. Object identity is the wrong axis — a collector that relocates surviving objects invalidates addresses, so you match by type, by path from a root, and by counts. The signature you are hunting is a type whose instance count climbed sharply while the number of *holders* stayed the same: one container is accumulating and nothing removes from it. Take both snapshots the same way, at the same point in the service's cycle, and after a collection, so that short-lived garbage does not swamp the delta.
go deeper
The idea to hold on to: one snapshot shows size, two snapshots show change. Change is what a leak looks like; size on its own is often just a big, correct data structure.
Explain why the comparison is done on aggregates — instances and bytes per type — rather than on individual objects, and what a relocating collector does to object addresses between two captures.
Show the capture discipline before the comparison: warm, same phase, after a collection, same settings. Then read the diff in order and name the signature — rising instance count with a flat holder count — and what you sample next.
Own the uncertainty: two snapshots cannot separate a cache filling toward its bound from an unbounded one, and capture itself perturbs the system. Decide how often the fleet is sampled and what evidence is required before an engineer is sent after a suspect.
## Why one snapshot is not enough A snapshot is an instant. It tells you what is large, not what is growing, and large is not a defect — a search index is supposed to hold a lot. The question a diff answers is different: *between these two moments, what got bigger, and by what shape?* That difference is what separates a genuine accumulation from a structure that has been the same size since startup. ## Capture discipline first A sloppy pair of snapshots produces a diff full of noise. Before comparing: - **Take both after a collection.** Otherwise the delta is dominated by short-lived objects that were merely uncollected at capture time, and every type appears to grow. - **Take both at the same point in the service's cycle** — the same phase of indexing, not one mid-merge and one idle — or the diff measures the phase, not the hour. - **Use the same capture settings** for both, including how references of clearable strength are treated, so retained figures are comparable. - **Let the service be warm.** The first minutes of a process fill caches and buffers that will never grow again; a diff spanning warm-up reports that filling as growth. ## Diff on aggregates, not identities It is tempting to ask which *objects* are new. Usually you cannot: - A collector that relocates surviving objects changes their addresses between captures, so an address is not a stable identity across an hour. - Even where addresses are stable, millions of new small objects per second make an object-level diff unreadable. So compare summaries: | Axis compared | What a rise means | |---|---| | Instances per type | more objects of that shape are alive than before | | Shallow bytes per type | the same objects got fatter, or there are simply more | | Holder count for that type | more containers exist, rather than one filling up | | Retained size along one root path | that specific subgraph grew, not the type everywhere | The last row is the one that turns a diff into a finding. A type can be up 30% for entirely legitimate reasons spread over the whole service; a single dominator path whose retained size went from 300 MB to 1.1 GB in an hour is a place. ## The signature to look for Run the comparison and read it in this order: 1. **Types sorted by delta in instance count.** Ignore the ones that scale with traffic in both directions. 2. **Holder count for the top type.** Flat holder count plus a rising instance count means one container is filling. Rising holder count means you are creating containers and should look one level up instead. 3. **The same root path in both files.** Confirm the growth lands under a specific dominator, not smeared across the heap. 4. **A sample of the newly abundant instances.** What do they have in common — a request identifier, a timestamp range, a document that should have been evicted? That common property usually names the code path that adds without removing. ## Reading the result honestly A few traps worth stating explicitly: - **Growth is not proof of a defect.** A cache filling toward its bound grows for an hour and then stops. Two snapshots cannot distinguish that from an unbounded structure; a third, later pair can. - **A flat diff is not proof of health.** An hour may be too short for a slow accumulation, and a capture that stalls the process can itself change what the next hour looks like. - **Absolute deltas beat percentages** when the base is small: a type going from 4 to 40 instances is a 900% rise and 36 objects. - **Compare like with like.** Two instances of a service that handle different shard sizes will differ in every aggregate; diff an instance against *itself*, an hour later. Done with that discipline, a two-snapshot diff answers a narrow question very well: which type, under which holder, is accumulating — which is exactly the input the dominator-path walk needs to point at a line of code.
- The diff shows a type up 40% but its holder count also went up 40%. What does that change?It moves the investigation one level up. The entries are not accumulating inside one container; containers themselves are being created and kept. Diff the holders' type next, and find what retains those containers — the accumulating structure is above them, not below.
- Both snapshots look nearly identical, yet the service's footprint climbed all hour. What would you check?Whether the growth is even in the snapshot's scope. A snapshot covers the managed object graph; memory held outside it — buffers owned by native code, runtime metadata, allocator caches, thread stacks — never appears in a diff of object graphs, and a climb with a flat object graph points there.
- Why not just diff the two files by object address and list the new objects?Because an address is not a stable identity. A collector that relocates survivors moves them between captures, so old objects reappear as new ones. Even with stable addresses the list is dominated by ordinary churn. Aggregates by type, holder and root path stay meaningful across relocation.
saying these in an interview costs you the question
- Diffs snapshots by object address across a relocating collector
- Captures one snapshot mid-warm-up and calls the filling a leak
- Treats any type that grew over an hour as the defect
- Compares two different instances instead of one instance twice
- Quotes percentage growth on a base of four objects