skip to content

A JVM shows frequent young collections that each promote a large volume of data, and old-generation occupancy climbs over hours even though the application's steady-state live set is stable. How do you confirm this is premature promotion rather than a leak, and what causes it and fixes it?

level: seniorimportance: should knowfreq 44%

answer

  1. flat post-GC floor ⇒ promotion, rising floor ⇒ leak
  2. threshold pinned at 1 + fat age-1 bucket
  3. lifetime measured in bytes allocated, not seconds
  4. fix order: allocate less → bigger young → bigger survivors → threshold
  5. promoted bytes per young GC is the key metric

basics

~20 s

Premature promotion means short-lived data is pushed into the old generation before it dies. Confirm it from GC logs: an age histogram with a collapsed tenuring threshold, large promoted bytes per young collection, and old-generation occupancy that drops back fully after an old/full collection — which distinguishes it from a leak, where occupancy after collection keeps rising.

solid answer

~60 s

**Confirm, don't guess.** Enable `-Xlog:gc*,gc+age=trace` and look at three things: the tenuring distribution (a threshold pinned at 1–2 with most bytes in the age-1 bucket), the amount promoted per young collection, and — decisively — the old-generation occupancy *after* each old or full collection. If post-collection occupancy returns to a flat baseline, the old generation is filling with garbage that arrived too early: premature promotion. If the post-collection floor rises monotonically, it is a leak and no GC tuning will fix it. **Causes.** Survivor spaces too small for the surviving volume, so overflow tenures regardless of age; the adaptive threshold collapsing under allocation bursts; a young generation too small relative to allocation rate, so collections fire before short-lived data has died; and workloads with a genuine medium-lifetime hump — request-scoped buffers, batch working sets — that outlive a young collection but not much longer. **Fixes,** in order of leverage: reduce or shorten allocation; enlarge the young generation so objects get time to die in eden; enlarge survivor space so the aging filter works; only then touch the tenuring threshold.

code

text · 3 lines
text
-Xlog:gc*,gc+age=trace,gc+heap=info:file=gc.log:time,uptime,level,tags
# look for: promoted bytes per young GC, "new threshold N (max 15)",
# and old-gen occupancy immediately AFTER each old/full collection

go deeper

for a junior

Recognise the term: short-lived objects reaching the old generation before they die, which makes expensive collections more frequent.

for a middle

Name the two mechanical causes — survivor overflow and a young generation too small for the allocation rate — and know that the age histogram in the GC log shows it.

for a senior

Run the full diagnosis: separate leak from promotion by the post-collection floor, read the tenuring distribution and promoted bytes, then fix in leverage order starting from allocation itself.

for a principal

Decide whether the workload's lifetime distribution fits a generational collector at all, and weigh sizing changes against pause budget, heap cost, and the engineering cost of reducing allocation in the application.

## What premature promotion is Promotion is meant to be a verdict: this object has survived enough young collections that it is probably long-lived, so move it to the region collected rarely. *Premature* promotion is that verdict being reached for objects that were about to die anyway. The consequence is not a correctness problem — it is a cost problem. Garbage that would have been reclaimed for free by a young collection is instead copied into the old generation, where reclaiming it requires the expensive collection the generational design exists to avoid. The symptom pattern is characteristic: young collections are frequent and each reports substantial promotion; old-generation occupancy climbs steadily between old collections; old or full collections become progressively more frequent; and the application's actual retained data is not growing. ## Step one: separate it from a leak These look identical on a naive heap-usage graph — both show a rising old generation. They are told apart by the **post-collection floor**: - **Premature promotion:** occupancy after each old/full collection returns to roughly the same baseline. The heap fills with garbage and the garbage is genuinely collectible. The sawtooth's *bottom* is flat. - **Memory leak:** the bottom of the sawtooth trends upward. Data is being retained, and the collector is behaving correctly. This single check should always precede tuning, because the fixes are completely different: a leak needs a code change; no combination of flags will help it. ## Step two: read the tenuring distribution With `-Xlog:gc+age=trace` the JVM prints per-collection which ages survived and what threshold it chose. Diagnostic patterns: - **Threshold pinned at 1** with a large age-1 bucket, and a desired-survivor-size far below the surviving total: survivors do not fit; overflow is going straight to old. The aging filter is not running at all. - **Threshold healthy but promoted bytes still large:** a genuine long-lived working set is being built — perhaps a cache warming up. That may be correct behaviour, not a defect. - **A long tail out to age 15:** the opposite problem — long-lived data being copied fifteen times before tenuring, inflating every young pause. Pair this with promoted-bytes-per-collection and young-pause duration, which in a copying collector tracks surviving volume. ## The four common causes 1. **Survivor space too small.** The most common. Overflow tenures data regardless of age, and the adaptive threshold collapses on top of it, compounding the problem. 2. **Young generation too small relative to allocation rate.** Object lifetime is measured in *bytes allocated*, not seconds: if a request holds objects for 50 ms and the process allocates a young generation's worth every 20 ms, collections fire while that data is still live and it is repeatedly copied and eventually promoted. Enlarging the young generation converts "survives a collection" back into "dies in eden". 3. **Allocation bursts.** A steady state that is fine collapses under a spike — a large batch import, a fan-out of parallel requests — producing one enormous survivor set that overflows and tenures wholesale. 4. **A genuine medium-lifetime hump.** Request-scoped buffers, per-batch working sets, pooled objects with a short residence time. These are neither short- nor long-lived, and the survivor spaces are exactly the structure meant to absorb them — but only if they are large enough. ## Fixes, in order of leverage **Allocate less, or hold shorter.** The highest-leverage fix is almost always application-side: stop materialising whole result sets, stream instead of buffering, cut the per-request retained footprint, shorten the window during which request data is reachable. Every downstream tuning knob is compensating for this. **Enlarge the young generation.** Give short-lived data time to die in eden before a collection fires. The cost is a longer (though less frequent) young pause and less heap left for the old generation. This is the single most effective flag-level change for the allocation-rate case. **Enlarge survivor space** (via the eden-to-survivor sizing on classic collectors, or by allowing a larger young size on region-based ones) so the aging filter actually operates and the adaptive threshold stops collapsing to 1. **Adjust the tenuring threshold — last.** Raising the maximum only helps if survivors *fit*; if they overflow, age is irrelevant. Lowering it is occasionally right when the age histogram proves survivors are genuinely long-lived and the repeated copying is pure waste. **Reconsider the collector.** If the workload keeps a large, mutable, medium-lived working set, the generational assumption is weak for it, and a collector whose old-generation work is concurrent may cost less than fighting the promotion rate. ## Verification Whatever you change, verify against the same three signals: promoted bytes per young collection, the chosen tenuring threshold, and the post-collection old-generation floor. And measure at steady state under representative load — the failure mode is workload-shaped, so a synthetic benchmark that allocates uniformly will not reproduce it.

  • Why can enlarging the young generation reduce promotion even without touching the tenuring threshold?
    Because object lifetime is effectively measured in bytes allocated between collections. A bigger eden means more time passes before a young collection fires, so more short-lived data is already dead when it does. Objects that previously survived one collection by bad timing now never survive at all, so they are never copied and never age toward promotion.
  • When is high promotion a normal, healthy signal rather than a defect?
    During warm-up, when a cache or long-lived working set is genuinely being built, promotion is exactly what should happen — the data is long-lived and belongs in the old generation. The distinguishing feature is that it is transient and bounded: promotion volume falls to a low steady state once the working set is populated, and the post-collection floor stabilises at the new level.
  • Why is raising MaxTenuringThreshold usually the wrong first move?
    Because the threshold only matters if survivors have somewhere to live. If the survivor space overflows, objects are promoted on overflow regardless of their age, and the adaptive calculation will keep driving the effective threshold down anyway. Fixing the capacity — young size or survivor size — is what makes any threshold setting take effect.

Promotion is meant to be tenure after a probation period. Premature promotion is granting tenure on day one because the probation office has no desks — the survivor space is full, so everyone is waved through.

saying these in an interview costs you the question

  • Diagnosing rising old-generation usage as premature promotion without checking the post-collection floor for a leak.
  • Reaching for MaxTenuringThreshold first while survivors are overflowing.
  • Assuming forcing more frequent full collections is a fix rather than a symptom amplifier.
  • Treating object lifetime as wall-clock time instead of bytes allocated between collections.
  • Concluding from a synthetic benchmark that the problem is gone without reproducing the real allocation shape.

context