skip to content

How do you compute an application's allocation rate and promotion rate from a HotSpot garbage-collection log, and what do those two numbers tell you?

level: seniorimportance: should knowfreq 45%

answer

  1. alloc = used_before(n) − used_after(n−1) ÷ Δt
  2. promotion needs gc+heap=debug (old occupancy / old regions)
  3. allocation → collection frequency; promotion → old-gen pressure
  4. gc+age=trace separates premature promotion from long-lived data
  5. humongous/large objects skip young and skew both numbers

basics

~20 s

Allocation rate = (heap used before a collection − heap used after the previous one) ÷ the interval between them, averaged over many collections. Promotion rate = old-generation growth per collection ÷ the same interval. High allocation means frequent young pauses; high promotion means old-generation pressure and eventual long collections.

solid answer

~50 s

**Allocation rate.** Between two consecutive young collections nothing is freed, so everything the heap gained is what the application allocated. Take `used_before(GC n) − used_after(GC n−1)`, divide by the timestamp difference, and average over a window: that is MB/s of allocation. It drives young-collection *frequency* — double the eden and, at the same rate, you halve the number of collections. **Promotion rate.** Enable `-Xlog:gc+heap=debug` so each collection prints old-generation occupancy (or G1's old-region counts). The old-generation growth across a young collection is what got promoted; divide by the interval for MB/s. This is the number that matters, because promotion is what eventually forces old-generation work — concurrent cycles, mixed collections, or full collections. A healthy profile is high allocation and near-zero promotion: objects die in eden. Allocation rate that a collector can absorb is a throughput question; promotion rate is a latency question. If promotion is high, ask whether objects are genuinely long-lived or are being promoted prematurely — `-Xlog:gc+age=trace` shows the tenuring distribution that separates the two.

code

text · 6 lines
text
[41.202s][info][gc            ] GC(88) Pause Young (Normal) (G1 Evacuation Pause) 1832M->612M(4096M) 11.402ms
[41.202s][debug][gc,heap       ] GC(88) Eden regions: 152->0(150)
[41.202s][debug][gc,heap       ] GC(88) Survivor regions: 8->10(20)
[41.202s][debug][gc,heap       ] GC(88) Old regions: 296->302
[41.202s][debug][gc,heap       ] GC(88) Humongous regions: 4->4
[41.958s][info][gc            ] GC(89) Pause Young (Normal) (G1 Evacuation Pause) 1820M->620M(4096M) 10.918ms

go deeper

for a junior

Know the two terms and that allocation rate governs how often young collections happen while promotion governs old-generation pressure.

for a middle

Do the arithmetic correctly from consecutive collections and know which log sub-tag exposes old-generation occupancy.

for a senior

Derive both from a real window, separate premature promotion from a large working set with the tenuring distribution, and state the caveats about humongous objects and bursty workloads.

for a principal

Use the pair as the workload's memory signature — the input to capacity planning, collector choice and the argument about whether the fix belongs in configuration or in the allocating code.

## Why these two numbers A GC log is mostly a record of two flows: how fast the application creates objects, and how fast objects escape the young generation into the old one. Nearly every GC symptom is a consequence of one of them, so deriving both is the first analytical step after confirming the log is complete. ## Deriving allocation rate Between the end of one collection and the start of the next, no memory is reclaimed — the heap only grows, and it grows exactly by what the application allocated (large objects that bypass eden are the exception, discussed below). So for consecutive collections *n−1* and *n*: ``` allocated = used_before(n) - used_after(n-1) interval = timestamp(n) - timestamp(n-1) rate = allocated / interval // MB/s ``` One pair is noise; average over dozens of collections in a steady-state window, and look at the distribution, not just the mean — a batch job that allocates in bursts has a peak rate the collector must absorb even if its average is modest. What the number means: allocation rate sets young-collection **frequency** for a given eden size. Frequency × pause = the throughput tax of young collection. A very high rate (say, several GB/s) is a signal to look at the code — churn in a hot loop, boxing, defensive copying, oversized log-message formatting — because reducing allocation is usually cheaper than tuning around it. ## Deriving promotion rate Promotion is invisible in the summary line, which reports total heap only. Enable the heap sub-tag: ``` -Xlog:gc*,gc+heap=debug:file=gc.log:time,uptime,level,tags ``` For Serial/Parallel you then get old-generation occupancy before and after each collection; for G1 you get region counts by class (`Eden regions: 24->0(25)`, `Survivor regions: 3->2`, `Old regions: 40->43`, `Humongous regions: 1->1`), and region count × region size gives bytes. ``` promoted(n) = old_used_after(n) - old_used_before(n) promotion rate = sum(promoted) / window duration ``` This number is the one that predicts pain. Everything promoted must eventually be traced by an old-generation cycle and, if it dies there, reclaimed by concurrent marking plus mixed collections, or by a full collection. Promotion rate therefore drives how often concurrent cycles run and how much they have to do, and it is the quantity that turns a healthy young-only rhythm into old-generation churn. ## Interpreting the pair - **High allocation, near-zero promotion** — the ideal generational profile. Objects die young; young collections stay cheap because cost tracks survivors, not garbage. If pauses are still too frequent, the lever is eden size or the collector, not the code. - **High promotion** — objects are surviving. Two very different causes: - *Genuinely long-lived data* — caches, session state, a large steady working set. The old generation must simply be able to hold it; the log will show a high, stable used-after floor. - *Premature promotion* — objects that would have died shortly after, pushed out because survivor space was too small or the tenuring threshold too low, so the collector gave up on ageing them. The tenuring distribution from `-Xlog:gc+age=trace` distinguishes the two: if the distribution shows a lot of bytes at age 1–2 and the desired survivor size exceeds the actual, objects are being flushed out before they get a chance to die. - **Rising used-after floor across the window** — the live set is growing. That is a retention question, pursued with heap-analysis tooling rather than with GC flags. ## Caveats that keep the arithmetic honest - **Humongous / large objects.** G1 allocates objects at least half a region in size directly into old-generation humongous regions, and other collectors can allocate large arrays straight into the old generation. Those bytes appear as promotion without ever having been young, and they distort both numbers. G1's humongous region counts in `gc+heap` output let you see the volume. - **TLAB slack.** Threads allocate in thread-local buffers; unused tail space and retired buffers make the heap-derived figure slightly conservative. It is close enough for tuning decisions. - **Only compare like intervals.** Mixed collections, concurrent-start collections and full collections change the accounting; compute allocation rate across consecutive *young* collections in a steady-state period, excluding startup, warmup and deployment spikes. - **Log-derived rates are averages.** They tell you the shape of the workload; they do not tell you which allocation sites produce it. For that you need allocation profiling, not the GC log. ## What you do with them The pair converts a vague complaint into a quantified statement: "we allocate 1.2 GB/s and promote 40 MB/s; at the current eden size that is a young collection every 250 ms at 12 ms each, and old occupancy grows fast enough to require a concurrent cycle every 90 seconds." That statement is what makes the next decision — resize, retune, change collector, or fix the allocating code — an argument from evidence instead of a guess.

  • The log shows a high promotion rate. How do you tell premature promotion from a genuinely large long-lived working set?
    Turn on `-Xlog:gc+age=trace` and read the tenuring distribution. Premature promotion looks like a large volume of bytes at low ages with the desired survivor size exceeding the actual survivor capacity — objects are evicted before they can die. A genuine working set shows a stable high used-after floor after collections and objects that survive many age increments regardless of survivor sizing.
  • Why can you not simply read the allocation rate off a single GC log line?
    One line reports used-before, used-after and the pause for a single collection; the allocation rate is a rate over time and needs the previous collection's used-after value plus the interval between the two timestamps. It also needs averaging: allocation is bursty, so one interval reflects a moment rather than the workload.

saying these in an interview costs you the question

  • Computing allocation rate from used-before minus used-after of the same collection — that is the amount reclaimed, not the amount allocated.
  • Believing the summary line alone reveals promotion; old-generation occupancy needs the `gc+heap` sub-tag.
  • Treating a high allocation rate as inherently bad without checking whether anything survives.
  • Ignoring humongous or large-array allocations that enter the old generation directly and inflate the apparent promotion.
  • Averaging across startup, deployment and steady state and then reasoning from the meaningless mean.

context