skip to content

A service allocates roughly 1 GB per second of mostly short-lived objects and its young space is 512 MB. How does allocation rate translate into collection frequency, and what levers change that?

level: seniorimportance: should knowfreq 45%

answer

  1. interval ≈ young size / allocation rate
  2. frequency from rate; duration from survivors
  3. rate from GC log: before(N+1) minus after(N)
  4. allocation sampling hooks buffer refills
  5. pooling converts cheap garbage into promoted data

basics

~20 s

Young collections happen roughly every time the young space fills: 1 GB/s into 512 MB is about two young collections per second. Levers are allocating fewer bytes, enlarging the young space, and reducing how much survives each cycle.

solid answer

~60 s

To a first approximation, young-collection interval equals young-space size divided by allocation rate. At 1 GB/s with 512 MB of young space that is one collection every half second, about two per second. Each collection's *duration*, by contrast, depends on the live set copied out, not on the bytes allocated. So there are three independent levers. **Allocate less.** Fewer or smaller objects per unit of work directly stretches the interval. This is usually the highest-value fix: boxing, per-request temporary collections, string concatenation in hot paths, defensive copies, and oversized buffers are the usual sources. **Make the young space bigger.** Doubling it halves the frequency, and it also helps quality: objects get more time to die before a collection sees them, so less is copied and less is promoted. **Reduce survivors.** If survivor space overflows or objects live just long enough, they get promoted prematurely, which loads the old generation and eventually causes more expensive collections there. Measure allocation rate from garbage-collection logs as heap occupancy after one collection subtracted from occupancy before the next, divided by the elapsed time.

code

text · 6 lines
text
[10.000s] Pause Young ... 1400M->200M(2048M) 12.3ms
[10.500s] Pause Young ... 1420M->205M(2048M) 12.9ms

allocated between the two = 1420M - 200M = 1220M
elapsed                   = 0.5 s
allocation rate           ≈ 2.4 GB/s

go deeper

for a junior

State the relationship: more bytes allocated per second means the young space fills sooner and collections run more often.

for a middle

Do the arithmetic, and separate what drives frequency from what drives pause duration.

for a senior

Measure the rate from logs, profile the call sites, and order the fixes from allocating less through sizing to promotion control.

for a principal

Treat allocation rate as a capacity input — relate it to memory bandwidth, headroom, collector CPU budget, and the cost of each remedy against the service's latency contract.

## The arithmetic Every byte a thread allocates advances its private buffer pointer, and every buffer refill advances the shared young-space pointer. When the young space cannot serve another buffer, a young collection runs. So: ``` interval between young collections ≈ young space size / allocation rate ``` At 1 GB/s with 512 MB usable young space, that is ~0.5 s, or about two collections per second. This relationship is why allocation rate is the single most predictive number about a JVM workload's collector behaviour. Crucially, frequency and duration have different drivers. Frequency comes from allocation rate versus young size. Duration comes from the live set at collection time, because a copying young collector traces and copies only survivors. That decoupling explains a common observation: a service can allocate enormously and still have millisecond pauses, as long as almost nothing survives. ## Measuring the rate Garbage-collection logs give it directly. Each young collection reports heap occupancy before and after. Occupancy before collection N+1 minus occupancy after collection N is the bytes allocated in between; divide by the wall time between them. Doing this over a representative window gives a steady-state allocation rate. Allocation profiling adds the second half of the picture — which call sites produce those bytes — and the JVM samples it cheaply by hooking the allocation slow path: it records a sample when a thread refills its buffer, and separately when an object is allocated outside a buffer. That is why allocation profiles are sampled per N bytes allocated rather than per object, and why they are nearly free to collect. ## The three levers in order of value **1. Allocate fewer bytes per unit of work.** This improves everything at once: fewer collections, less zeroing, less memory bandwidth, better cache behaviour. Typical wins come from removing per-call temporary objects in hot loops, avoiding autoboxing of primitives in high-throughput paths, reusing buffers where the lifetime is clearly scoped, streaming instead of materializing large intermediate collections, and not building log strings that are then discarded by the level check. **2. Enlarge the young space.** Frequency falls in proportion. There is a second-order benefit: with a longer interval, more objects die naturally before a collection observes them, so both the copied volume and the promotion rate fall. The costs are a bigger memory footprint and, for collectors whose pause scales with survivors, potentially longer individual pauses if the live set grows with the space. **3. Reduce what survives.** Survivors are the expensive part. Objects that outlive a couple of collections get promoted to the old generation, and a high promotion rate is what eventually drives the expensive old-generation work. Caches with unbounded or poorly-tuned lifetimes, per-request state kept alive by asynchronous continuations, and buffers held across an await are common causes of objects surviving just long enough to be promoted. ## What does not help Calling `System.gc()` does not reduce allocation rate; it adds work. Switching collectors changes the shape of pauses, not the volume of bytes your code allocates — a low-pause collector will do the same amount of copying, just concurrently and with more CPU. And object pooling to dodge allocation is usually counterproductive for small short-lived objects, because pooled objects survive by construction, converting cheap young garbage into promoted long-lived data that the old generation must manage. ## Framing an answer A strong answer states the arithmetic, separates frequency from duration, names measurement (logs for rate, sampled allocation profiling for call sites), and orders the fixes: reduce bytes allocated, then size the young space, then attack survivor and promotion volume — with pooling explicitly called out as a last resort rather than a default.

  • If pauses are already short, why care about a high allocation rate at all?
    Because it costs throughput even when pauses are invisible. Zeroing and copying consume memory bandwidth, frequent collections add fixed per-collection overhead such as stopping threads and scanning roots, and a higher rate leaves less headroom before objects start being promoted prematurely. It is also the main knob that still works when pause targets cannot be lowered further.
  • Why can enlarging the young space reduce promotion, not just collection frequency?
    Objects are promoted when they survive collections. A longer interval between collections gives short-lived objects more time to become unreachable before any collection observes them, so fewer are copied and fewer accumulate the age that triggers promotion. Larger survivor spaces likewise reduce overflow-driven premature promotion.

saying these in an interview costs you the question

  • Treating pause duration as proportional to allocation rate; duration tracks the live set.
  • Proposing object pooling as the first fix for high allocation of small objects.
  • Assuming switching to a low-pause collector reduces the amount of garbage produced.
  • Reporting allocation rate from heap-used samples without subtracting collections, which understates it badly.

context