A service allocates roughly 1 GB per second of mostly short-lived objects and its young space is 512 MB. How does allocation rate translate into collection frequency, and what levers change that?
answer
- interval ≈ young size / allocation rate
- frequency from rate; duration from survivors
- rate from GC log: before(N+1) minus after(N)
- allocation sampling hooks buffer refills
- pooling converts cheap garbage into promoted data
basics
~20 sYoung collections happen roughly every time the young space fills: 1 GB/s into 512 MB is about two young collections per second. Levers are allocating fewer bytes, enlarging the young space, and reducing how much survives each cycle.
solid answer
~60 sTo a first approximation, young-collection interval equals young-space size divided by allocation rate. At 1 GB/s with 512 MB of young space that is one collection every half second, about two per second. Each collection's *duration*, by contrast, depends on the live set copied out, not on the bytes allocated. So there are three independent levers. **Allocate less.** Fewer or smaller objects per unit of work directly stretches the interval. This is usually the highest-value fix: boxing, per-request temporary collections, string concatenation in hot paths, defensive copies, and oversized buffers are the usual sources. **Make the young space bigger.** Doubling it halves the frequency, and it also helps quality: objects get more time to die before a collection sees them, so less is copied and less is promoted. **Reduce survivors.** If survivor space overflows or objects live just long enough, they get promoted prematurely, which loads the old generation and eventually causes more expensive collections there. Measure allocation rate from garbage-collection logs as heap occupancy after one collection subtracted from occupancy before the next, divided by the elapsed time.
code
text · 6 lines[10.000s] Pause Young ... 1400M->200M(2048M) 12.3ms
[10.500s] Pause Young ... 1420M->205M(2048M) 12.9ms
allocated between the two = 1420M - 200M = 1220M
elapsed = 0.5 s
allocation rate ≈ 2.4 GB/sgo deeper
State the relationship: more bytes allocated per second means the young space fills sooner and collections run more often.
Do the arithmetic, and separate what drives frequency from what drives pause duration.
Measure the rate from logs, profile the call sites, and order the fixes from allocating less through sizing to promotion control.
Treat allocation rate as a capacity input — relate it to memory bandwidth, headroom, collector CPU budget, and the cost of each remedy against the service's latency contract.
## The arithmetic Every byte a thread allocates advances its private buffer pointer, and every buffer refill advances the shared young-space pointer. When the young space cannot serve another buffer, a young collection runs. So: ``` interval between young collections ≈ young space size / allocation rate ``` At 1 GB/s with 512 MB usable young space, that is ~0.5 s, or about two collections per second. This relationship is why allocation rate is the single most predictive number about a JVM workload's collector behaviour. Crucially, frequency and duration have different drivers. Frequency comes from allocation rate versus young size. Duration comes from the live set at collection time, because a copying young collector traces and copies only survivors. That decoupling explains a common observation: a service can allocate enormously and still have millisecond pauses, as long as almost nothing survives. ## Measuring the rate Garbage-collection logs give it directly. Each young collection reports heap occupancy before and after. Occupancy before collection N+1 minus occupancy after collection N is the bytes allocated in between; divide by the wall time between them. Doing this over a representative window gives a steady-state allocation rate. Allocation profiling adds the second half of the picture — which call sites produce those bytes — and the JVM samples it cheaply by hooking the allocation slow path: it records a sample when a thread refills its buffer, and separately when an object is allocated outside a buffer. That is why allocation profiles are sampled per N bytes allocated rather than per object, and why they are nearly free to collect. ## The three levers in order of value **1. Allocate fewer bytes per unit of work.** This improves everything at once: fewer collections, less zeroing, less memory bandwidth, better cache behaviour. Typical wins come from removing per-call temporary objects in hot loops, avoiding autoboxing of primitives in high-throughput paths, reusing buffers where the lifetime is clearly scoped, streaming instead of materializing large intermediate collections, and not building log strings that are then discarded by the level check. **2. Enlarge the young space.** Frequency falls in proportion. There is a second-order benefit: with a longer interval, more objects die naturally before a collection observes them, so both the copied volume and the promotion rate fall. The costs are a bigger memory footprint and, for collectors whose pause scales with survivors, potentially longer individual pauses if the live set grows with the space. **3. Reduce what survives.** Survivors are the expensive part. Objects that outlive a couple of collections get promoted to the old generation, and a high promotion rate is what eventually drives the expensive old-generation work. Caches with unbounded or poorly-tuned lifetimes, per-request state kept alive by asynchronous continuations, and buffers held across an await are common causes of objects surviving just long enough to be promoted. ## What does not help Calling `System.gc()` does not reduce allocation rate; it adds work. Switching collectors changes the shape of pauses, not the volume of bytes your code allocates — a low-pause collector will do the same amount of copying, just concurrently and with more CPU. And object pooling to dodge allocation is usually counterproductive for small short-lived objects, because pooled objects survive by construction, converting cheap young garbage into promoted long-lived data that the old generation must manage. ## Framing an answer A strong answer states the arithmetic, separates frequency from duration, names measurement (logs for rate, sampled allocation profiling for call sites), and orders the fixes: reduce bytes allocated, then size the young space, then attack survivor and promotion volume — with pooling explicitly called out as a last resort rather than a default.
- If pauses are already short, why care about a high allocation rate at all?Because it costs throughput even when pauses are invisible. Zeroing and copying consume memory bandwidth, frequent collections add fixed per-collection overhead such as stopping threads and scanning roots, and a higher rate leaves less headroom before objects start being promoted prematurely. It is also the main knob that still works when pause targets cannot be lowered further.
- Why can enlarging the young space reduce promotion, not just collection frequency?Objects are promoted when they survive collections. A longer interval between collections gives short-lived objects more time to become unreachable before any collection observes them, so fewer are copied and fewer accumulate the age that triggers promotion. Larger survivor spaces likewise reduce overflow-driven premature promotion.
saying these in an interview costs you the question
- Treating pause duration as proportional to allocation rate; duration tracks the live set.
- Proposing object pooling as the first fix for high allocation of small objects.
- Assuming switching to a low-pause collector reduces the amount of garbage produced.
- Reporting allocation rate from heap-used samples without subtracting collections, which understates it badly.