For HotSpot's Parallel collector (-XX:+UseParallelGC), describe what happens during a young collection versus a full collection — which algorithm each phase uses, and why the resulting heap has no fragmentation.
answer
- Eden + two survivors + old
- Young = parallel copy, cost ∝ survivors
- Promotion at tenuring threshold
- Full = mark, summary, sliding compact
- Moving collector = contiguous free space = bump allocation
basics
~20 sA young collection is a parallel copying collection: live objects in eden and the active survivor space are copied to the other survivor space or promoted to the old generation. A full collection is a parallel mark, summary and sliding-compaction pass over the entire heap. Both move objects, so free space stays contiguous.
solid answer
~60 s**Young collection.** The heap's young generation is eden plus two survivor spaces. When eden fills, all threads stop and the GC threads scan the roots plus dirty cards (old objects holding references into young) and **copy** every live object out: to the empty survivor space, or straight to the old generation if it has already survived enough collections or the survivor space cannot hold it. Eden and the source survivor space are then declared empty wholesale. Cost is proportional to *surviving* data, not to garbage, which is why a high infant-mortality workload gets cheap young collections. **Full collection.** Three parallel phases over the whole heap: **mark** live objects from the roots; **summary**, which computes for each region of the old space where its live data will end up after compaction; and **compact**, which slides live objects into place and fixes every reference to them. Both halves are moving collectors, so after either one the free space is a single contiguous block. Allocation is then just a pointer bump, and there is no fragmentation and no free-list bookkeeping.
code
text · 2 lines[GC (Allocation Failure) [PSYoungGen: 262112K->21489K(305664K)] 262112K->21497K(1005056K), 0.0221ms]
[Full GC (Ergonomics) [PSYoungGen: 21489K->0K(305664K)] [ParOldGen: 8K->21150K(699392K)] 21497K->21150K(1005056K), [Metaspace: 8721K->8721K(1056768K)], 0.4133 secs]go deeper
Be able to name the spaces (eden, two survivors, old) and say that young collections copy survivors out while full collections compact the whole heap.
Explain why copying costs scale with survivors, what promotion and the tenuring threshold are, and why compaction eliminates fragmentation.
Add the operational reading: which log line is which, what a full collection that fails to reclaim tells you, and how large-object allocation bypasses eden.
Discuss the trade the design encodes — moving collection buys bump allocation and zero fragmentation at the cost of requiring a stop, which is the fork in the road every later collector had to navigate.
## The heap layout this collector works on The Parallel collector uses a classic generational layout with contiguous spaces: - **Eden** — where almost all new objects are allocated, via a pointer bump inside each thread's own local allocation buffer. - **Two survivor spaces**, conventionally called *from* and *to*. At any moment one is empty. - **Old generation** — one contiguous space holding objects that have survived long enough to be promoted. In legacy GC logs these appear as `PSYoungGen` and `ParOldGen`; `PS` stands for "parallel scavenge", the name of the young collector. ## Young collection: parallel copying When eden has no room for an allocation, the JVM brings every application thread to a safepoint and runs a *scavenge*. 1. **Find the roots.** Thread stacks, static fields, JNI handles — plus any old-generation objects that hold a reference into the young generation. Rather than scanning the whole old generation for those, the collector consults the card table: reference stores dirty a card, and only dirty cards are scanned. This is what makes a young collection cheap even with a huge old generation. 2. **Copy, don't sweep.** Every reachable young object is copied out of eden and the occupied survivor space. Its destination is either the empty survivor space or, if the object's survival count has reached the tenuring threshold or the survivor space is too full, the old generation — this is *promotion*. GC threads copy into their own promotion buffers to avoid contending on a shared allocation pointer, and they steal work from each other's queues so no thread finishes early while another has a deep graph left. 3. **Forward the references.** Each copied object leaves a forwarding pointer behind so that other references to it are updated to the new address rather than duplicating the object. 4. **Reclaim wholesale.** Eden and the source survivor space are now entirely garbage and are declared empty by resetting their allocation pointers. The survivor roles swap. The key property: the work is proportional to the volume of *surviving* data. Garbage costs nothing to reclaim. If 98% of a generation dies before collection, a young collection touches 2% of it. ## Full collection: parallel mark-summary-compact When the old generation cannot satisfy a promotion, or ergonomics decide the heap needs reorganising, the collector runs a full collection over young *and* old generations in one stop, using all GC threads across three phases: - **Marking.** Starting from the roots, GC threads traverse the object graph in parallel and mark everything reachable. Reachability information is recorded in bitmaps. - **Summary.** The old space is divided into fixed-size chunks; the collector computes, per chunk, how much live data it holds and therefore what the post-compaction destination address of that data will be. Because the density of live data usually increases toward the start of the old generation, the phase identifies a prefix that is so dense that compacting it would cost more than it recovers, and leaves it alone. - **Compaction.** Live objects are slid toward one end of the space in parallel, and every reference pointing at them is updated. Class metadata is unloaded and weak references processed in the same stop. ## Why there is no fragmentation A mark-**sweep** collector leaves the survivors in place and adds the gaps between them to a free list. Over time the free space becomes many small non-adjacent holes: you can have plenty of free bytes and still be unable to satisfy one large allocation, and allocation requires searching the free list. That is fragmentation. The Parallel collector never sweeps. Its young half copies survivors out and empties the source spaces; its full half slides survivors together. In both cases the outcome is *all live data at one end, all free space in one contiguous block at the other*. Consequences: - Allocation is a pointer bump and a bounds check — a handful of instructions. - There is no free-list metadata to maintain, and no allocation failure caused by fragmentation. - Objects allocated together stay adjacent after copying, which helps cache locality. The price is that moving objects means fixing every reference to them, which is precisely why all of this must happen with the application stopped — no application thread may observe an object mid-move. ## Reading it in a log `Pause Young (Allocation Failure)` is the scavenge; `Pause Full (Ergonomics)` or `Pause Full (Allocation Failure)` is the mark-summary-compact. The occupancy figures before and after tell you how much survived, and a full collection whose "after" figure stays high is the classic sign of a live set that no longer fits. ## What interviewers listen for That you distinguish *copying, cost proportional to survivors* in the young generation from *mark-summary-compact over the whole heap* for a full collection, and that you can say why compaction removes fragmentation and makes allocation a pointer bump.
- Why is a young collection cheap even when the young generation is large?Because a copying collector's cost is proportional to the data that survives, not to the space being collected. Most objects die young, so the collector typically copies a small percentage of eden and then reclaims the entire space by resetting a pointer. Doubling eden while keeping the survival rate constant does not double the pause — it mostly buys you fewer collections.
- What happens when an object is too large to fit in eden?It is allocated directly in the old generation. A very large array skips the young generation entirely, which means it will only ever be reclaimed by a full collection. A workload that churns large arrays therefore drives full collections even though the objects are short-lived, which is a classic cause of unexpected long pauses under this collector.
saying these in an interview costs you the question
- Describing the young collection as mark-and-sweep — it is a copying collector and never touches the garbage
- Saying the Parallel collector's old generation is swept and therefore fragments; it is compacted
- Believing survivor spaces are both in use at once — one is always empty at the start of a scavenge
- Thinking cost scales with the size of the generation rather than with the volume of live data
- Assuming a full collection collects only the old generation; it collects the entire heap and unloads classes