The HotSpot Garbage-First (G1) collector divides the Java heap into regions instead of contiguous young and old spaces. Explain how that layout works and what eden, survivor, old, and humongous regions mean.
answer
- heap/2048, 1-32 MB, power of two
- role is a tag, not an address range
- eden / survivor / old / humongous
- evacuate then free whole region
- humongous = >= half a region
basics
~20 sG1 splits the heap into equal-sized regions (a power of two, 1-32 MB, roughly 2048 of them). Each region is tagged at runtime as eden, survivor, old, or humongous. A generation is the current set of regions with that tag, not a contiguous address range.
solid answer
~60 sG1 reserves the heap as a grid of equal-sized regions. HotSpot aims for about 2048 regions, so region size is heap-size/2048 rounded to a power of two and clamped to 1-32 MB; `-XX:G1HeapRegionSize` overrides it. A region has no permanent generation: it is tagged **free**, **eden**, **survivor**, **old**, or **humongous** as it is used, and the tag changes over its life. The young generation is simply the current set of eden plus survivor regions, scattered anywhere in the address space, so G1 can grow or shrink young size between pauses without moving a contiguous boundary. Threads allocate through TLABs carved from eden regions by bump pointer. A collection **evacuates** live objects out of the chosen regions into survivor or old regions and returns the emptied regions to the free list whole, so compaction is a byproduct of collecting rather than a separate phase. An object needing at least half a region is **humongous**: it is allocated directly into a contiguous run of regions tagged humongous, accounted as old, and is not moved by ordinary evacuation.
code
text · 4 linesjava -XX:+UseG1GC -Xmx16g -XX:G1HeapRegionSize=8m -Xlog:gc,gc+heap=debug MyApp
# 16 GB / 2048 = 8 MB, so the explicit flag above matches the default choice.
# A 4 GB heap defaults to 2 MB regions; anything above 32 GB stays clamped at 32 MB.go deeper
Be able to say the heap is cut into many equal regions and that each region is labelled eden, survivor, old, or humongous, and that labels change over time.
Add the sizing rule (about 2048 regions, power of two, 1-32 MB), the humongous threshold of half a region, and that collecting means copying live objects out and freeing whole regions.
Explain why the layout enables adaptive young sizing and incremental old-generation work, and name the costs: per-region metadata, remembered sets, and the need to keep free regions in reserve to copy into.
Frame it as a design trade: paying continuous bookkeeping and a copy reserve to convert a heap-size-proportional pause into many bounded, schedulable units of work, and note where that trade breaks down (humongous-heavy or allocation-burst workloads).
## The layout G1 replaced The classic HotSpot heap is contiguous: an eden plus two survivor spaces form the young generation, and the old generation sits next to it as one block. Because each space is a contiguous address range, resizing it means moving a boundary, and reclaiming the old generation means processing the whole old block at once. As heaps grew to tens of gigabytes, the cost of touching the entire old space in one operation made pauses scale with heap size. ## The region grid G1 divides the reserved heap into equal-sized regions. HotSpot picks the size so that there are roughly 2048 regions: heap-size/2048, rounded down to a power of two, clamped between 1 MB and 32 MB. `-XX:G1HeapRegionSize=<n>m` forces a value. Regions are the unit of allocation, of collection, and of accounting. Each region carries a role, and roles are reassigned constantly: - **Free** - on the free list, not yet in use. - **Eden** - fresh allocation lands here, through thread-local allocation buffers (TLABs) carved from a region and filled by bump pointer. - **Survivor** - holds objects that survived at least one collection but have not yet been promoted. - **Old** - holds promoted objects, i.e. those that survived enough collections (or were forced out because survivor space was full). - **Humongous** - holds a single object of at least half a region, spanning one or more contiguous regions; counted as old. Because roles are per-region tags, "the young generation" is just the set of regions currently tagged eden or survivor. Those regions can be anywhere in the address space. This is what lets G1 change young capacity from one pause to the next: it simply decides how many free regions to hand out as eden before triggering the next collection, instead of resizing a contiguous space. ## How a collection uses the layout G1 collects by **evacuation**: it picks a set of regions, copies their live objects into fresh survivor or old regions, and then frees the source regions in their entirety. Since surviving objects are packed densely into destination regions, compaction happens as a side effect - G1 never sweeps garbage in place in its normal cycles, so there is no free-list fragmentation inside collected regions. The price is that copying requires somewhere to copy to, so G1 must keep free regions in reserve. ## Humongous objects An allocation whose size is at least 50% of one region cannot be handled by the ordinary TLAB path. G1 finds a contiguous run of free regions, marks the first as humongous-start and the rest as humongous-continues, and places the object there. The tail of the last region is wasted. Humongous regions are treated as old, are not copied by normal evacuation, and are reclaimed as whole regions when the object is found unreachable - modern HotSpot can reclaim them eagerly during a young pause when nothing points into them. ## Consequences to state in an interview The region grid buys incremental old-generation reclamation (collect a subset of old regions per pause), adaptive generation sizing, and compaction for free. It costs extra bookkeeping - per-region liveness accounting and per-region remembered sets recording incoming references - plus the possibility of running out of free regions to evacuate into.
- What happens if a single object is larger than one region?It is still allocated as a humongous object, but G1 must find a contiguous run of free regions large enough to hold it. The first region is tagged humongous-start and the rest humongous-continues. If no contiguous run is available, G1 must collect or expand the heap first, and repeated large allocations can fail even when total free memory looks sufficient because the free space is not contiguous.
- Does a smaller region size always give shorter pauses?No. Smaller regions give finer granularity when choosing a collection set, which can help, but they multiply the number of remembered sets and per-region metadata, raise scan and bookkeeping overhead, and lower the humongous threshold so more ordinary arrays become humongous. Region size is rarely the right first tuning knob; the pause-time goal is.
Think of a warehouse of identical pallets rather than two fixed rooms. Any pallet can be labelled 'new stock' or 'long-term storage' today and relabelled tomorrow; you clear a pallet by moving the few good items onto a fresh pallet and wheeling the old one away empty.
saying these in an interview costs you the question
- Saying G1 has a contiguous young generation that is simply split into regions - roles are per-region tags, and young regions can be anywhere.
- Claiming a region is permanently young or old; the tag is reassigned every time a region is recycled.
- Thinking humongous means 'very large heap' rather than 'object occupying at least half a region'.
- Assuming G1 sweeps garbage in place and therefore fragments like a free-list collector; normal G1 cycles evacuate and free whole regions.
- Believing region size is fixed at 1 MB or is chosen by the application rather than derived from heap size.