When is a fully concurrent low-pause collector such as ZGC or Shenandoah the right choice for a JVM service, and what does it cost you compared with G1?
answer
- Concurrent marking AND concurrent relocation → pause decoupled from heap size
- Reach for it when a tail-latency SLO + big heap collide
- Cost: access barriers (~10% throughput), CPU for concurrent work, headroom
- Failure mode = allocation stall (allocation outran the collector)
- Production-ready from JDK 15; ZGC generational by default from JDK 23
basics
~20 sChoose them when a strict tail-latency SLO must hold on a large heap — pauses stay sub-millisecond to low-millisecond largely independent of heap size, because marking and relocation run concurrently. You pay barrier overhead on application code, extra CPU for concurrent work, and more heap headroom.
solid answer
~60 sZGC and Shenandoah do essentially all their work — marking *and* object relocation — concurrently with application threads, leaving only short fixed-cost stops (root scanning and handshakes). The headline property is that pause time is largely **decoupled from heap size and live-set size**: a 4 GB heap and a 500 GB heap both see pauses in roughly the same sub-millisecond to low-millisecond band, where G1's pauses grow with the collection set and its full-compaction fallback grows with the live set. Choose one when: - a p99/p99.9 latency SLO is tight enough that even a well-tuned 100–200 ms G1 pause is unacceptable; - the heap is large (tens to hundreds of GB) and G1's pauses scale badly on it; - pause *predictability* matters more than a few percent of throughput. The costs are real: load/read barriers on reference access slow application code (commonly on the order of ~10% throughput versus a throughput collector, workload-dependent); concurrent GC threads need spare CPU; and relocation needs free heap headroom — run too close to the ceiling and allocation stalls appear, which are the low-pause equivalent of falling behind. On a small heap with a relaxed SLO, they buy nothing.
go deeper
Know that ZGC and Shenandoah do marking and relocation concurrently so pauses stay very short regardless of heap size, at some throughput cost.
Explain where the cost comes from — access barriers, concurrent CPU, headroom — and that they are chosen for latency, not throughput.
Own the decision: state the SLO numerically, measure G1 first, trial under real load, budget CPU and heap headroom, and recognise allocation stalls as the failure mode.
Weigh it as a fleet economics question — extra CPU and memory per node versus tail-latency risk and failover behaviour — and set the criteria under which a service is allowed to switch.
## What "fully concurrent" buys Every collector must trace live objects and, if it compacts, move them. G1 traces concurrently but *evacuates* in stop-the-world pauses, which is why its pause duration scales with how much it decided to collect. ZGC and Shenandoah move relocation itself into the concurrent phase: application threads keep running while objects are being copied, and reference reads (or accesses) are intercepted by barriers so a thread never sees a stale address. The consequence is the property that matters operationally: **pause time is bounded by fixed work, not by heap or live-set size**. What remains stop-the-world is root scanning and short handshakes — work proportional to thread count and root set, not to the gigabytes on the heap. So the collectors keep pauses in the sub-millisecond to low-millisecond range across heaps from small to enormous, and the pause profile stays flat as the heap grows. ## When to reach for them - **Tight tail-latency SLOs.** When p99.9 is a contractual number and a 150 ms stop would breach it. Trading a few percent of throughput for eliminating a class of tail spike is a good deal for a request-serving tier. - **Large heaps.** Tens to hundreds of gigabytes — in-memory caches, large data-serving nodes, big JVM-based stores. G1 remains workable well into that range, but its pauses grow and its full-compaction fallback becomes catastrophic. - **Pause predictability over peak throughput.** Systems with timeouts, leases, heartbeats and coordination protocols where a long stop causes not just slow responses but *failover* — a stopped node looks dead. - **When G1 has already been tuned and still misses.** Not as a first move: reaching for ZGC before understanding the current profile skips the diagnosis. ## What they cost **Barrier overhead.** Every access through a reference executes extra instructions so the collector can maintain its invariants while relocating. This is a tax on ordinary application code, not on the collector, and it is why raw throughput on a fixed workload is typically a few to about ten percent below a throughput collector — pointer-chasing, reference-heavy workloads pay the most; array/primitive-heavy compute pays little. **CPU.** Concurrent marking and relocation run on real cores. If the machine (or the container's CPU quota) has no headroom, that work competes with request handling, and you can end up with worse latency than G1 despite shorter pauses. These collectors want spare CPU. **Footprint and headroom.** Relocation needs somewhere to copy to, and concurrent collection needs to finish before allocation exhausts the heap. Both collectors also carry per-object or per-heap metadata overhead. Run them at 95% occupancy and you get *allocation stalls*: application threads blocked waiting for memory because the collector could not keep up. That is the low-pause failure mode, and it is the equivalent of G1's evacuation failure — the same root cause, allocation outrunning collection, wearing different clothes. **Operational maturity.** Fewer engineers have deep experience with them, tooling and folklore are thinner, and vendor builds differ in what they ship. ## Choosing between them, without marketing Both are region-based, fully concurrent, and production-ready (from JDK 15). ZGC is present in mainstream OpenJDK builds and runs generationally by default from JDK 23 (the older single-generation mode was deprecated and then removed), which substantially improved its throughput and CPU efficiency by collecting short-lived objects cheaply instead of tracing the whole heap every cycle. Shenandoah is available in most OpenJDK-based distributions, though some vendor builds omit it. The honest engineering answer is that on a given workload the difference between them is usually smaller than the difference between either and G1, so decide by measurement on your workload and by what your JDK distribution actually ships — not by benchmark folklore. ## How to make the call defensibly 1. **State the requirement numerically**: the latency percentile and its budget, the heap size, the allocation rate. "Low latency" is not a requirement. 2. **Measure the status quo** from GC logs: pause percentiles, Full GC count, throughput. 3. **If G1 is untuned, tune it first** — headroom, pause target, allocation reduction. Many "we need ZGC" cases are really "our heap is too small" or "we allocate humongous objects". 4. **Trial the concurrent collector on the real workload at real load**, with GC logging on. Compare pause percentiles *and* throughput *and* CPU utilisation, and watch specifically for allocation stalls. 5. **Budget the resources**: give it CPU headroom and heap headroom, or it will underperform the collector you replaced. ## Version note ZGC and Shenandoah became production-ready in JDK 15. Generational ZGC arrived in JDK 21 and became ZGC's default mode from JDK 23, with the non-generational mode subsequently removed — so on JDK 21+ these collectors are a far more practical default choice than they were on JDK 11.
- What is an allocation stall, and what does it tell you?It is a thread blocked waiting for memory because the concurrent collector has not freed space fast enough — the low-pause collectors' equivalent of G1's evacuation failure. It means allocation rate outran concurrent collection, usually because the heap has too little headroom, the collector has too few CPU cycles, or the application allocates too fast. The remedy is headroom, CPU, or less allocation, not a smaller pause target.
- A team switches from G1 to ZGC and sees pauses drop but p99 request latency get worse. What is likely happening?The concurrent work has nowhere to run. Marking and relocation need real CPU, and in a container with a tight quota those threads compete with request-handling threads, so requests slow even though nothing is stopped. Barrier overhead adds to it. The fix is CPU headroom (or a larger quota) and heap headroom, and if neither is available the switch was the wrong trade.
- Why is the choice between ZGC and Shenandoah usually less important than the choice to leave G1?Both are region-based, fully concurrent designs with the same headline property — pauses bounded by fixed work rather than heap size — so on a given workload they typically land far closer to each other than either does to G1. The practical deciders are what your JDK distribution ships, what your team can operate, and measured results on your own workload rather than published benchmarks.
saying these in an interview costs you the question
- Calling low-pause collectors strictly better and recommending them by default, ignoring the throughput and CPU cost.
- Claiming they are pauseless — short stop-the-world root scanning and handshakes remain.
- Deploying one into a CPU-constrained container and expecting latency to improve.
- Running them near the heap ceiling and being surprised by allocation stalls.
- Switching collectors before diagnosing whether the current one is simply untuned or the heap undersized.