You are picking a garbage collector for a latency-sensitive JVM service and are considering ZGC, the sub-millisecond colored-pointer collector. What do you gain, what do you pay for it, and what failure mode should you size capacity around?
answer
- Buy: pause independent of heap/live set; concurrent compaction
- Pay: load (+store) barrier throughput, GC CPU, headroom, no compressed oops
- Failure mode = ALLOCATION STALLS, not long pauses
- Alert on stalls, not just pause histograms
- Sub-ms GC pause ≠ sub-ms p99.9; safepoints, JIT, OS remain
basics
~20 sYou gain GC pauses under a millisecond regardless of heap or live-set size. You pay a barrier cost on reference loads, CPU for concurrent GC threads, extra heap headroom, and the loss of compressed references. The failure mode is allocation stalls, not long pauses — size CPU and heap for that.
solid answer
~50 s**Gain:** pause time decoupled from heap and live-set size. That is what makes multi-hundred-gigabyte or terabyte heaps operationally viable and what flattens the GC contribution to p99.9 latency. **Pay:** - A load barrier on every heap reference read — a modest throughput tax, worst for pointer-chasing workloads, negligible for I/O-bound or primitive-heavy ones. - Concurrent GC threads competing with application threads for CPU; on a saturated machine, GC work becomes application slowdown. - Heap headroom, because the application keeps allocating throughout a cycle and relocation needs somewhere to copy into. - No compressed object references, so pointer-dense heaps carry a footprint penalty versus G1 below ~32 GB. **Failure mode:** when allocation outruns reclamation, threads **stall** waiting for memory. Plan capacity so the collector always wins that race: generous headroom, spare cores, generational ZGC, and alerting on stall events in the GC log rather than only on pause histograms. And be clear that sub-millisecond *GC* pauses do not by themselves buy a sub-millisecond service tail.
go deeper
Know the shape of the trade: ZGC gives very short pauses even on huge heaps, and costs some throughput and extra memory.
Be able to list the concrete costs — barrier overhead, GC CPU, heap headroom, no compressed references — and name allocation stalls as the failure mode.
Turn it into an operational plan: establish the live set from logs, size headroom and CPU, alert on stalls, and reduce allocation rate when they appear.
Drive the decision from the SLO and cost model: state what tail budget the pause reduction protects, price the extra CPU and RAM, insist on a measured comparison, and set the boundary of the claim so the team does not expect ZGC to fix non-GC latency.
## Frame the decision honestly A collector choice is a resource trade, not a ranking. ZGC buys one property — **pause time independent of heap size and live-set size** — and charges for it in throughput, CPU, and memory. The question to answer is whether that property is on your critical path. It usually is when: your service has a strict tail-latency SLO (p99.9 in single-digit or low tens of milliseconds); your live set is large enough that a stop-the-world compaction would be measured in hundreds of milliseconds or seconds; or your heap is large enough (hundreds of GB and up) that other collectors' pauses become unmanageable. It usually is not when the workload is batch or throughput-oriented, where a throughput-first collector gets more work done per core and nobody is watching a latency histogram. ## What you gain, precisely - **Bounded pauses.** Sub-millisecond stop-the-world phases whose cost tracks the root set, not the heap. A 500 GB heap pauses like a 5 GB one. - **Concurrent compaction.** Defragmentation happens every cycle without a pause, so you do not accumulate fragmentation debt that eventually forces an emergency full collection. - **Headroom to grow the heap.** Because pauses do not lengthen, adding memory is a viable answer to allocation pressure, which is not true of collectors whose pauses scale with the live set. - **Very little tuning surface.** Practically: choose the collector, set a heap size, provide CPU. There is no generation-sizing exercise. ## What you pay, precisely **Throughput.** Every reference load from the heap carries a barrier check. Measured cost is workload-dependent: often a few percent, more for code that chases pointers through large data structures in tight loops, near zero for services dominated by I/O or primitive computation. Generational ZGC also adds a store barrier. If your service is CPU-bound on object graph traversal, benchmark before committing. **CPU.** Concurrent collection means GC threads run *alongside* your threads. On a box already at 90 % CPU, the collector cannot get the cycles it needs and the effect surfaces as latency, then as stalls. ZGC deployments need genuine CPU headroom; treat GC threads as a first-class consumer in capacity planning, not as free background work. **Memory.** Two distinct costs. First, **headroom**: the application allocates during the whole cycle and relocation needs empty pages to copy survivors into, so a heap sized to just fit the live set will stall. Second, **no compressed oops**: every reference is 8 bytes, where G1 below roughly 32 GB uses 4. For reference-dense heaps that is a real footprint increase and some cache-density loss. Generational ZGC substantially reduced the headroom requirement, but it did not remove it. ## The failure mode to design around This is the part that separates a considered answer from a memorized one. Collectors that pause degrade *gracefully-looking* — pauses simply get longer. ZGC does not: it holds pauses at sub-millisecond and instead, when it cannot free memory fast enough, makes allocating threads **stall** until memory becomes available. A stall can be orders of magnitude longer than any pause. So the operational discipline is: 1. **Alert on allocation stalls**, not only on pause times. A dashboard showing max pause = 0.08 ms tells you nothing about whether threads are waiting for memory. 2. **Provision headroom deliberately.** Establish the steady-state live set from GC logs and size the heap well above it, with enough margin to absorb traffic spikes and any known bursty allocation (large request payloads, batch jobs colocated with request handling). 3. **Provision CPU.** If GC threads cannot run, nothing else in the design saves you. 4. **Use generational ZGC** (the default on current JDKs) — it needs materially less CPU and headroom for the same allocation rate. 5. **Attack allocation rate itself** when stalls appear. Reducing garbage produced per request is often cheaper than adding memory, and it is the only fix that scales. ## Do not oversell the latency claim A sub-millisecond GC pause is one component of tail latency. The remainder still contains: other safepoint operations unrelated to GC; JIT compilation and deoptimization; OS-level effects such as page faults, swap, CPU steal on shared hosts, and NUMA misplacement; lock contention and queueing in your own code. Teams that migrate to ZGC and still see 50 ms tails are usually looking at one of those. State the boundary of the claim: ZGC removes GC pauses from your tail budget; it does not make the rest of the system real-time. ## How to decide in practice Measure, do not assume. Run the workload with the current collector and with ZGC, at realistic allocation rates and heap sizes, and compare three numbers: throughput (requests/sec at fixed CPU), the latency tail (p99.9, not the mean), and stability under a spike. Then price the difference — extra cores and extra RAM against the SLO the pause reduction protects. If nobody can name the SLO the change protects, the answer is that you do not need it yet.
- Your ZGC service starts showing multi-hundred-millisecond latency spikes under peak load, with GC pauses still logged at well under a millisecond. What is your first hypothesis and what would you change?Allocation stalls: the collector is losing the race against allocation, so threads block waiting for free memory. Confirm it in the GC log, which reports stall events distinctly from pauses. Remedies in order of cost: give the heap more headroom, give the JVM more CPU so concurrent GC threads can keep up, ensure generational ZGC is in use, and then reduce the allocation rate per request, which is the only fix that scales with traffic.
- A 6 GB-heap service with a 20 ms p99 SLO is currently on G1 and meeting it. Would you move it to ZGC?Probably not by default. At that heap size G1's pauses are typically well inside the budget, and moving costs throughput from the barriers plus footprint from losing compressed references. Move only if measurements show G1 pauses actually consuming a meaningful share of the tail budget, or if the heap is expected to grow substantially. The decision should be driven by a measured tail-latency contribution, not by the collector's headline pause number.
Hiring a cleaning crew that works while the shop stays open: customers are never locked out, but the crew occupies floor space and staff, and if the shop gets busier than the crew can handle, the queue at the till — not a closed door — is what customers feel.
saying these in an interview costs you the question
- Recommending ZGC purely because it has the lowest advertised pause, with no SLO or measurement behind the choice.
- Assuming concurrent collection is free rather than paid for in CPU that competes with application threads.
- Sizing the heap to just fit the live set, ignoring that the application allocates throughout a whole cycle.
- Monitoring only pause histograms and never allocation stalls, then being surprised by latency spikes.
- Forgetting the footprint cost of losing compressed references on mid-size heaps.
- Promising sub-millisecond end-to-end latency because GC pauses are sub-millisecond.