skip to content

Low-Latency Colored-Pointer Collector (ZGC)

A concurrent collector aiming at sub-millisecond pauses regardless of heap size, using colored pointers and load barriers to relocate objects while the application keeps running. Interviewers bring it up for latency-sensitive or very large-heap services, and to see whether you also know its throughput and footprint cost.

on this pageshow

questions

6

The ZGC collector in HotSpot advertises garbage-collection pauses under a millisecond that do not grow with heap size. Which parts of its collection cycle run concurrently with application threads, and what work is still done inside a stop-the-world pause?

level: middleimportance: must knowfreq 45%

answer

  1. Three pauses: mark start, mark end, relocate start
  2. Pause work = roots only, never heap-proportional
  3. Stack watermarks → concurrent stack scanning (JDK 16)
  4. Remap is folded into the next cycle's mark
  5. Sub-ms pause ≠ sub-ms latency: allocation stalls, CPU

basics

~20 s

ZGC marks, relocates and remaps concurrently with running threads. Only three short root-oriented pauses remain: mark start, mark end, relocate start. Their cost scales with the number of thread roots, not with heap or live-set size, so pauses stay sub-millisecond.

solid answer

~50 s

ZGC's design rule is that **no stop-the-world pause does work proportional to the heap, the live set, or the object count**. Marking the object graph, processing weak references, selecting which pages to evacuate, copying live objects, and repairing references to moved objects all run concurrently with application threads. What stays in a pause is essentially root handling and phase transitions: `Pause Mark Start` (flip the marking color, arm the barriers, start root scanning), `Pause Mark End` (terminate marking, hand off weak references), and `Pause Relocate Start` (flip to the new remap color and fix roots). Since JDK 16 thread stacks are processed concurrently via stack watermarks, so even a process with thousands of threads keeps these pauses in the tens-to-hundreds of microseconds range. The consequence: heap size affects GC *cycle duration* and CPU cost, not pause length. And sub-millisecond GC pauses are not the same as sub-millisecond service latency — allocation stalls and CPU contention with the collector can still show up in the tail.

code

text · 7 lines
text
[3.155s][info][gc,phases] GC(2) Pause Mark Start                     0.023ms
[3.198s][info][gc,phases] GC(2) Concurrent Mark                      42.559ms
[3.198s][info][gc,phases] GC(2) Pause Mark End                       0.021ms
[3.199s][info][gc,phases] GC(2) Concurrent Process Non-Strong Refs    0.674ms
[3.203s][info][gc,phases] GC(2) Concurrent Select Relocation Set      3.401ms
[3.203s][info][gc,phases] GC(2) Pause Relocate Start                  0.019ms
[3.216s][info][gc,phases] GC(2) Concurrent Relocate                  12.930ms

go deeper

for a junior

Know the headline: ZGC does almost all its work while the application runs, so pauses are tiny and roughly constant regardless of heap size.

for a middle

Name the three pauses and what stays in them (roots and phase transitions), and state that marking, relocation and remapping are concurrent.

for a senior

Explain why the pauses are O(roots) — handshakes and concurrent stack scanning — and be honest that the cost moved to CPU, headroom and allocation stalls.

for a principal

Frame it as a latency-vs-throughput and capacity decision: you are buying a bounded pause with CPU and memory headroom, and you must design monitoring and capacity around the stall failure mode rather than around pause histograms.

## What "pause time independent of heap size" actually claims Every collector has to do three kinds of work: find what is live (marking), reclaim what is not, and eventually defragment memory (compaction). Older collectors do at least one of those with all application threads — *mutators*, in GC vocabulary — frozen at a safepoint. If marking or copying happens while frozen, then a bigger live set means a longer freeze; pause time grows with the heap. ZGC inverts the priority. Its stated goal is that **no pause performs work proportional to the heap size, the live-set size, or the number of objects**. Pauses touch only *roots* — the references held in thread stacks, registers, static fields and a few JVM-internal tables. That set is bounded by how many threads you run, not by how much data you keep, so a pause on an 8 GB heap and on an 8 TB heap costs roughly the same. ## The shape of a cycle A ZGC cycle alternates short pauses with long concurrent phases: 1. **Pause Mark Start (STW).** The JVM flips the *marking color* used for this cycle, arms the load barriers, and begins root scanning. Thread stacks are not fully walked here: since JDK 16, each thread carries a *stack watermark*, and frames below it are scanned lazily/concurrently when the thread returns into them. This is why the pause does not scale with thread count or stack depth. 2. **Concurrent Mark / Remap.** Worker threads traverse the object graph from the roots and record liveness by setting mark bits inside the object references themselves. The same traversal simultaneously *remaps* any reference still pointing at an object that was relocated during the previous cycle — remapping is not a separate pass, it is folded into the next cycle's marking. 3. **Pause Mark End (STW).** Marking terminates: drain remaining work, decide that the mark stacks are truly empty, and hand off weak/soft/phantom reference processing. 4. **Concurrent reference processing, relocation-set selection.** Non-strong references are processed, and the collector picks the *relocation set* — the pages holding the most garbage, i.e. the ones where copying the few survivors buys the most free memory. 5. **Pause Relocate Start (STW).** Flip to the new "remapped" color and fix the roots so that references held by threads point into the post-relocation world. 6. **Concurrent Relocate.** Live objects in the relocation set are copied out while mutators keep running; the load barrier makes sure a mutator never works on a stale copy. ## Why the pauses stay small Three ingredients: - **Barriers, not freezing.** Correctness during concurrent marking and copying comes from a barrier compiled into application code (see the load barrier), not from suspending the application. - **Handshakes instead of global safepoints where possible.** ZGC leans on per-thread handshakes, so a slow or blocked thread does not hold the whole VM hostage. - **Concurrent stack scanning.** Root scanning, historically the pause floor for concurrent collectors, was moved out of the pause with stack watermarks. ## Reading it in a log The phase names appear directly in `-Xlog:gc+phases`, which is the quickest way to confirm on a real service that the STW slices really are microseconds while the concurrent phases are tens of milliseconds. ## What the guarantee does not cover A candidate who stops at "pauses are sub-millisecond" misses the tradeoff. ZGC moves the cost, it does not delete it: - Concurrent phases run on GC threads that **compete with your application for CPU**. On a saturated box, GC work becomes application slowdown rather than an application pause. - If allocation outruns reclamation, threads hit **allocation stalls**: a thread that cannot get memory waits until the collector frees some. That is ZGC's failure mode, and it can be far longer than a millisecond. - Total heap needs headroom, because the application keeps allocating throughout the cycle. - Non-GC pauses still exist. Safepoints happen for other reasons, the JIT deoptimizes, the OS pages memory. A sub-millisecond GC pause is a necessary, not sufficient, condition for a sub-millisecond p99.9. ## Cycle duration vs pause duration One more distinction worth stating explicitly: heap size does not affect pause length, but it very much affects how long a *cycle* takes and how much CPU it burns. On a multi-terabyte heap, concurrent marking may run for many seconds. That is fine as long as the collector still finishes before the application exhausts free memory — which is exactly what heap headroom and GC thread count buy you.

  • If the pauses don't depend on heap size, what does depend on heap size?
    The duration and CPU cost of the concurrent phases. Marking a larger live set takes longer and consumes more GC-thread CPU, and a larger heap needs proportionally more relocation work. That matters because the collector must complete a cycle before the application exhausts free memory, so bigger heaps and higher allocation rates need more GC threads and more headroom — not longer pauses.
  • A service on ZGC shows 40 ms latency spikes even though every logged GC pause is under 0.1 ms. Where would you look?
    First at allocation stalls: if the collector cannot keep up, threads block waiting for memory, and that appears in the GC log as stall events rather than as pauses. Next at CPU contention — concurrent GC threads competing with application threads on a saturated machine. Then at non-GC causes: other safepoint operations, JIT deoptimization, OS paging or CPU steal.

Repaving a motorway without closing it: traffic keeps flowing lane by lane, and the only full stops are the seconds needed to move the cones at the start and end of each stage — that cost is the same whether the road is one kilometre or a hundred.

saying these in an interview costs you the question

  • Claiming ZGC has no stop-the-world pauses at all; it has three short ones per cycle.
  • Assuming sub-millisecond GC pauses guarantee sub-millisecond request latency, ignoring allocation stalls and CPU competition.
  • Saying a bigger heap makes ZGC pauses longer — it lengthens concurrent phases, not pauses.
  • Believing the concurrency is free rather than paid for in barrier overhead, GC CPU and heap headroom.

context

open as a page

Generational ZGC shipped as an option in JDK 21 and became the default in JDK 23. What does splitting a concurrent, colored-pointer collector into young and old generations actually change in its implementation, and which production problem motivated the change?

level: middleimportance: should knowfreq 35%

basics

~20 s

It lets ZGC collect a small young space frequently instead of marking the whole heap each cycle. That required a new store barrier and remembered sets to track old-to-young references. The payoff: much less CPU and heap headroom for a given allocation rate, and fewer allocation stalls.

open as a page

ZGC stores garbage-collection metadata inside the 64-bit object reference itself — the technique known as colored pointers. What is kept in those bits, how does the JVM still dereference such a pointer correctly, and what constraints does the technique impose on the platform and on memory footprint?

level: seniorimportance: should knowfreq 40%

basics

~20 s

ZGC puts GC state — marked, remapped, and in the generational design generation and remembered-set bits — into unused bits of every 64-bit heap reference. Barriers interpret and strip the color before use. It requires 64-bit addressing, bounds the address space, and rules out compressed oops.

open as a page

ZGC compacts the heap by moving live objects while application threads keep reading and writing them. Describe how it chooses what to move, how it prevents threads from working on a stale copy, and at what point the old memory can be handed back for reuse.

level: seniorimportance: should knowfreq 35%

basics

~20 s

ZGC selects the pages holding the most garbage as the relocation set, then copies their live objects concurrently. Off-heap forwarding tables map old to new addresses, and load barriers redirect — or perform — each move via compare-and-set. Pages are freed immediately; stale references get remapped lazily.

open as a page

HotSpot's ZGC compiles a load barrier into application code. When does that barrier run, what does its fast path test, what can happen on the slow path, and what does it mean that the barrier is 'self-healing'?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Every read of a reference field from the heap runs a load barrier. The fast path tests the pointer's color against the current good color; a bad color takes a slow path that marks the object or looks up its new location, then writes the corrected pointer back into the field it was read from — self-healing.

open as a page

You are picking a garbage collector for a latency-sensitive JVM service and are considering ZGC, the sub-millisecond colored-pointer collector. What do you gain, what do you pay for it, and what failure mode should you size capacity around?

level: principalimportance: should knowfreq 30%

basics

~20 s

You gain GC pauses under a millisecond regardless of heap or live-set size. You pay a barrier cost on reference loads, CPU for concurrent GC threads, extra heap headroom, and the loss of compressed references. The failure mode is allocation stalls, not long pauses — size CPU and heap for that.

open as a page