skip to content

The ZGC collector in HotSpot advertises garbage-collection pauses under a millisecond that do not grow with heap size. Which parts of its collection cycle run concurrently with application threads, and what work is still done inside a stop-the-world pause?

level: middleimportance: must knowfreq 45%

answer

  1. Three pauses: mark start, mark end, relocate start
  2. Pause work = roots only, never heap-proportional
  3. Stack watermarks → concurrent stack scanning (JDK 16)
  4. Remap is folded into the next cycle's mark
  5. Sub-ms pause ≠ sub-ms latency: allocation stalls, CPU

basics

~20 s

ZGC marks, relocates and remaps concurrently with running threads. Only three short root-oriented pauses remain: mark start, mark end, relocate start. Their cost scales with the number of thread roots, not with heap or live-set size, so pauses stay sub-millisecond.

solid answer

~50 s

ZGC's design rule is that **no stop-the-world pause does work proportional to the heap, the live set, or the object count**. Marking the object graph, processing weak references, selecting which pages to evacuate, copying live objects, and repairing references to moved objects all run concurrently with application threads. What stays in a pause is essentially root handling and phase transitions: `Pause Mark Start` (flip the marking color, arm the barriers, start root scanning), `Pause Mark End` (terminate marking, hand off weak references), and `Pause Relocate Start` (flip to the new remap color and fix roots). Since JDK 16 thread stacks are processed concurrently via stack watermarks, so even a process with thousands of threads keeps these pauses in the tens-to-hundreds of microseconds range. The consequence: heap size affects GC *cycle duration* and CPU cost, not pause length. And sub-millisecond GC pauses are not the same as sub-millisecond service latency — allocation stalls and CPU contention with the collector can still show up in the tail.

code

text · 7 lines
text
[3.155s][info][gc,phases] GC(2) Pause Mark Start                     0.023ms
[3.198s][info][gc,phases] GC(2) Concurrent Mark                      42.559ms
[3.198s][info][gc,phases] GC(2) Pause Mark End                       0.021ms
[3.199s][info][gc,phases] GC(2) Concurrent Process Non-Strong Refs    0.674ms
[3.203s][info][gc,phases] GC(2) Concurrent Select Relocation Set      3.401ms
[3.203s][info][gc,phases] GC(2) Pause Relocate Start                  0.019ms
[3.216s][info][gc,phases] GC(2) Concurrent Relocate                  12.930ms

go deeper

for a junior

Know the headline: ZGC does almost all its work while the application runs, so pauses are tiny and roughly constant regardless of heap size.

for a middle

Name the three pauses and what stays in them (roots and phase transitions), and state that marking, relocation and remapping are concurrent.

for a senior

Explain why the pauses are O(roots) — handshakes and concurrent stack scanning — and be honest that the cost moved to CPU, headroom and allocation stalls.

for a principal

Frame it as a latency-vs-throughput and capacity decision: you are buying a bounded pause with CPU and memory headroom, and you must design monitoring and capacity around the stall failure mode rather than around pause histograms.

## What "pause time independent of heap size" actually claims Every collector has to do three kinds of work: find what is live (marking), reclaim what is not, and eventually defragment memory (compaction). Older collectors do at least one of those with all application threads — *mutators*, in GC vocabulary — frozen at a safepoint. If marking or copying happens while frozen, then a bigger live set means a longer freeze; pause time grows with the heap. ZGC inverts the priority. Its stated goal is that **no pause performs work proportional to the heap size, the live-set size, or the number of objects**. Pauses touch only *roots* — the references held in thread stacks, registers, static fields and a few JVM-internal tables. That set is bounded by how many threads you run, not by how much data you keep, so a pause on an 8 GB heap and on an 8 TB heap costs roughly the same. ## The shape of a cycle A ZGC cycle alternates short pauses with long concurrent phases: 1. **Pause Mark Start (STW).** The JVM flips the *marking color* used for this cycle, arms the load barriers, and begins root scanning. Thread stacks are not fully walked here: since JDK 16, each thread carries a *stack watermark*, and frames below it are scanned lazily/concurrently when the thread returns into them. This is why the pause does not scale with thread count or stack depth. 2. **Concurrent Mark / Remap.** Worker threads traverse the object graph from the roots and record liveness by setting mark bits inside the object references themselves. The same traversal simultaneously *remaps* any reference still pointing at an object that was relocated during the previous cycle — remapping is not a separate pass, it is folded into the next cycle's marking. 3. **Pause Mark End (STW).** Marking terminates: drain remaining work, decide that the mark stacks are truly empty, and hand off weak/soft/phantom reference processing. 4. **Concurrent reference processing, relocation-set selection.** Non-strong references are processed, and the collector picks the *relocation set* — the pages holding the most garbage, i.e. the ones where copying the few survivors buys the most free memory. 5. **Pause Relocate Start (STW).** Flip to the new "remapped" color and fix the roots so that references held by threads point into the post-relocation world. 6. **Concurrent Relocate.** Live objects in the relocation set are copied out while mutators keep running; the load barrier makes sure a mutator never works on a stale copy. ## Why the pauses stay small Three ingredients: - **Barriers, not freezing.** Correctness during concurrent marking and copying comes from a barrier compiled into application code (see the load barrier), not from suspending the application. - **Handshakes instead of global safepoints where possible.** ZGC leans on per-thread handshakes, so a slow or blocked thread does not hold the whole VM hostage. - **Concurrent stack scanning.** Root scanning, historically the pause floor for concurrent collectors, was moved out of the pause with stack watermarks. ## Reading it in a log The phase names appear directly in `-Xlog:gc+phases`, which is the quickest way to confirm on a real service that the STW slices really are microseconds while the concurrent phases are tens of milliseconds. ## What the guarantee does not cover A candidate who stops at "pauses are sub-millisecond" misses the tradeoff. ZGC moves the cost, it does not delete it: - Concurrent phases run on GC threads that **compete with your application for CPU**. On a saturated box, GC work becomes application slowdown rather than an application pause. - If allocation outruns reclamation, threads hit **allocation stalls**: a thread that cannot get memory waits until the collector frees some. That is ZGC's failure mode, and it can be far longer than a millisecond. - Total heap needs headroom, because the application keeps allocating throughout the cycle. - Non-GC pauses still exist. Safepoints happen for other reasons, the JIT deoptimizes, the OS pages memory. A sub-millisecond GC pause is a necessary, not sufficient, condition for a sub-millisecond p99.9. ## Cycle duration vs pause duration One more distinction worth stating explicitly: heap size does not affect pause length, but it very much affects how long a *cycle* takes and how much CPU it burns. On a multi-terabyte heap, concurrent marking may run for many seconds. That is fine as long as the collector still finishes before the application exhausts free memory — which is exactly what heap headroom and GC thread count buy you.

  • If the pauses don't depend on heap size, what does depend on heap size?
    The duration and CPU cost of the concurrent phases. Marking a larger live set takes longer and consumes more GC-thread CPU, and a larger heap needs proportionally more relocation work. That matters because the collector must complete a cycle before the application exhausts free memory, so bigger heaps and higher allocation rates need more GC threads and more headroom — not longer pauses.
  • A service on ZGC shows 40 ms latency spikes even though every logged GC pause is under 0.1 ms. Where would you look?
    First at allocation stalls: if the collector cannot keep up, threads block waiting for memory, and that appears in the GC log as stall events rather than as pauses. Next at CPU contention — concurrent GC threads competing with application threads on a saturated machine. Then at non-GC causes: other safepoint operations, JIT deoptimization, OS paging or CPU steal.

Repaving a motorway without closing it: traffic keeps flowing lane by lane, and the only full stops are the seconds needed to move the cones at the start and end of each stage — that cost is the same whether the road is one kilometre or a hundred.

saying these in an interview costs you the question

  • Claiming ZGC has no stop-the-world pauses at all; it has three short ones per cycle.
  • Assuming sub-millisecond GC pauses guarantee sub-millisecond request latency, ignoring allocation stalls and CPU competition.
  • Saying a bigger heap makes ZGC pauses longer — it lengthens concurrent phases, not pauses.
  • Believing the concurrency is free rather than paid for in barrier overhead, GC CPU and heap headroom.

context