The ZGC collector in HotSpot advertises garbage-collection pauses under a millisecond that do not grow with heap size. Which parts of its collection cycle run concurrently with application threads, and what work is still done inside a stop-the-world pause?
answer
- Three pauses: mark start, mark end, relocate start
- Pause work = roots only, never heap-proportional
- Stack watermarks → concurrent stack scanning (JDK 16)
- Remap is folded into the next cycle's mark
- Sub-ms pause ≠ sub-ms latency: allocation stalls, CPU
basics
~20 sZGC marks, relocates and remaps concurrently with running threads. Only three short root-oriented pauses remain: mark start, mark end, relocate start. Their cost scales with the number of thread roots, not with heap or live-set size, so pauses stay sub-millisecond.
solid answer
~50 sZGC's design rule is that **no stop-the-world pause does work proportional to the heap, the live set, or the object count**. Marking the object graph, processing weak references, selecting which pages to evacuate, copying live objects, and repairing references to moved objects all run concurrently with application threads. What stays in a pause is essentially root handling and phase transitions: `Pause Mark Start` (flip the marking color, arm the barriers, start root scanning), `Pause Mark End` (terminate marking, hand off weak references), and `Pause Relocate Start` (flip to the new remap color and fix roots). Since JDK 16 thread stacks are processed concurrently via stack watermarks, so even a process with thousands of threads keeps these pauses in the tens-to-hundreds of microseconds range. The consequence: heap size affects GC *cycle duration* and CPU cost, not pause length. And sub-millisecond GC pauses are not the same as sub-millisecond service latency — allocation stalls and CPU contention with the collector can still show up in the tail.
code
text · 7 lines[3.155s][info][gc,phases] GC(2) Pause Mark Start 0.023ms
[3.198s][info][gc,phases] GC(2) Concurrent Mark 42.559ms
[3.198s][info][gc,phases] GC(2) Pause Mark End 0.021ms
[3.199s][info][gc,phases] GC(2) Concurrent Process Non-Strong Refs 0.674ms
[3.203s][info][gc,phases] GC(2) Concurrent Select Relocation Set 3.401ms
[3.203s][info][gc,phases] GC(2) Pause Relocate Start 0.019ms
[3.216s][info][gc,phases] GC(2) Concurrent Relocate 12.930msgo deeper
Know the headline: ZGC does almost all its work while the application runs, so pauses are tiny and roughly constant regardless of heap size.
Name the three pauses and what stays in them (roots and phase transitions), and state that marking, relocation and remapping are concurrent.
Explain why the pauses are O(roots) — handshakes and concurrent stack scanning — and be honest that the cost moved to CPU, headroom and allocation stalls.
Frame it as a latency-vs-throughput and capacity decision: you are buying a bounded pause with CPU and memory headroom, and you must design monitoring and capacity around the stall failure mode rather than around pause histograms.
## What "pause time independent of heap size" actually claims Every collector has to do three kinds of work: find what is live (marking), reclaim what is not, and eventually defragment memory (compaction). Older collectors do at least one of those with all application threads — *mutators*, in GC vocabulary — frozen at a safepoint. If marking or copying happens while frozen, then a bigger live set means a longer freeze; pause time grows with the heap. ZGC inverts the priority. Its stated goal is that **no pause performs work proportional to the heap size, the live-set size, or the number of objects**. Pauses touch only *roots* — the references held in thread stacks, registers, static fields and a few JVM-internal tables. That set is bounded by how many threads you run, not by how much data you keep, so a pause on an 8 GB heap and on an 8 TB heap costs roughly the same. ## The shape of a cycle A ZGC cycle alternates short pauses with long concurrent phases: 1. **Pause Mark Start (STW).** The JVM flips the *marking color* used for this cycle, arms the load barriers, and begins root scanning. Thread stacks are not fully walked here: since JDK 16, each thread carries a *stack watermark*, and frames below it are scanned lazily/concurrently when the thread returns into them. This is why the pause does not scale with thread count or stack depth. 2. **Concurrent Mark / Remap.** Worker threads traverse the object graph from the roots and record liveness by setting mark bits inside the object references themselves. The same traversal simultaneously *remaps* any reference still pointing at an object that was relocated during the previous cycle — remapping is not a separate pass, it is folded into the next cycle's marking. 3. **Pause Mark End (STW).** Marking terminates: drain remaining work, decide that the mark stacks are truly empty, and hand off weak/soft/phantom reference processing. 4. **Concurrent reference processing, relocation-set selection.** Non-strong references are processed, and the collector picks the *relocation set* — the pages holding the most garbage, i.e. the ones where copying the few survivors buys the most free memory. 5. **Pause Relocate Start (STW).** Flip to the new "remapped" color and fix the roots so that references held by threads point into the post-relocation world. 6. **Concurrent Relocate.** Live objects in the relocation set are copied out while mutators keep running; the load barrier makes sure a mutator never works on a stale copy. ## Why the pauses stay small Three ingredients: - **Barriers, not freezing.** Correctness during concurrent marking and copying comes from a barrier compiled into application code (see the load barrier), not from suspending the application. - **Handshakes instead of global safepoints where possible.** ZGC leans on per-thread handshakes, so a slow or blocked thread does not hold the whole VM hostage. - **Concurrent stack scanning.** Root scanning, historically the pause floor for concurrent collectors, was moved out of the pause with stack watermarks. ## Reading it in a log The phase names appear directly in `-Xlog:gc+phases`, which is the quickest way to confirm on a real service that the STW slices really are microseconds while the concurrent phases are tens of milliseconds. ## What the guarantee does not cover A candidate who stops at "pauses are sub-millisecond" misses the tradeoff. ZGC moves the cost, it does not delete it: - Concurrent phases run on GC threads that **compete with your application for CPU**. On a saturated box, GC work becomes application slowdown rather than an application pause. - If allocation outruns reclamation, threads hit **allocation stalls**: a thread that cannot get memory waits until the collector frees some. That is ZGC's failure mode, and it can be far longer than a millisecond. - Total heap needs headroom, because the application keeps allocating throughout the cycle. - Non-GC pauses still exist. Safepoints happen for other reasons, the JIT deoptimizes, the OS pages memory. A sub-millisecond GC pause is a necessary, not sufficient, condition for a sub-millisecond p99.9. ## Cycle duration vs pause duration One more distinction worth stating explicitly: heap size does not affect pause length, but it very much affects how long a *cycle* takes and how much CPU it burns. On a multi-terabyte heap, concurrent marking may run for many seconds. That is fine as long as the collector still finishes before the application exhausts free memory — which is exactly what heap headroom and GC thread count buy you.
- If the pauses don't depend on heap size, what does depend on heap size?The duration and CPU cost of the concurrent phases. Marking a larger live set takes longer and consumes more GC-thread CPU, and a larger heap needs proportionally more relocation work. That matters because the collector must complete a cycle before the application exhausts free memory, so bigger heaps and higher allocation rates need more GC threads and more headroom — not longer pauses.
- A service on ZGC shows 40 ms latency spikes even though every logged GC pause is under 0.1 ms. Where would you look?First at allocation stalls: if the collector cannot keep up, threads block waiting for memory, and that appears in the GC log as stall events rather than as pauses. Next at CPU contention — concurrent GC threads competing with application threads on a saturated machine. Then at non-GC causes: other safepoint operations, JIT deoptimization, OS paging or CPU steal.
Repaving a motorway without closing it: traffic keeps flowing lane by lane, and the only full stops are the seconds needed to move the cones at the start and end of each stage — that cost is the same whether the road is one kilometre or a hundred.
saying these in an interview costs you the question
- Claiming ZGC has no stop-the-world pauses at all; it has three short ones per cycle.
- Assuming sub-millisecond GC pauses guarantee sub-millisecond request latency, ignoring allocation stalls and CPU competition.
- Saying a bigger heap makes ZGC pauses longer — it lengthens concurrent phases, not pauses.
- Believing the concurrency is free rather than paid for in barrier overhead, GC CPU and heap headroom.