skip to content

Modern JVM collectors perform most of their work concurrently with the running application but still take short stop-the-world pauses. What work is placed in those pauses, and what does moving work out of them cost the system?

level: principalimportance: should knowfreq 34%

answer

  1. pause keeps: root capture, phase/barrier flip, marking termination
  2. handshakes move per-thread root scanning off the global pause
  3. concurrency paid in barriers, CPU, floating garbage, headroom
  4. failure mode shifts: long pause → allocation stall / fallback
  5. sub-ms pauses: the remaining latency is elsewhere

basics

~20 s

Pauses retain the work that needs global agreement or a stable per-thread view: capturing roots from thread stacks, switching collector phase and barrier state, and finishing marking. Moving the rest off the pause costs throughput through barriers on every reference access, CPU shared with the application, floating garbage, and heap headroom so the concurrent cycle can finish before the heap fills.

solid answer

~60 s

**What stays inside a pause.** Work requiring either a globally agreed instant or a stable snapshot of a thread's own state: enumerating roots from thread stacks and registers, flipping collector phase and barrier mode so all threads observe the change consistently, and terminating marking (draining what the application produced while marking ran). Collectors with sub-millisecond pauses push root scanning onto per-thread handshakes so no single global pause scales with heap size, but a global agreement point remains. **What the concurrency costs.** - **Barriers.** Tracing or relocating while the application mutates references requires the application to notify the collector on reference writes, or to be corrected on reads. Every such access pays. - **CPU.** Concurrent threads take cores the application wanted; throughput drops even though pauses shrink. - **Floating garbage.** Objects that die after the snapshot are retained until the next cycle, so the heap must be larger. - **Headroom and pacing.** The cycle must finish before allocation exhausts the heap; if it loses the race, you get allocation stalls or a stop-the-world fallback. So the choice is not "shorter pauses, free", it is pause time bought with CPU and memory.

go deeper

for a junior

Know that concurrent phases run alongside the application and that short pauses remain for work that needs all threads stopped.

for a middle

Name the work that stays in the pause — root capture, phase transitions, marking termination — and that barriers are what make concurrent tracing possible.

for a senior

Price the tradeoff concretely: barrier cost on every reference access, CPU contention, floating garbage, pacing and headroom, and the shift in failure mode to allocation stalls.

for a principal

Turn it into a provisioning decision against a latency objective, note that per-thread handshakes are what decouple pauses from heap size, and redirect effort once pauses are sub-millisecond to safepoint arrival, VM operations and host effects.

## The question behind the question Asking why any pause remains is really asking which parts of collection are *inherently* global, and what price the runtime pays to shrink them. A principal-level answer names the residual work, explains why it resists concurrency, and then prices the alternative honestly. ## What genuinely resists being made concurrent ### 1. Capturing roots Marking starts from roots — references in thread stacks and registers, static fields, JNI handles. Reading a thread's stack requires that thread's state to be describable and stable, which means the thread must be at a safepoint. Historically this was done for all threads inside one global pause, and its cost grew with thread count and stack depth. The modern move is to scan each thread's roots via a **per-thread handshake**: the thread stops at its own next poll, its roots are processed, and it resumes — no global stop. This removes root scanning from the pause budget almost entirely and is a large part of why some collectors advertise pauses independent of heap size. What remains is a much cheaper global agreement. ### 2. Phase and barrier transitions A concurrent collector's correctness depends on all threads agreeing about which phase is active, because the barrier behaviour differs by phase — what a write barrier must record while marking is not what it must do while idle, and a relocation phase changes what a read must do with a reference. Switching that state needs an instant that all threads observe consistently. A short global safepoint is the simplest and cheapest way to establish it. ### 3. Terminating marking While marking runs concurrently, the application keeps mutating references, so the collector must consume whatever the barriers recorded. Because the application can generate new work while the collector drains it, marking needs a termination step where mutation is halted briefly to drain the residue and prove completion. Collectors keep this short by bounding how much can be outstanding, but a final short pause is the standard mechanism. ### 4. Relocation bookkeeping Moving collectors must ensure no thread uses a stale address. Concurrent relocation is possible using read barriers that forward accesses to the new copy, but installing and retiring the forwarding state is again a phase transition, and reference updating must be coordinated with the mutator. ## The costs you pay for concurrency ### Barrier overhead on the application This is the biggest and most persistent cost. To trace concurrently, the collector needs the application to report reference changes; to relocate concurrently, it needs reads to be corrected. Either way, ordinary field accesses in application code carry extra instructions. The cost is distributed across all execution rather than concentrated in a pause — which is exactly the point, but it means a fully stop-the-world collector can achieve higher raw throughput than a low-pause one on the same hardware. ### CPU contention Concurrent work runs on cores. On a saturated host, giving the collector threads means taking them from request handling, so latency can worsen even as pauses shrink. It also feeds back into safepoint arrival times: a starved application thread is slow to reach its next poll. Concurrency assumes spare CPU; without headroom it degrades. ### Floating garbage and the need for headroom Concurrent marking works against a logical snapshot. Objects that become unreachable after the snapshot is taken are not identified in that cycle — they *float* and survive to the next one. So a concurrent collector needs a larger heap for the same workload. Additionally, the cycle must complete before allocation exhausts the remaining space, which requires the collector to start early enough (pacing based on observed allocation rate) and to be given headroom to absorb prediction error. ### A worse failure mode A stop-the-world collector's bad day is a long pause. A concurrent collector's bad day is that the cycle loses the race: threads are throttled or stalled waiting for memory, or the collector falls back to a degenerate or full stop-the-world collection — which is both long *and* unexpected. Failure is less frequent but sharper, which matters for how you alarm and how much headroom you keep. ### Complexity and observability More moving parts, more heuristics, more failure modes. The relevant metric shifts from "pause length" to "did the cycle keep up", and monitoring must follow — allocation rate versus reclamation rate, cycle start time, stall counts. ## How to reason about the tradeoff The honest framing is a three-way budget between **pause time, throughput and memory**, at a given CPU allocation: - A batch job that cares only about total completion time is often best served by putting all the work in pauses: no barriers, no floating garbage, minimal heap. - A latency-sensitive service with tail-latency objectives buys short pauses with barrier overhead, extra heap and CPU headroom, and accepts lower peak throughput. - Under-provisioning CPU or heap while selecting a concurrent collector is the common mistake: it produces the worst of both — reduced throughput *and* stalls when the cycle cannot finish. And the residual pause is rarely the binding constraint once it is sub-millisecond. At that point the practical limits on stopped time are usually elsewhere: arrival at safepoints under CPU pressure, non-GC VM operations, and host-level effects such as page faults. A principal answer says so, because it directs effort where the remaining latency actually lives. ## Summary Pauses retain the work needing global agreement or stable per-thread state: root capture, phase and barrier transitions, and marking termination — with handshakes moving root scanning out of the global pause. The concurrency is paid for in barrier overhead, CPU, floating garbage and headroom, and it changes the failure mode from a long pause to an allocation stall. Choose the point on that curve deliberately, and provision for it.

  • Why does a concurrent collector generally need a larger heap than a fully stop-the-world one for the same workload?
    Two reasons. Objects that die after the marking snapshot is taken are not identified during that cycle and float until the next one, so more dead objects are resident at any instant. And the collector must start its cycle early enough to finish before allocation exhausts free space, which requires headroom to absorb errors in its allocation-rate prediction. Too little headroom turns into allocation stalls or a stop-the-world fallback.
  • If a collector's pauses are already well under a millisecond, where does the application's remaining stopped time usually come from?
    Mostly from outside the collector's own work: threads being slow to reach safepoints under CPU throttling or oversubscription, page faults on a swapping host, non-GC VM operations such as deoptimization or class redefinition, and the sheer rate of safepoints. Continuing to tune the collector at that point yields nothing; the leverage is in CPU headroom, memory residency and reducing safepoint-requiring operations.
  • When is a fully stop-the-world collector the better engineering choice?
    When total throughput or footprint matters more than tail latency — batch processing, short-lived jobs, CPU- or memory-constrained environments. Without barriers the application runs faster, the heap can be smaller because there is no floating garbage, and the collector needs no spare cores. The cost is pauses that grow with the live set, which such workloads can absorb.

Renovating a shop while it trades: closing for a week is fastest and cheapest, but customers see a long outage. Staying open means barriers, extra staff and slower work — you trade total efficiency for never being fully shut, and you must keep enough spare capacity that trading never outruns the renovation.

saying these in an interview costs you the question

  • Claiming a modern collector eliminates pauses entirely.
  • Presenting low-pause collectors as strictly better, without naming barrier overhead, CPU cost or heap headroom.
  • Assuming concurrent phases suspend application threads.
  • Believing pause length is determined only by heap size, ignoring safepoint arrival and non-GC VM operations.
  • Selecting a concurrent collector while provisioning CPU and heap as if it were a stop-the-world one.

context