HotSpot's Concurrent Mark-Sweep collector (CMS) performed most of its old-generation marking while application threads kept running. Walk through the phases of one of its old-generation cycles and say which of them stopped the world.
answer
- initial mark (STW) → concurrent mark → preclean → remark (STW) → sweep → reset
- only two pauses; young ParNew pauses were separate
- remark was the long one; scheduled after a young GC
- sweep = free lists, no compaction
- floating garbage + CPU cost + write barrier are the price
basics
~20 sInitial mark (stop-the-world, marks roots), concurrent mark, concurrent preclean and abortable preclean, final remark (stop-the-world, catches what mutation changed), concurrent sweep, concurrent reset. Only initial mark and remark paused the application; young collections were separately stop-the-world. Sweeping reclaimed space in place without compacting.
solid answer
~50 sA CMS old-generation cycle ran in six phases: 1. **Initial mark** — stop-the-world, short: mark objects directly reachable from roots and from the young generation. 2. **Concurrent mark** — application running: trace the object graph from those roots. 3. **Concurrent preclean / abortable preclean** — application running: revisit regions the mutator dirtied during marking, and wait for a convenient point to schedule the pause (ideally just after a young collection, so remark has less to scan). 4. **Final remark** — stop-the-world, usually the longest CMS pause: re-scan roots and dirtied cards to fix up everything that changed while marking ran concurrently. 5. **Concurrent sweep** — application running: reclaim unmarked objects onto free lists, **in place, without compaction**. 6. **Concurrent reset** — clear state for the next cycle. Young-generation collection was separate and fully stop-the-world (ParNew, a copying collector). CMS therefore reduced *old-gen* pause time, not pause count — and paid for it with extra CPU, floating garbage, and fragmentation.
code
text · 8 lines[GC (CMS Initial Mark) [1 CMS-initial-mark: 2048000K(4194304K)] 2260311K(6291456K), 0.0231 secs]
[CMS-concurrent-mark-start]
[CMS-concurrent-mark: 1.042/1.055 secs]
[CMS-concurrent-preclean: 0.031/0.033 secs]
[CMS-concurrent-abortable-preclean: 0.402/1.921 secs]
[GC (CMS Final Remark) [1 CMS-remark: 2048000K(4194304K)] 2891004K(6291456K), 0.1843 secs]
[CMS-concurrent-sweep: 0.879/0.892 secs]
[CMS-concurrent-reset: 0.010/0.010 secs]go deeper
Name the two stop-the-world phases and say the heavy marking ran alongside the application.
Give the full phase list in order, explain why remark must stop the world, and note that sweeping uses free lists without moving objects.
Add the operational angle: remark duration tracks young-generation size and mutation rate, preclean schedules the pause after a young collection, and floating garbage plus CPU contention are the standing costs.
Position CMS as the origin of the concurrent-collector skeleton later collectors refined, and discuss the latency-versus-throughput and headroom budget a concurrent collector demands from capacity planning.
## What CMS was for Before CMS, an old-generation collection meant tracing the whole old generation with the application stopped, so the pause scaled with live-set size — seconds on a multi-gigabyte heap. CMS's bet was that most of that tracing could run *concurrently with the application*, leaving only two short stop-the-world phases. It was a generational collector: the young generation was collected by a separate parallel copying collector (ParNew) with ordinary stop-the-world pauses, and CMS handled only the old generation. ## The phases in detail **Initial mark (stop-the-world).** The collector needs a starting set. It marks objects directly reachable from thread stacks, static fields, and other roots, plus objects in the old generation referenced from the young generation. This is a shallow scan, so the pause is short — though it grows if the young generation is large and full of references into the old generation, which is why the collector preferred to run it right after a young collection. **Concurrent mark.** Marking threads trace the reachable graph from the initial set while application threads keep mutating it. This is the phase that buys the low pause: the expensive transitive walk is overlapped with useful work. It also creates the central problem — the heap is changing underneath the marker, so the mark can become wrong. A write barrier records mutations (by dirtying the card containing the modified object) so they can be revisited. **Preclean and abortable preclean.** These concurrent phases process the mutation record that accumulated during marking, shrinking the work left for the pause that follows. The *abortable* variant additionally waits, up to a time limit, for a young collection to happen — remark right after a young collection has far fewer young objects to scan, so scheduling matters. **Final remark (stop-the-world).** Everything that changed during concurrent marking has to be reconciled with the application stopped, otherwise the collector could free a live object. Remark re-scans roots and dirty cards and finishes the marking. This was typically the longest CMS pause — not close to a full-GC pause, but sensitive to young-generation size and mutation rate, which is why remark times were the first thing to look at when CMS pauses grew. **Concurrent sweep.** Unmarked objects are reclaimed and their space is returned to **free lists**. Crucially, live objects are **not moved**. That choice is what makes sweeping concurrent-safe — no application thread can hold a reference that the collector is about to invalidate — and it is also the origin of CMS's defining weakness, fragmentation. **Concurrent reset.** Marking state is cleared, ready for the next cycle. ## What "concurrent" did and did not buy It is worth being precise about the tradeoffs, because interviewers probe them: - **Pauses got shorter, not fewer.** Young collections were still fully stop-the-world, and the two CMS pauses were added on top of them. - **Throughput went down.** Marking threads compete with the application for CPU, and the write barrier adds cost to every reference store. On a CPU-saturated machine, giving cycles to the collector directly slows the application — CMS traded throughput for latency by design. - **Floating garbage.** Objects that die *after* being marked in a concurrent cycle are not collected until the next cycle, so the heap must carry that extra headroom. - **Allocation got slower.** Because free space lives in free lists rather than one contiguous block, old-generation allocation and promotion involve free-list management rather than a pointer bump. ## Reading the log CMS logs named its phases explicitly — `CMS-initial-mark`, `CMS-concurrent-mark`, `CMS-concurrent-abortable-preclean`, `CMS Final Remark`, `CMS-concurrent-sweep` — with concurrent phases showing both CPU time and wall time. Two numbers mattered: the durations of the two stop-the-world phases, and whether cycles were completing before the old generation filled. ## Why the design is still worth knowing CMS was deprecated in JDK 9 and removed in JDK 14, so you will not tune it on a new service. It is asked about because its structure is the ancestor of everything that followed: a short root-scanning pause, a concurrent trace, a fix-up pause, and a reclamation phase is exactly the skeleton of the region-based G1 collector, and the later fully-concurrent collectors are best understood as answers to the specific weaknesses CMS exposed — above all, that a collector which never compacts eventually loses to fragmentation.
- Why did a stop-the-world remark phase exist at all if marking was concurrent?While marking runs concurrently, the application keeps changing references, so a marked graph can become stale: a reference could be moved into an already-scanned object, hiding a live object from the marker. CMS recorded those mutations via a write barrier and reconciled them in remark with all threads stopped, which is the only way to reach a consistent final mark. Without it the collector could free a live object.
- CMS reduced pause duration, but what did it cost?Throughput and headroom. Concurrent marking threads compete with application threads for CPU, and the write barrier adds cost to reference stores, so total work per unit of application progress rises. It also produces floating garbage — objects that die after being marked survive until the next cycle — so the heap must be provisioned larger than the live set to keep cycles finishing in time.
saying these in an interview costs you the question
- Saying CMS was pause-free or had no stop-the-world phases
- Forgetting that young-generation collections were still fully stop-the-world
- Claiming concurrent collection is free rather than a throughput-for-latency trade
- Believing CMS compacted the old generation during its sweep
- Describing remark as trivial — in practice it was the dominant CMS pause