skip to content

Concurrent Mark-Sweep Collector (Historical)

The historical low-pause collector that marked concurrently but never compacted, so fragmentation could trigger a concurrent-mode failure and a long full collection. It is deprecated and removed, but interviewers still ask because it explains precisely which problem the region-based collector was built to solve.

on this pageshow

questions

4

HotSpot's Concurrent Mark-Sweep collector reclaimed old-generation space without compacting it. Explain what that caused over time, and what a "concurrent mode failure" in its logs meant for the application.

level: seniorimportance: must knowfreq 48%

answer

  1. sweep in place → free lists → fragmentation
  2. free bytes ≠ contiguous space; watch largest free chunk
  3. promotion failed / concurrent mode failure
  4. fallback = serial stop-the-world mark-sweep-compact, seconds
  5. mitigate with earlier cycles + headroom; never eliminated

basics

~20 s

Sweeping in place left free space scattered across free lists, so the old generation fragmented. Eventually a promotion needed a contiguous block bigger than any free chunk even though total free bytes were ample. That triggered a fallback to a single-threaded stop-the-world mark-sweep-compact full collection — a pause of many seconds, the exact thing CMS existed to avoid.

solid answer

~1 min

CMS's concurrent sweep freed dead objects **in place**: it returned their space to free lists and never moved surviving objects. That is what makes sweeping safe to run alongside the application — no reference an application thread holds ever becomes stale — but it means the old generation's free space slowly becomes many small, non-contiguous chunks. The failure mode is a promotion that needs a contiguous block larger than any available chunk, even with plenty of total free bytes. Related, a cycle can also simply lose the race: allocation fills the old generation before the concurrent cycle finishes. Either way the JVM falls back to a **full collection performed by the serial mark-sweep-compact collector** — stop-the-world, and in older JDKs single-threaded, so on a large old generation the pause could run into many seconds or worse. It logged as `concurrent mode failure` (or `promotion failed`). The irony is structural: a collector chosen for low pauses had, as its worst case, one of the longest pauses in HotSpot. Mitigations — starting cycles earlier, more headroom, occasional compaction at a full GC — reduced the frequency but could never remove the failure mode, which is a large part of why a compacting, region-based collector replaced it.

code

text · 3 lines
text
[GC (Allocation Failure) [ParNew (promotion failed): 1887488K->1887488K(1887488K), 0.412 secs]
  [CMS (concurrent mode failure): 4180221K->2103998K(4194304K), 6.7412 secs]
  6067709K->2103998K(6081792K), [Metaspace: 98213K->98213K(1140736K)], 7.1585 secs]

go deeper

for a junior

Say that not moving objects leaves free space in scattered pieces, so a big object may not fit even when total free space looks fine.

for a middle

Distinguish promotion failure from losing the race to fill the old generation, and name the fallback as a stop-the-world compacting full collection.

for a senior

Describe the diagnosis from GC logs, the largest-free-chunk metric, and why the standard mitigations only reduce probability; connect it to the tail-latency incident shape.

for a principal

Argue that capacity planning must be driven by the fallback's worst case rather than the median pause, and use that to justify moving to a collector that compacts as part of normal operation.

## Why CMS did not compact Compaction moves live objects together so free space becomes one contiguous block. Moving an object means updating every reference to it — which, if the application is running, means the application could observe a stale reference mid-move. Doing this concurrently requires machinery CMS did not have (the later concurrent-compaction collectors solve it with load or store barriers and forwarding information). CMS made the simpler choice: **mark and sweep, never move**. Dead objects' space is returned to size-segregated free lists, and live objects stay exactly where they are. That single decision explains most of CMS's operational character. ## Fragmentation An allocator that never moves anything ends up with free space interleaved with live objects. Over hours or days of mixed-size allocation and promotion, the old generation's free space becomes many small chunks rather than one big one. The collector coalesces adjacent free chunks where it can, and it keeps statistics to split chunks sensibly, but it cannot manufacture contiguity that the object layout does not permit. The crucial consequence is that **free bytes stop predicting allocation success**. A promotion of a large object — say a several-hundred-kilobyte array being copied out of the young generation — needs a contiguous run. If the largest free chunk is smaller than the object, the promotion fails, even if the old generation reports gigabytes free. This is why teams running CMS learned to watch the *largest free chunk* rather than just the occupancy percentage. ## The two failure paths **Promotion failed.** A young collection is under way, survivors must be promoted into the old generation, and there is no chunk big enough. The young collection cannot complete as planned. **Concurrent mode failure.** The old generation fills before the concurrent cycle finishes. This happens when the cycle started too late for the allocation rate, when fragmentation left less usable space than the occupancy figure suggested, or when the concurrent threads were starved of CPU. ## What the fallback cost In both cases the JVM abandons the concurrent strategy for that moment and performs a **full collection with the serial mark-sweep-compact collector**: all application threads stopped, the entire heap traced, and — finally — compacted. Compaction is the point: it restores contiguity so the application can continue. But in older JDKs this fallback was single-threaded, so its duration scaled with the whole live set with no parallelism to help. On a multi-gigabyte old generation that meant pauses measured in many seconds, sometimes tens of seconds. Operationally this is the worst possible shape: a service that normally shows tens of milliseconds of pause occasionally shows a multi-second freeze. Health checks time out, load balancers eject the instance, queues back up, and the incident looks like a hang rather than a garbage-collection event. Diagnosing it means finding `concurrent mode failure` or `promotion failed` in the GC log alongside the long full-GC record. ## The mitigations, and why they were never a fix Teams running CMS reached for a familiar set of levers: - **Start the cycle earlier** so it finishes before the old generation fills — done by lowering the initiating occupancy threshold and pinning it, at the cost of more cycles and more CPU spent collecting. - **Provision more headroom.** A larger heap buys time for cycles to complete and for fragmentation to matter less. - **Force compaction at full GCs and compact periodically**, using the collector's options to compact before a full collection or every so many full collections — which trades a scheduled long pause for an unscheduled one. - **Reduce large-object churn** in the application, since big contiguous promotions are what fragmentation punishes hardest. Every one of these lowers the probability of the failure. None removes it, because the failure is inherent to a non-compacting old generation with a live-object population that keeps changing shape. That is the structural argument for the collectors that followed: a region-based collector reclaims and compacts by *evacuating* regions — copying the live objects out of a region and freeing it whole — so contiguity is restored as a normal part of every collection, not as an emergency fallback. ## The lesson that outlives CMS CMS is removed, but the pattern generalises: when a collector's steady state and its fallback have wildly different pause characteristics, capacity planning must be driven by the fallback, not the average. "Low pause" is only meaningful if the worst case is also bounded. Judging a collector by its p50 pause while its p99.99 is a full compaction is precisely the mistake CMS taught the industry to stop making.

  • Why can a promotion fail when the old generation reports plenty of free memory?
    Because the free memory is not contiguous. A non-compacting collector returns dead objects' space to free lists, so free space ends up scattered between live objects. A promoted object needs a single run of memory at least its size, so if the largest free chunk is smaller than the object, promotion fails despite a large total free figure. This is why the largest-free-chunk metric matters more than occupancy under CMS.
  • Why did lowering the initiating occupancy threshold reduce concurrent-mode failures but not eliminate them?
    Starting the cycle earlier gives concurrent marking more time to finish before allocation fills the old generation, so it directly addresses the losing-the-race path. It does nothing about fragmentation, which can deny a large contiguous promotion at any occupancy level, and it costs CPU because cycles run more often. The failure mode is inherent to never compacting, so tuning changes only its probability.
  • How was the region-based collector that replaced CMS structurally different in this respect?
    It divides the heap into regions and reclaims by evacuation: live objects in a chosen region are copied into another region and the source region is freed whole. That makes compaction part of ordinary collection rather than an emergency fallback, so contiguous space is continually restored and free memory means allocatable memory. It still has a full-GC fallback, but the steady state does not accumulate fragmentation the way an in-place sweeper does.

saying these in an interview costs you the question

  • Believing CMS compacted the old generation as part of its concurrent sweep
  • Treating free bytes in the old generation as equivalent to allocatable space
  • Saying a concurrent-mode failure just means the cycle restarts, missing the stop-the-world compacting full GC
  • Assuming the CMS fallback full collection was parallel and therefore fast
  • Claiming the right initiating-occupancy setting removes the failure mode entirely

context

open as a page

HotSpot's Concurrent Mark-Sweep collector (CMS) performed most of its old-generation marking while application threads kept running. Walk through the phases of one of its old-generation cycles and say which of them stopped the world.

level: middleimportance: should knowfreq 40%

basics

~20 s

Initial mark (stop-the-world, marks roots), concurrent mark, concurrent preclean and abortable preclean, final remark (stop-the-world, catches what mutation changed), concurrent sweep, concurrent reset. Only initial mark and remark paused the application; young collections were separately stop-the-world. Sweeping reclaimed space in place without compacting.

open as a page

HotSpot's Concurrent Mark-Sweep collector had to begin its concurrent old-generation cycle well before the old generation was actually full. Why could it not simply wait until the space ran out, and what were the costs of starting too early or too late?

level: seniorimportance: should knowfreq 32%

basics

~20 s

Its marking and sweeping ran concurrently with the application, which keeps allocating and promoting the whole time. If the cycle starts when the old generation is already full, there is no space to serve those allocations before it finishes, so the JVM falls back to a long stop-the-world full collection. Starting too early wastes CPU on extra cycles.

open as a page

HotSpot's Concurrent Mark-Sweep collector was deprecated in JDK 9 and removed entirely in JDK 14. If you owned a latency-sensitive service still pinned to an old JDK because of it, how would you reason about the migration and what would you expect to change operationally?

level: principalimportance: should knowfreq 34%

basics

~20 s

Treat it as unavoidable: staying pinned costs security updates and language features for a collector nobody maintains. Move to the region-based collector (G1) as the default — it compacts as part of normal collection, removing the fragmentation failure mode. Drop CMS-specific flags rather than translating them, set a pause-time goal and heap size, then validate against production-shaped load.

open as a page