HotSpot's Concurrent Mark-Sweep collector was deprecated in JDK 9 and removed entirely in JDK 14. If you owned a latency-sensitive service still pinned to an old JDK because of it, how would you reason about the migration and what would you expect to change operationally?
answer
- deprecated JDK 9, removed JDK 14 — maintenance cost, no owner
- G1 evacuates regions → compaction is normal, no fragmentation cliff
- delete CMS flags; start from heap size + pause goal
- compare p99/p99.9/max, not averages
- unified -Xlog:gc* breaks old dashboards and parsers
basics
~20 sTreat it as unavoidable: staying pinned costs security updates and language features for a collector nobody maintains. Move to the region-based collector (G1) as the default — it compacts as part of normal collection, removing the fragmentation failure mode. Drop CMS-specific flags rather than translating them, set a pause-time goal and heap size, then validate against production-shaped load.
solid answer
~60 s**Frame it as risk, not preference.** Pinning a JDK to keep a removed collector means forgoing security updates and every runtime improvement since; the collector itself had no maintainers, which is why it was removed. **Target selection.** G1, the region-based collector, is the default and the intended successor. It keeps the concurrent-marking structure but reclaims by *evacuating* regions, so compaction happens during normal collection — which retires CMS's defining failure mode, the fragmentation-driven fallback to a long full GC. If measured pause requirements are genuinely in the low-millisecond range on a large heap, the concurrent-compaction collectors (ZGC, Shenandoah) are the next step, at some throughput and footprint cost. **Migration mechanics.** Delete the CMS flag set rather than translating it — there is no per-flag mapping, and the old numbers encode assumptions that no longer hold. Start from heap size plus a pause-time goal, and add only what measurement justifies. **Expect a different shape.** Typically slightly higher median pauses than CMS's best case, but a far better tail, no fragmentation cliff, and different log records — so recalibrate alerts, dashboards, and any log parsing against the unified GC logging format.
code
text · 10 lines# before (JDK 8, CMS)
-XX:+UseConcMarkSweepGC -XX:+UseParNewGC
-XX:CMSInitiatingOccupancyFraction=70 -XX:+UseCMSInitiatingOccupancyOnly
-XX:+CMSParallelRemarkEnabled -XX:+UseCMSCompactAtFullCollection
# after (modern JDK, region-based collector)
-XX:+UseG1GC
-Xms8g -Xmx8g
-XX:MaxGCPauseMillis=100
-Xlog:gc*,gc+heap=info:file=gc.log:time,uptime,level,tags:filecount=10,filesize=20mgo deeper
Know that CMS is gone and that the region-based collector is the default replacement; do not attempt the tuning judgment.
Explain why compaction during normal collection removes the fragmentation failure mode, and that migration means dropping the old flags and setting heap size plus a pause goal.
Own the validation: production-shaped load, full pause distributions, GC CPU, re-tooled unified-logging dashboards, and a rollback plan.
Lead with risk framing — the cost of staying pinned versus days of tuning — and articulate collector choice as a negotiation between pause targets, throughput, headroom, and operational simplicity, planned against the worst case rather than the median.
## Why the removal happened CMS was deprecated in JDK 9 and removed in JDK 14. The stated reasoning was maintenance economics rather than any single defect: CMS was a large, intricate, separately-maintained code path that complicated changes to the shared GC infrastructure, and nobody stepped forward to maintain it once G1 became the default. The technical case was already strong — a collector that never compacts carries a fragmentation-driven fallback to a long stop-the-world compacting full collection, which no amount of tuning eliminates. ## The decision to make The real question a service owner faces is not "which collector is best" but "what am I paying to stay where I am". Staying on a JDK old enough to have CMS means: - no security patches beyond that release's support window; - no runtime improvements — later JDKs bring compiler, startup, and footprint gains that often outweigh the collector difference; - an unbounded worst-case pause that remains one production traffic spike away. Against that, the migration cost is a tuning-and-validation exercise measured in days. Framed that way, the decision is straightforward; the engineering work is in doing it without a latency regression. ## Choosing the successor **G1 (region-based, the default)** is the right first target for nearly everyone. It preserves the parts of CMS that mattered — concurrent marking of the old generation, short root-scanning and fix-up pauses — and changes the reclamation strategy: the heap is divided into equal-sized regions, and collection *evacuates* live objects out of chosen regions into others, freeing each source region whole. Because evacuation is a copy, compaction is intrinsic to ordinary collection. Free memory therefore means allocatable memory, and the CMS fragmentation cliff simply does not exist in the steady state. G1 also takes a pause-time goal and adapts how many regions it collects per pause to meet it, which replaces much of the hand-tuning CMS demanded. **The concurrent-compaction collectors (ZGC, Shenandoah)** go further: they perform relocation concurrently using barriers, so pause times become largely independent of heap size. They are the answer when the requirement is genuinely single-digit-millisecond pauses on a large heap, and they cost some throughput and some footprint for it. Choose them from measured requirements, not aspiration. **The parallel throughput collector** remains the right answer for batch or throughput-bound workloads where pause time is not the constraint — worth stating explicitly, because a service that migrated off CMS may not have needed low-pause behaviour in the first place. ## How to run the migration **Delete the old flags.** There is no per-flag translation from a CMS configuration; the flags encode assumptions about in-place sweeping, initiating occupancy, and permanent generation that do not carry over. Carrying them forward at best produces warnings and at worst pins the new collector into a bad configuration. Start clean. **Start with a minimal configuration**: maximum heap size, and a pause-time goal. Resist adding knobs until measurement demands them. **Validate under production-shaped load.** GC behaviour is a function of allocation rate, promotion rate, object size distribution, and live-set size — a synthetic load test that misses any of these will mislead. Replay real traffic if you can, and run long enough to see steady state, not just warm-up. **Watch the right metrics.** Compare full pause distributions, not averages: p50, p99, p99.9, and maximum. Also compare CPU spent in GC, heap occupancy after collection, and allocation rate. The honest expected outcome is that G1's median pause may be a little higher than CMS at its best, while the tail is dramatically better because the fallback cliff is gone. **Re-tool the observability.** Modern JDKs use unified logging (`-Xlog:gc*`), whose records look nothing like the old CMS log lines. Any dashboard, alert, or log-parsing script keyed on CMS phase names must be rewritten, and pause thresholds recalibrated against the new distribution. This is routinely underestimated and is the most common source of post-migration noise. **Know the new failure modes.** G1 is not free of sharp edges: allocations of very large objects occupy contiguous *humongous* regions and can drive collection frequency, and G1 still has a stop-the-world full-GC fallback if it cannot keep up. The point is not that the fallback vanished, but that reaching it now requires genuine over-subscription rather than the slow accumulation of fragmentation. ## The judgment to articulate What an interviewer is listening for at this level is the reasoning frame: that a collector choice is a negotiation between pause-time targets, throughput, heap headroom, and operational simplicity; that the worst case, not the median, should drive capacity planning for a latency-sensitive service; and that migrating a runtime is a measurement exercise with a rollback plan, not a flag swap. The specific answer — move to the region-based collector, validate, escalate to a concurrent-compaction collector only if measurement shows you must — falls out of that frame.
- Why does moving to a region-based evacuating collector remove CMS's worst failure mode?CMS swept in place and never moved live objects, so old-generation free space fragmented until a large promotion could not find a contiguous chunk, forcing a stop-the-world compacting full collection. A region-based collector reclaims by copying live objects out of selected regions and freeing each region whole, so compaction happens on every ordinary collection. Free memory stays contiguous at region granularity, and the fragmentation cliff disappears from the steady state.
- A team migrates off CMS and reports that median pause time got slightly worse. How would you respond?Ask for the full distribution rather than the median. The relevant comparison is the tail: CMS's median excluded the multi-second compacting full collections it eventually hits, so a slightly higher p50 alongside a much better p99.9 and maximum is a win for a latency-sensitive service. If the tail did not improve, then investigate real causes — heap sizing, humongous allocations, or an unrealistic pause-time goal.
- When is moving to a concurrent-compaction collector such as ZGC or Shenandoah the right call instead?When measurement shows the region-based collector cannot meet a genuine pause requirement, typically single-digit milliseconds on a heap large enough that evacuation pauses scale badly. Those collectors relocate objects concurrently using barriers, so pause time becomes largely independent of heap size. The price is throughput overhead from the barriers and additional footprint, so the requirement should be measured and stated, not assumed.
saying these in an interview costs you the question
- Proposing to stay pinned on an old JDK indefinitely to keep CMS
- Translating CMS flags one-for-one into the new collector's flags
- Judging the migration on median pause time while ignoring the tail
- Assuming the new collector has no stop-the-world full-GC fallback at all
- Choosing a low-latency concurrent-compaction collector without a measured pause requirement, ignoring its throughput and footprint cost