skip to content

What does adopting a concurrent-compaction garbage collector such as Shenandoah actually cost in CPU and memory, and how would you validate that a given service genuinely benefits from it?

level: principalimportance: nice to knowfreq 21%

answer

  1. Buying latency; paying in throughput + CPU + headroom
  2. Barrier tax on reference loads, always on
  3. Precondition: spare cores and free heap
  4. Prove the tail is GC-driven before switching
  5. Accept on p99.9 + zero full GC + flat live set + CPU/RSS budget

basics

~20 s

It costs throughput (barriers on reference loads run on every request path), CPU (collector threads run alongside the application), and memory headroom (it must have free regions to evacuate into). Validate with real traffic: compare caller-side p99.9, CPU-seconds per request, RSS, and the absence of degenerated or full collections.

solid answer

~1 min

**The costs are real and three-fold:** - **Throughput**: the load reference barrier executes on reference loads throughout application code. The tax is workload dependent — pointer-chasing code pays most — and typically lands in the single-digit to low-double-digit percent range. - **CPU**: concurrent marking, evacuation and reference updating run on collector threads while requests are served. Without spare cores that time comes straight out of the application, and low-pause behaviour degrades toward the failure modes. - **Memory**: it compacts by copying, so it needs free regions to evacuate into plus slack for allocation during the cycle. It cannot be run at near-full occupancy. **Validation must be empirical and end-to-end:** 1. Measure the baseline first: caller-side p99/p99.9, CPU-seconds per request, RSS, and the existing pause distribution. 2. Confirm GC pauses actually dominate the tail — often they do not, and the latency comes from locks, I/O or a downstream dependency. 3. Run a canary with production traffic, not a benchmark, for long enough to see peak load. 4. Accept only if tail latency improves materially **and** there are zero full GCs, rare-to-absent degenerated GCs, stable live set, and an acceptable CPU and RSS delta. If the tail is not GC-driven, the adoption is a pure cost.

code

text · 11 lines
text
-XX:+UseShenandoahGC
-Xmx6g                       # real headroom: do not run near full occupancy
-XX:ConcGCThreads=3          # only with spare cores
-Xlog:gc,gc+ergo:file=gc.log:time,uptime

# accept only if all hold over a peak-load window:
#   caller-side p99.9 materially lower than baseline
#   grep -c 'Pause Full' gc.log            == 0
#   grep -c 'Degenerated' gc.log           ~ 0
#   occupancy after each cycle flat (no rising live set)
#   CPU-seconds/request and container RSS within budget

go deeper

for a junior

Know that low-pause collectors trade throughput and CPU for shorter pauses, and that you must measure rather than assume.

for a middle

Name the three costs — barrier throughput, collector CPU, evacuation headroom — and the need for real-traffic testing.

for a senior

Own the validation method: baseline, correlate outliers with pauses, canary, and accept on a set of criteria including zero full GCs.

for a principal

Set the adoption policy — which SLO tier justifies the tax, the CPU and heap floors that make it viable, and the pinned configuration and alerts that keep the decision valid over time.

## Frame it as a purchase, not an upgrade A low-pause collector does not reduce the amount of garbage-collection work. It **relocates** that work out of stop-the-world pauses and into concurrent execution, and it adds new work — barriers — to make concurrency safe. You are buying tail latency and paying in throughput, CPU and memory. The senior-and-above answer is to state the price precisely and then insist on evidence that the thing being bought is the thing the service actually needs. ## Cost 1 — throughput, via barriers To move objects while the application runs, the JIT inserts a **load reference barrier** wherever the application loads a reference out of the heap: it checks global GC state and, during evacuation, resolves or copies the referent and heals the source field. A snapshot-at-the-beginning write barrier additionally supports concurrent marking. These execute in normal application code, not only during collection. The overhead is therefore proportional to how reference-load-heavy the workload is: graph traversal, linked structures, deep object models and interpreter-like code pay more; numeric and array-oriented code pays little. Published and measured figures generally sit in the single-digit to low-double-digit percentage range, but the honest position in an interview is that the number is workload specific and must be measured on your own service. ## Cost 2 — CPU, via concurrent threads Marking, evacuating and updating references run on collector threads *while the application runs*. On a machine with idle cores, that is a good trade. On a CPU-saturated or quota-limited container it is not: collector threads and application threads contend, effective throughput drops, and — worse — the collector may not finish a cycle before the heap fills, producing degenerated or full stop-the-world collections. A collector chosen for latency then delivers the worst pauses in the system. This gives a hard adoption precondition: **spare CPU headroom**. A service pinned to one core is not a candidate; a service already running at 90% CPU is not a candidate until it is given room. ## Cost 3 — memory headroom Compaction by copying requires somewhere to copy to. Live objects in the collection set must land in free regions, and the application keeps allocating throughout the cycle, so the collector needs free space for both. Racing evacuations also waste the losing copies until regions are recycled. Practically, the collector cannot be operated at the near-full occupancy a stop-the-world compactor tolerates; you must budget real free heap, and the container must have the memory to back it plus off-heap collector metadata. ## Validating a benefit — the method **Establish the baseline before changing anything.** Record caller-side latency percentiles (p50, p99, p99.9), CPU-seconds per request, container RSS, and the current pause distribution from GC logs. Measuring server-side only is a common mistake: under CPU quotas, throttling extends the stall a client sees beyond the logged pause. **Prove the tail is GC-driven.** Correlate latency outliers with pause timestamps. In a large fraction of investigations the p99.9 is dominated by lock contention, connection-pool waits, downstream calls or thread-pool queueing, and a collector change will not move it. This step prevents the most expensive kind of wrong answer — adopting a throughput tax that buys nothing. **Canary with production traffic.** Synthetic benchmarks misrepresent allocation profiles and object lifetimes badly. Route a slice of real traffic to instances running the candidate configuration and let them experience peak load and a full deploy cycle. **Read the acceptance criteria as a set, not a single metric:** - Tail latency (p99.9 measured at the caller) improves materially, not marginally. - **Zero** `Pause Full` events; degenerated collections rare or absent. - Live set flat across cycles — a rising one indicates a leak, which must be fixed independently. - CPU-seconds per request within an agreed budget, understanding it will rise. - RSS within the container limit with margin, counting off-heap collector metadata. **Then pin the configuration.** Record the collector, heap size and headroom in the deployment so a later capacity change does not silently invalidate the validation. ## Additional considerations worth raising - **Root-set size drives the remaining pauses.** Thousands of threads with deep stacks lengthen the init/final mark pauses. Reducing thread count is a genuine tuning lever, unlike shrinking the heap. - **Generational mode.** A generational Shenandoah mode (`-XX:ShenandoahGCMode=generational`, JEP 404) collects young objects far more cheaply and materially reduces CPU cost on typical workloads; if the JDK in use has it, evaluate it as part of the adoption rather than after. - **Availability.** Not every JDK build ships Shenandoah; confirm the target runtime supports it before designing around it. - **Portfolio discipline.** At fleet scale, define which SLO tier justifies a low-pause collector and the CPU/memory floor that makes it viable, rather than letting each team choose by reputation. ## The judgment being tested The interviewer wants to hear that you can name the price, identify the precondition (spare CPU and heap headroom), insist on proving the tail is GC-driven before paying it, and validate with caller-side measurements on real traffic. "It has shorter pauses so we should use it" is the answer that fails.

  • A team reports that switching to a low-pause collector made their p99.9 worse. What is the most likely explanation?
    The container had no spare CPU, so collector threads competed with application threads and the concurrent cycle could not finish before the heap filled. That produces degenerated or full stop-the-world collections whose pauses are far longer than the collector's normal ones. A secondary possibility is that the tail was never GC-driven, and the added barrier overhead simply slowed everything down.
  • How would you decide the heap headroom to give a service using this collector?
    Empirically, from allocation rate and cycle duration: the free space must cover what the application allocates during a full concurrent cycle, plus the destination space for evacuated live objects, plus margin for spikes. I would start generously, observe the trigger reasons and whether any degeneration occurs at peak, and tighten only while degenerated collections stay at zero.

saying these in an interview costs you the question

  • Presenting a low-pause collector as strictly better, with no throughput or CPU cost
  • Switching collectors without first proving the latency tail is caused by GC pauses
  • Validating on a synthetic benchmark rather than production traffic
  • Running it at near-full heap occupancy because 'the old collector coped'
  • Measuring pause length only in the GC log and never at the caller

context