You own a fleet of JVM services with mixed workloads — request-serving APIs, streaming consumers, nightly batch jobs. How do you decide which garbage collector each should run, and how do you validate the decision?
answer
- Throughput vs latency vs footprint — name the governing metric first
- G1 = default; others need an argument
- Exhaust cheaper levers: heap headroom, Full GC causes, allocation reduction
- Validate on the real workload: pause percentiles + throughput + CPU + stalls
- Pin flags explicitly; ergonomics varies with instance shape
basics
~20 sClassify each service by its governing metric — tail latency, throughput, or footprint — then map: latency SLO plus a large heap goes to a concurrent collector, batch to a throughput collector, everything else to the pause-targeting default. Validate by running the real workload under load with GC logging and comparing pause percentiles, throughput and CPU.
solid answer
~60 s**Classify before configuring.** For each service, name the metric it is judged on and the numbers behind it: p99/p99.9 latency budget, required throughput, heap and live-set size, allocation rate, CPU headroom, and cost sensitivity. **Then map, with a default:** - APIs and anything with a latency SLO: start on **G1** (the default) — it is well understood, well tooled, and adequate for most services. - Services where a tight tail SLO collides with a large heap, or where a long stop triggers failover: **ZGC/Shenandoah**, budgeting CPU and heap headroom. - Batch, ETL, offline compute with no waiting caller: **Parallel** for the throughput win. - Tiny sidecars and one-CPU containers: **Serial**, whose footprint and simplicity fit. **Validate empirically.** Collector choice is not decidable from first principles on someone else's benchmark. Run the real workload at realistic load with GC logging on, over a window that includes peak and any batch overlap, and compare pause percentiles, throughput, CPU utilisation, and failure signals (Full GCs, evacuation failures, allocation stalls). **Then standardise.** Pin the collector and heap explicitly in the base image so ergonomics does not vary by instance shape, and treat exceptions as documented decisions with evidence.
go deeper
Know that the choice follows the workload — batch favours throughput collectors, latency-sensitive services favour pause-bounded ones — and that G1 is the sensible default.
Add the specific mapping and the measurable inputs: heap size, live set, allocation rate, latency budget.
Run the trial properly — real workload, real load, identical logging, compare percentiles and CPU, watch for Full GCs and allocation stalls — and fix sizing and allocation before switching.
Own it as fleet policy and economics: a default with documented exceptions, flags pinned in the platform, GC telemetry shipped and alerted on, and a re-evaluation trigger on JDK upgrades.
## The decision is about the workload, not the collector Every collector is a point on a three-way tradeoff: **throughput** (application CPU share), **latency** (length and frequency of stops), **footprint** (heap headroom and metadata). No collector wins all three, so the question "which collector is best" is unanswerable and the useful question is "which of the three does this service's contract actually price?" ## Step 1: classify each service For every service, write down: - **Governing metric.** Is it judged on p99.9 latency, on requests per second, on job completion time, or on cost per instance? - **The numbers.** The latency budget in milliseconds; the throughput target; the heap size and, importantly, the *live set*, since that is what long collections scale with; the allocation and promotion rates from GC logs. - **Failure semantics.** Does a long stop merely slow a response, or does it trip a health check, expire a lease, lose a partition assignment, and cause a failover? Streaming consumers and clustered stores often care about the second, and that changes the calculus far more than average latency does. - **Resource envelope.** CPU quota and memory limit per instance. A concurrent collector without spare CPU is a downgrade. ## Step 2: map to a default, with named exceptions A workable fleet policy: | Class | Collector | Rationale | |---|---|---| | Request-serving API, moderate heap | G1 | Pause target, mature tooling, good default balance | | Latency-critical service, large heap, or stop-triggers-failover | ZGC / Shenandoah | Pauses bounded by construction, not by heap size | | Batch, ETL, offline compute | Parallel | No waiting caller; a few percent throughput is free money | | Tiny sidecar, one-CPU container | Serial | Smallest footprint, no coordination overhead | | Benchmark / allocation-regression harness | Epsilon | Zero-GC control group | The important discipline is that **G1 is the default and the others require an argument**. Most "we need ZGC" requests turn out to be an undersized heap, humongous allocations, or an application allocating far more than it needs — problems a collector switch masks rather than solves. ## Step 3: exhaust the cheaper levers first Before changing collector, check in this order: is the heap sized for the live set with evacuation headroom; are there Full GCs and what causes them; is anything calling `System.gc()`; are there humongous allocations; is the allocation rate absurd for what the service does; is the container CPU-throttled. Reducing allocation is usually the highest-leverage change available and helps under every collector. ## Step 4: validate on the real workload Published benchmarks describe someone else's object graph. Run yours: - **Realistic load, realistic duration.** Include peak, and include the window where batch overlaps online traffic — that overlap is where fleets actually break. - **GC logging on** for every arm of the trial, with the same flags, so comparisons are like for like. - **Compare the full picture:** pause percentiles (not means), application-level latency percentiles, throughput, CPU utilisation, memory footprint, and failure signals — Full GC count, evacuation failures, allocation stalls. - **Watch for the second-order effects:** a collector that shortens pauses but raises CPU can worsen end-to-end latency under a quota; a collector that raises throughput can raise tail latency. - **Decide against the SLO**, not against the prettiest GC graph. The GC log is an instrument; the service's own latency and error metrics are the verdict. ## Step 5: standardise and keep it observable - **Pin explicitly.** Put the collector and heap flags in the base image or platform config rather than relying on ergonomics, which varies with instance shape and container limits — the same artifact silently picking Serial on a small pod and G1 on a large one is a fleet-consistency defect. - **Ship GC logs** with rotation and wall-clock decorators, retained long enough to analyse an incident after the fact. - **Alert on the signals that mean the choice stopped fitting:** Full GCs under a concurrent collector, allocation stalls, pause percentiles crossing a threshold, GC CPU share climbing. - **Re-examine on JDK upgrades.** Defaults and collector maturity move — the JDK 8 → 9 default change from Parallel to G1, CMS's removal in JDK 14, ZGC becoming generational by default from JDK 23. An upgrade is a legitimate trigger to re-run the comparison. ## Step 6: price it At fleet scale this is an economics decision. A concurrent collector may need more CPU and more heap per instance; multiplied across hundreds of instances that is a real bill, set against the cost of tail-latency breaches, failovers and the engineering time spent tuning around a collector that does not fit. Stating that tradeoff in money and risk — rather than in benchmark percentages — is what distinguishes a principal-level answer.
- A team asks to move to ZGC because their p99 is bad. What do you check before approving?Their GC logs first: pause percentiles, Full GC count and causes, allocation and promotion rates, and how much of the p99 is actually GC rather than downstream calls or lock contention. Then heap headroom and humongous allocations, since both produce G1 pathologies that a switch would merely mask. Only if pauses are genuinely the cause, G1 is already sized and tuned, and the instance has CPU and memory headroom does the switch make sense — validated on the real workload.
- Why pin collector and heap flags rather than relying on JVM ergonomics across a fleet?Ergonomics derives the collector and heap from the CPU and memory the JVM detects, which in containers means the pod's limits. The same image can therefore run G1 on a large node and Serial on a small one, with heap ceilings that shift when limits change — variability that shows up as unexplained latency differences between instances. Pinning makes behaviour a property of the deployment rather than an accident of scheduling.
- How would you handle a mixed service that serves requests during the day and runs heavy batch at night?Prefer to split them into separate processes or tiers so each gets the collector its metric deserves, since one JVM cannot be both throughput-optimal and latency-optimal. If they must share, the latency requirement wins — a pause-bounded collector with enough headroom to absorb the batch allocation burst — and the trial must cover the overlap window, which is exactly where the configuration will be tested in production.
saying these in an interview costs you the question
- Recommending a collector from published benchmarks without trialling the actual workload.
- Switching collectors before checking heap sizing, Full GC causes and allocation behaviour.
- Treating low pause time as the goal instead of the service's own latency and error SLOs.
- Leaving the choice to ergonomics across a heterogeneous fleet and being surprised by inconsistent behaviour.
- Ignoring the CPU and memory budget a concurrent collector needs, then blaming the collector when latency worsens.