skip to content

You are choosing a garbage collector for a JVM workload. Under what conditions would you deliberately pick the throughput-oriented Parallel collector (-XX:+UseParallelGC) instead of the modern default region-based collector, and how would you justify that call with evidence?

level: principalimportance: should knowfreq 38%

answer

  1. Which number is the workload judged on?
  2. Batch / CI / short-lived → Parallel
  3. 1–2 core containers: no background GC threads to steal CPU
  4. Small heap = short compaction; big live set = seconds
  5. Prove it: runtime, GC overhead %, p99.9, CPU-seconds

basics

~20 s

Choose it when the success metric is completion time or cost rather than tail latency, and when CPU is scarce: batch and ETL jobs, CI and build workers, short-lived tasks, and small containers with one or two cores where concurrent GC threads would compete with application threads. Justify with measured GC overhead and end-to-end runtime, not with a preference.

solid answer

~60 s

The decision hinges on which number the business is paying for. **Pick Parallel when the metric is finishing time or cost.** Batch jobs, ETL, report generation, compilers, CI workers, message-drain jobs, short-lived functions. Nobody's request latency is affected by a 700 ms stop in the middle of a 20-minute job — but 3–10% less total CPU spent on GC directly shortens the job and reduces the bill. **Pick it when CPU is scarce.** In a container with one or two cores, a concurrent collector's background threads take cycles away from the application threads. The Parallel collector has no background threads and a near-free write barrier, so the constrained case is exactly where it wins most. **Pick it when the heap is small** (roughly single-digit gigabytes with a modest live set) so a whole-heap compaction is a few hundred milliseconds, not seconds. **Do not pick it** for latency-SLO services, very large heaps, or anything where a whole-heap compaction pause would breach an objective. The justification must be measured: run both collectors on the same workload and compare end-to-end runtime, GC overhead percentage, p99/p99.9 latency and CPU consumed. Choose from the numbers.

go deeper

for a junior

Know the headline: Parallel for batch jobs and small containers where finishing fast matters more than a smooth latency profile.

for a middle

Be able to name the trade explicitly — throughput and low footprint bought with whole-heap pauses — and say which workloads can absorb it.

for a senior

Bring the measurement plan: same workload, both collectors, compare runtime, GC overhead, tail latency and CPU, and read full-collection frequency from the log.

for a principal

Treat it as a fleet-level policy question: a small number of documented collector profiles, each tied to a metric and an explicit revisit condition, rather than per-team flag folklore.

## Reframe the question before answering it The wrong instinct is to compare collectors feature by feature. The right one is to ask *what number does this workload get judged on*, because collectors are just different points on the same three-way trade between throughput, pause time and footprint. You cannot maximise all three; you decide which one you are buying. - **Throughput** — proportion of time running application code. What a batch job is judged on. - **Pause time** — length of an individual stop. What a latency SLO is judged on. - **Footprint and CPU** — what your cloud bill is judged on. The Parallel collector maximises throughput and keeps footprint modest, and pays for both with pause length. ## The cases where Parallel is the right call **1. Completion-time workloads.** Batch pipelines, ETL, nightly reports, data-processing jobs, compilers, CI and build agents, scheduled drains of a queue. The user-visible metric is "the job finished at 04:20 instead of 04:23." A stop-the-world pause inside such a job is invisible; total GC overhead is not. This is the flagship case. **2. CPU-constrained containers.** This one is underappreciated. Concurrent collectors buy latency by spending CPU: background marking threads, refinement work, and more expensive read/write barriers on the application's own path. Give the JVM one or two cores and those background threads are directly stealing from the application. The Parallel collector has none of them, and its only barrier is an unconditional card mark. On a small container it often shows both better throughput *and*, counter-intuitively, competitive pauses, because the heap it is compacting is small. **3. Small heaps with modest live sets.** The Parallel collector's weakness scales with live data. At a 2–4 GB heap with a few hundred megabytes live, a full compaction is a few hundred milliseconds and happens rarely. The weakness has not bitten yet. **4. Short-lived processes.** A JVM that runs for 90 seconds may never need a full collection at all. Optimising for the steady-state pause profile of a collector that will never reach steady state is wasted. **5. Cost-optimised fleets.** If you run thousands of instances of a non-latency-critical workload, a few percent less CPU in GC is a real line item, and lower memory overhead means smaller instances. ## The cases where it is the wrong call - **Any service with a tail-latency objective.** A whole-heap compaction stops every thread at once, so every in-flight request pays simultaneously. One 800 ms stop wrecks p99.9 for that whole window. - **Large heaps.** Beyond roughly 8–16 GB with a substantial live set, full-collection pauses become seconds. Region-based and concurrent-compaction collectors exist precisely because that became intolerable. - **Large, growing live sets.** Caches, in-memory indexes, session state — the pause is proportional to the live data you deliberately keep. - **Workloads with big, churny arrays.** Objects too large for eden go straight to the old generation, so a large-array workload drives full collections even though the data is short-lived. ## How to justify the choice with evidence A principal-level answer never stops at "batch, so Parallel." It specifies the experiment: 1. **Fix the workload.** A representative, repeatable run — real data volume, real concurrency, warmed up. 2. **Run both collectors** with the same heap sizing, changing only the collector flag. For benchmark stability, consider pinning generation sizes so ergonomics does not add variance between runs. 3. **Collect the same four numbers each time:** end-to-end wall-clock runtime; GC overhead (sum of pause durations divided by wall clock, from `-Xlog:gc`); the latency distribution if there is one, at p50/p99/p99.9; and CPU-seconds and peak RSS consumed. 4. **Compare against the objective, not against a preference.** If runtime is 6% shorter and there is no latency SLO, Parallel wins and the maximum observed pause is simply a fact you record. If p99.9 blows past the SLO, no throughput gain rescues it. 5. **Record the decision and its expiry condition.** "Parallel, because the job is batch and the live set is under 500 MB. Revisit if the heap exceeds 8 GB or if this process starts serving synchronous requests." A collector choice that nobody can re-derive later becomes cargo cult. ## The organisational dimension A fleet with two collector profiles — a latency profile for request-serving services and a throughput profile for batch — is a defensible standard. A fleet with nine hand-tuned flag sets copied between teams is not. Whatever you choose, make it a named, documented profile, and keep the number of profiles small enough that people can reason about them. ## The sentence to land "Parallel when the metric is completion time or cost and CPU is scarce; the default region-based collector when the metric is tail latency or the heap is large. Then prove it: same workload, both collectors, compare runtime, GC overhead, p99.9 and CPU — and write down the condition that would make me revisit."

  • A team says 'we picked Parallel because it's fastest.' What would you ask them?
    Fastest at what, and measured how. I would ask which metric their service is judged on, whether they ran the comparison on a representative workload, and what their maximum observed pause was. If it is a request-serving service with a latency objective, the throughput win is irrelevant next to a whole-heap compaction stop; if it is a batch job, they may well be right but should still be able to show the numbers and name the condition that would make them revisit.
  • Does choosing the Parallel collector change how you size the heap?
    Yes. Because the full-collection pause scales with live data, a very large heap is actively harmful under this collector — you are enlarging the worst case rather than avoiding it. I would size the heap to comfortably hold the live set with room for allocation churn, rather than reflexively giving it everything the container has, and I would watch full-collection frequency and duration as the primary signal that the heap or the collector needs to change.

saying these in an interview costs you the question

  • Choosing a collector by reputation or defaults instead of by which metric the workload is judged on
  • Assuming the newest collector is always the right one, including on one-core containers and batch jobs
  • Giving a throughput collector an enormous heap and being surprised by multi-second full collections
  • Claiming Parallel is unsuitable for production; it is unsuitable for latency SLOs, which is not the same thing
  • Comparing collectors without holding heap size, workload and warm-up constant, then trusting the numbers

context