skip to content

When would you choose HotSpot's throughput-oriented Parallel collector (-XX:+UseParallelGC) over G1, and what do you give up by doing so?

level: middleimportance: must knowfreq 60%

answer

  1. Parallel = all stop-the-world, no barriers, best throughput
  2. G1 = concurrent mark + incremental evacuation, pause target
  3. G1 costs: write barriers, remembered sets, concurrent CPU, headroom
  4. Parallel full GC scales with live set → seconds on big heaps
  5. Batch/ETL → Parallel; request SLO → G1

basics

~20 s

Choose Parallel when total work per unit time matters and pauses do not — batch jobs, ETL, offline analytics, benchmarks. It usually delivers a few percent more throughput because it does no concurrent work and no write barriers. You give up pause control: its collections stop everything for a time that scales with the heap.

solid answer

~60 s

**Parallel** collects with multiple threads but entirely stop-the-world, in a classic copying young generation plus a mark-sweep-compact old generation. It maintains no concurrent marking, does very little barrier work, and never competes with application threads for CPU. That makes it the throughput champion on the same hardware — typically a few percent, occasionally more, over G1 — and it has the simplest behaviour to reason about. **G1** is region-based and incremental: it marks concurrently, evacuates young regions and batches of garbage-rich old regions in pauses sized to `-XX:MaxGCPauseMillis` (default 200 ms). It pays for that with write barriers on reference stores, remembered-set maintenance, and concurrent threads that consume CPU alongside the application. So the choice is: **Parallel for jobs measured in completed work** — nightly batch, data pipelines, compile farms, benchmark harnesses, anything where a two-second stop costs nothing — and **G1 for anything with a request latency SLO**, because Parallel's full collection scales with the live set and a multi-gigabyte heap means seconds of silence. G1 is the default for good reason; picking Parallel is a deliberate trade of tail latency for a few percent of throughput.

go deeper

for a junior

Know that Parallel is the throughput collector with fully stop-the-world collections and G1 the pause-targeting default, and which kind of application suits each.

for a middle

Explain where G1's overhead actually comes from and why Parallel's full-collection pause scales with the live set.

for a senior

Frame it as a measured trade against the service's SLO, and run both under the real workload with GC logging before committing.

for a principal

Set the fleet rule — request-serving tiers on a pause-bounded collector, batch tiers free to take the throughput win — and require evidence for exceptions.

## Two different objective functions Garbage collectors optimise for different things, and "which is better" is meaningless without naming the metric. The three that matter are **throughput** (share of CPU time doing application work rather than collection), **latency** (how long, and how often, application threads are stopped), and **footprint** (heap headroom and metadata overhead needed to run well). You can usually have two. ## What Parallel is The Parallel collector is a straightforward generational design executed by many threads, all stop-the-world: - **Young collections** copy survivors between eden and survivor spaces and promote long-lived objects to the old generation. Cost is proportional to surviving data. - **Full collections** mark, sweep and compact the whole old generation, again with all application threads stopped. Cost is proportional to the live set. Because nothing runs concurrently, Parallel needs almost no coordination with application threads: no concurrent marking, no expensive write barrier maintaining remembered sets, no snapshot bookkeeping. Every CPU cycle it spends is spent collecting, and every cycle it does not spend is available to the application. It also has adaptive sizing (`-XX:+UseAdaptiveSizePolicy`, on by default) that tunes generation sizes toward a throughput goal expressed by `-XX:GCTimeRatio`, and it accepts `-XX:MaxGCPauseMillis` as a soft hint, though it has far less machinery to honour one than G1 does. ## What G1 is, and what it costs G1 divides the heap into equal-size regions and treats generations as *sets of regions* rather than contiguous spaces. It marks the old generation concurrently, then reclaims it incrementally: each pause evacuates the young regions plus a batch of the most garbage-rich old regions, with the batch sized so the pause fits the target. That incrementality is what decouples pause duration from heap size, and it is why a 32 GB G1 heap can still keep pauses in the tens of milliseconds while a 32 GB Parallel full collection cannot. The price is real and worth stating in an interview: - **Write barriers.** Every reference store executes extra code to record cross-region references, so G1 can collect a subset of regions without scanning the whole heap. Pure application code therefore runs marginally slower. - **Remembered sets.** The metadata recording those cross-region references costs memory and CPU to maintain and to refine. - **Concurrent threads.** Marking runs while the application runs, taking CPU that would otherwise serve requests. - **Headroom.** Evacuation needs free regions to copy into; run G1 too close to the ceiling and it degrades into evacuation failures and full compactions. Net effect on a like-for-like benchmark: Parallel typically finishes a fixed workload a few percent faster. ## When Parallel is the right answer - **Batch and offline work.** A nightly ETL, a Spark executor, a report generator, a compiler farm, a scientific job. Nobody is waiting on a response; the metric is wall-clock completion of the whole job, and a two-second stop in the middle costs nothing. - **Short-lived processes with modest heaps** where a full collection is quick anyway. - **CPU-constrained environments** where you cannot afford concurrent GC threads competing with application threads — a small container with a hard CPU quota may do better with a collector that does all its work in a stop rather than one that constantly runs marking threads it has no spare cores for. - **Benchmarking**, where deterministic stop-the-world behaviour is easier to reason about than concurrent work overlapping the measurement. ## When Parallel is the wrong answer Any service with a latency SLO. The killer is the full collection: its duration scales with live data, so a service with 8 GB of live objects can stop for multiple seconds. That blows p99 latency, trips health checks and load-balancer timeouts, and causes cascading retries in a service mesh. Young collections alone might be fine; the full collections are not, and Parallel has no mechanism to make old-generation reclamation incremental. ## How to decide in practice Ask what the service is measured on. If the answer is requests per second at a latency percentile, G1 (or lower still). If it is "the job must finish by 6 a.m.", Parallel is a legitimate few-percent win. Then *verify* rather than assume: run the real workload under both, with GC logging on, and compare throughput and pause percentiles over a representative window. The gap between collectors on your workload is frequently smaller — or larger — than the folklore. ## Version note Parallel remains supported and is not deprecated; it was the default on JDK 8 and must be requested explicitly (`-XX:+UseParallelGC`) from JDK 9 onward. CMS, the old low-pause alternative, was removed in JDK 14, so on modern JDKs the low-pause options are G1, ZGC and Shenandoah.

  • Roughly how much throughput does G1 give up compared with Parallel, and where does it go?
    On like-for-like workloads it is commonly in the low single-digit percent, though it varies widely with allocation and mutation rates. It goes into write barriers executed on reference stores, maintaining and refining remembered sets, and CPU consumed by concurrent marking threads that would otherwise run application code. Workloads that mutate references heavily pay more; mostly-read workloads pay less.
  • A batch job on Parallel with a 24 GB heap occasionally stops for six seconds. Is that a problem?
    Only if something outside the job cares. For a pure batch process with no callers, a six-second full collection is simply part of the job's wall-clock time and may still be the fastest overall option. It becomes a problem the moment a health check, a lease renewal, a coordination heartbeat or a downstream timeout is involved, at which point pause-bounded collection matters more than the few percent of throughput.

saying these in an interview costs you the question

  • Calling G1 strictly better than Parallel; G1 trades throughput for pause control.
  • Believing Parallel is deprecated or removed — that was CMS, in JDK 14.
  • Ignoring that Parallel's full-collection pause scales with the live set, and recommending it for a latency-sensitive service.
  • Overlooking G1's costs — write barriers, remembered sets, concurrent CPU, required headroom.
  • Choosing on folklore rather than measuring both under the real workload.

context