skip to content

Parallel Collector

The throughput collector: many GC threads, still fully stop-the-world, minimizing total collection time rather than individual pause length. Interviewers use it to probe the throughput-versus-latency trade-off, since it is often the better answer for batch jobs even though it is no longer the default.

on this pageshow

questions

5

HotSpot's Parallel collector (enabled with -XX:+UseParallelGC) is described as a "throughput collector." What does that mean concretely about how it does its work, and what does it give up in exchange?

level: middleimportance: must knowfreq 55%

answer

  1. Throughput = % of time not in GC
  2. All phases STW, many GC threads, work stealing
  3. No concurrent threads = no CPU stolen between pauses
  4. Only a card-mark barrier, no SATB
  5. Full GC compacts whole heap in one stop

basics

~20 s

Every phase is stop-the-world but executed by many GC threads at once, minimising total time spent collecting and maximising the share of time the application runs. The cost: full stops whose length grows with live-set size, and no collection work overlapping the application.

solid answer

~60 s

Throughput here means the fraction of process time spent running application code rather than collecting. The Parallel collector optimises exactly that: it does **no** work concurrently with the application. It lets the mutator run at full speed until a collection is needed, then stops every application thread and throws all its GC threads at the heap at once — parallel copying in the young generation, parallel mark-summary-compact for a full collection — with work stealing to keep the threads busy. Because nothing is concurrent, the mutator pays almost no ongoing tax: no concurrent marking threads competing for CPU, no expensive snapshot-at-the-beginning barriers, just a cheap card-mark on reference stores. That is why it usually posts the best raw throughput of any HotSpot collector on a given machine. What it gives up is pause behaviour. A full collection compacts the whole heap in one stop, so the pause scales with the amount of live data — multi-second pauses on multi-GB heaps are normal. There is no incremental or partial old-generation collection to spread that cost out.

code

text · 3 lines
text
[2.145s][info][gc] GC(3) Pause Young (Allocation Failure) 512M->96M(1024M) 18.311ms
[5.004s][info][gc] GC(7) Pause Young (Allocation Failure) 608M->188M(1024M) 24.870ms
[7.902s][info][gc] GC(9) Pause Full (Ergonomics) 880M->210M(1024M) 742.115ms

go deeper

for a junior

Know the one-liner: many GC threads, all work done while the application is stopped, tuned to spend the least total time collecting.

for a middle

Explain throughput versus pause time as two different measurements, and name the price — a whole-heap compaction in a single stop whose length grows with live data.

for a senior

Connect it to workload shape: quote GC overhead from the log, and be able to argue which services can absorb a multi-hundred-millisecond stop and which cannot.

for a principal

Frame it as a resource-allocation decision — concurrent collectors buy latency with CPU and footprint; on constrained or batch workloads that purchase is a net loss, and Parallel is the deliberate choice.

## Two different clocks GC discussions constantly confuse two measurements. **Throughput** is the fraction of total process time spent executing application code rather than collecting. If a process runs for 100 seconds and 2 of them are inside the collector, throughput is 98%. HotSpot even exposes this as a goal: `-XX:GCTimeRatio=N` asks for at most `1/(1+N)` of time in GC, and its default of 99 is a request for 1%. **Pause time** is the length of one individual stop. A collector can have excellent throughput and terrible pauses (one 4-second stop per hour) or mediocre throughput and lovely pauses (constant small stops plus background threads burning CPU). The Parallel collector is built to win the first number and is explicitly willing to lose the second. ## How it wins throughput **Nothing runs concurrently.** There are no background marking threads, no refinement threads, no concurrent compaction. Between collections, 100% of the CPU belongs to the application. **When it does collect, it collects with everybody.** All application threads are brought to a safepoint and stopped; then N GC threads work the heap in parallel, using per-thread local allocation buffers for copying and work-stealing queues so no thread idles while another still has a deep object graph to scan. That converts what a single-threaded collector would do in T seconds into roughly T/N seconds of wall clock, at the same total CPU cost. **The mutator tax is tiny.** Concurrent collectors need the application to cooperate: snapshot-at-the-beginning or incremental-update barriers on reference writes, load barriers on reads, extra bookkeeping so a moving collector can relocate objects while threads are running. The Parallel collector needs none of that. Its only ongoing cost is a cheap unconditional card mark on reference stores, so that a young collection can find old-to-young references without scanning the whole old generation. Allocation itself is a pointer bump in a thread-local buffer. **The heap stays simple.** Classic contiguous young generation (eden plus two survivor spaces) and a contiguous old generation. No region metadata, no remembered-set-per-region maintenance, no mixed-collection heuristics. Less machinery means less overhead and very predictable, cache-friendly copying. ## What it gives up **Full collections are one big stop.** The old generation is collected by a parallel mark-summary-compact pass over the whole heap in a single stop-the-world event. Compaction physically slides live objects, so the pause is roughly proportional to the amount of *live* data (plus a marking cost proportional to the live object graph). On a 1 GB heap that might be a few hundred milliseconds; on an 8–16 GB heap with a large live set, seconds. There is no way to do "part of" the old generation. **Pause goals are best-effort only.** You can set `-XX:MaxGCPauseMillis`, and the ergonomics will shrink generations to try to hit it, but there is no mechanism that can interrupt a full compaction once it starts. Treat the flag as a hint, never a guarantee. **No concurrent class unloading or reference processing.** Those happen inside the stop as well. **Latency-sensitive services suffer.** If your service has a p99 latency SLO measured in tens of milliseconds, a periodic multi-hundred-millisecond stop shows up directly in the tail, and every in-flight request pays for it simultaneously. ## Reading it in a log With unified logging (`-Xlog:gc`), Parallel collections appear as `Pause Young (Allocation Failure)` and `Pause Full (Ergonomics)`. There are no `Concurrent ...` lines at all — their absence is itself the signature of this collector. Older-format logs name the spaces `PSYoungGen` and `ParOldGen` (`PS` for "parallel scavenge"). ## Where the trade pays off Batch jobs, ETL, data processing, compilers, CI workloads, benchmarks, short-lived jobs, and any process whose success metric is "finish sooner" rather than "answer this request in under X ms." It also does well in small CPU-constrained containers, where a concurrent collector's background threads would be competing with the very application threads you are trying to keep fed. ## The sentence to say in an interview "Parallel does all its work stop-the-world with many threads. That maximises the proportion of time the app runs, because nothing is stolen from the app between collections and the barrier overhead is minimal — but a full collection compacts the entire heap in one stop, so pause length scales with live data."

  • How would you actually measure whether the Parallel collector is giving you the throughput you want?
    Sum the pause durations from the GC log over a window and divide by the wall-clock length of that window; that is your GC overhead, and 1 minus it is throughput. `-Xlog:gc` plus a log parser, or the GC time counters exposed by the JVM's memory-management beans and jstat, both give you this. Compare it against the implicit goal of `-XX:GCTimeRatio` (default 99, i.e. 1% overhead) and against the job's total runtime, which is the number that actually matters for a batch workload.
  • Does the Parallel collector use write barriers at all, given that nothing runs concurrently?
    Yes, but only a card-marking barrier. Every reference store into an object dirties the corresponding card so that a young collection can find old-to-young references without scanning the whole old generation. It is an unconditional store to a byte array — a couple of instructions with no branch. That is far cheaper than the snapshot-at-the-beginning or load barriers a concurrent collector needs, which is a large part of why Parallel's raw throughput is hard to beat.
  • If Parallel has the best throughput, why isn't it still the default collector?
    Because the common case shifted. Most JVMs today run long-lived services with latency objectives and heaps of several gigabytes, where a whole-heap compaction pause is unacceptable even if it is rare. The default was changed to the region-based collector, which trades a few percent of throughput for pauses that scale with a bounded region set instead of with the whole live heap. Parallel remains the better pick where finishing time, not tail latency, is the metric.

A road crew that closes the whole motorway at 3 a.m. and repaves it with fifty workers at once. Total disruption is minimal and the work finishes fastest, but while it happens nobody drives at all — unlike a crew that coneS off one lane and works all week alongside traffic.

saying these in an interview costs you the question

  • Saying the Parallel collector runs concurrently with the application — every one of its phases is stop-the-world
  • Thinking 'parallel' refers to the application being multi-threaded rather than to the GC threads
  • Claiming -XX:MaxGCPauseMillis guarantees a maximum pause; it is a best-effort ergonomic hint
  • Assuming its full collection is single-threaded — modern Parallel does the old-generation mark-compact with all GC threads
  • Dismissing it as obsolete; on CPU-constrained containers and batch jobs it still posts the best throughput

context

open as a page

Which JVM flag selects HotSpot's Parallel garbage collector, how many collector threads does it use by default, and was it ever the JVM's own default choice?

level: juniorimportance: should knowfreq 32%

basics

~20 s

-XX:+UseParallelGC selects it. Thread count defaults to the number of available processors up to 8, then grows more slowly above that, and is overridable with -XX:ParallelGCThreads. It was the default on server-class machines through JDK 8; from JDK 9 the region-based G1 collector became the default.

open as a page

For HotSpot's Parallel collector (-XX:+UseParallelGC), describe what happens during a young collection versus a full collection — which algorithm each phase uses, and why the resulting heap has no fragmentation.

level: middleimportance: should knowfreq 40%

basics

~20 s

A young collection is a parallel copying collection: live objects in eden and the active survivor space are copied to the other survivor space or promoted to the old generation. A full collection is a parallel mark, summary and sliding-compaction pass over the entire heap. Both move objects, so free space stays contiguous.

open as a page

HotSpot's Parallel collector runs with adaptive size policy enabled by default. What exactly is it adapting, in what order of priority, and when would you deliberately disable it with -XX:-UseAdaptiveSizePolicy?

level: seniorimportance: should knowfreq 32%

basics

~20 s

It continuously resizes eden, the survivor spaces and the old generation, and adjusts the tenuring threshold, using measured collection statistics. Goals in priority order: the pause-time goal, then the throughput goal, then minimum footprint. Disable it when you need stable, explicitly-set generation sizes for reproducible behaviour.

open as a page

You are choosing a garbage collector for a JVM workload. Under what conditions would you deliberately pick the throughput-oriented Parallel collector (-XX:+UseParallelGC) instead of the modern default region-based collector, and how would you justify that call with evidence?

level: principalimportance: should knowfreq 38%

basics

~20 s

Choose it when the success metric is completion time or cost rather than tail latency, and when CPU is scarce: batch and ETL jobs, CI and build workers, short-lived tasks, and small containers with one or two cores where concurrent GC threads would compete with application threads. Justify with measured GC overhead and end-to-end runtime, not with a preference.

open as a page