HotSpot's Parallel collector (enabled with -XX:+UseParallelGC) is described as a "throughput collector." What does that mean concretely about how it does its work, and what does it give up in exchange?
answer
- Throughput = % of time not in GC
- All phases STW, many GC threads, work stealing
- No concurrent threads = no CPU stolen between pauses
- Only a card-mark barrier, no SATB
- Full GC compacts whole heap in one stop
basics
~20 sEvery phase is stop-the-world but executed by many GC threads at once, minimising total time spent collecting and maximising the share of time the application runs. The cost: full stops whose length grows with live-set size, and no collection work overlapping the application.
solid answer
~60 sThroughput here means the fraction of process time spent running application code rather than collecting. The Parallel collector optimises exactly that: it does **no** work concurrently with the application. It lets the mutator run at full speed until a collection is needed, then stops every application thread and throws all its GC threads at the heap at once — parallel copying in the young generation, parallel mark-summary-compact for a full collection — with work stealing to keep the threads busy. Because nothing is concurrent, the mutator pays almost no ongoing tax: no concurrent marking threads competing for CPU, no expensive snapshot-at-the-beginning barriers, just a cheap card-mark on reference stores. That is why it usually posts the best raw throughput of any HotSpot collector on a given machine. What it gives up is pause behaviour. A full collection compacts the whole heap in one stop, so the pause scales with the amount of live data — multi-second pauses on multi-GB heaps are normal. There is no incremental or partial old-generation collection to spread that cost out.
code
text · 3 lines[2.145s][info][gc] GC(3) Pause Young (Allocation Failure) 512M->96M(1024M) 18.311ms
[5.004s][info][gc] GC(7) Pause Young (Allocation Failure) 608M->188M(1024M) 24.870ms
[7.902s][info][gc] GC(9) Pause Full (Ergonomics) 880M->210M(1024M) 742.115msgo deeper
Know the one-liner: many GC threads, all work done while the application is stopped, tuned to spend the least total time collecting.
Explain throughput versus pause time as two different measurements, and name the price — a whole-heap compaction in a single stop whose length grows with live data.
Connect it to workload shape: quote GC overhead from the log, and be able to argue which services can absorb a multi-hundred-millisecond stop and which cannot.
Frame it as a resource-allocation decision — concurrent collectors buy latency with CPU and footprint; on constrained or batch workloads that purchase is a net loss, and Parallel is the deliberate choice.
## Two different clocks GC discussions constantly confuse two measurements. **Throughput** is the fraction of total process time spent executing application code rather than collecting. If a process runs for 100 seconds and 2 of them are inside the collector, throughput is 98%. HotSpot even exposes this as a goal: `-XX:GCTimeRatio=N` asks for at most `1/(1+N)` of time in GC, and its default of 99 is a request for 1%. **Pause time** is the length of one individual stop. A collector can have excellent throughput and terrible pauses (one 4-second stop per hour) or mediocre throughput and lovely pauses (constant small stops plus background threads burning CPU). The Parallel collector is built to win the first number and is explicitly willing to lose the second. ## How it wins throughput **Nothing runs concurrently.** There are no background marking threads, no refinement threads, no concurrent compaction. Between collections, 100% of the CPU belongs to the application. **When it does collect, it collects with everybody.** All application threads are brought to a safepoint and stopped; then N GC threads work the heap in parallel, using per-thread local allocation buffers for copying and work-stealing queues so no thread idles while another still has a deep object graph to scan. That converts what a single-threaded collector would do in T seconds into roughly T/N seconds of wall clock, at the same total CPU cost. **The mutator tax is tiny.** Concurrent collectors need the application to cooperate: snapshot-at-the-beginning or incremental-update barriers on reference writes, load barriers on reads, extra bookkeeping so a moving collector can relocate objects while threads are running. The Parallel collector needs none of that. Its only ongoing cost is a cheap unconditional card mark on reference stores, so that a young collection can find old-to-young references without scanning the whole old generation. Allocation itself is a pointer bump in a thread-local buffer. **The heap stays simple.** Classic contiguous young generation (eden plus two survivor spaces) and a contiguous old generation. No region metadata, no remembered-set-per-region maintenance, no mixed-collection heuristics. Less machinery means less overhead and very predictable, cache-friendly copying. ## What it gives up **Full collections are one big stop.** The old generation is collected by a parallel mark-summary-compact pass over the whole heap in a single stop-the-world event. Compaction physically slides live objects, so the pause is roughly proportional to the amount of *live* data (plus a marking cost proportional to the live object graph). On a 1 GB heap that might be a few hundred milliseconds; on an 8–16 GB heap with a large live set, seconds. There is no way to do "part of" the old generation. **Pause goals are best-effort only.** You can set `-XX:MaxGCPauseMillis`, and the ergonomics will shrink generations to try to hit it, but there is no mechanism that can interrupt a full compaction once it starts. Treat the flag as a hint, never a guarantee. **No concurrent class unloading or reference processing.** Those happen inside the stop as well. **Latency-sensitive services suffer.** If your service has a p99 latency SLO measured in tens of milliseconds, a periodic multi-hundred-millisecond stop shows up directly in the tail, and every in-flight request pays for it simultaneously. ## Reading it in a log With unified logging (`-Xlog:gc`), Parallel collections appear as `Pause Young (Allocation Failure)` and `Pause Full (Ergonomics)`. There are no `Concurrent ...` lines at all — their absence is itself the signature of this collector. Older-format logs name the spaces `PSYoungGen` and `ParOldGen` (`PS` for "parallel scavenge"). ## Where the trade pays off Batch jobs, ETL, data processing, compilers, CI workloads, benchmarks, short-lived jobs, and any process whose success metric is "finish sooner" rather than "answer this request in under X ms." It also does well in small CPU-constrained containers, where a concurrent collector's background threads would be competing with the very application threads you are trying to keep fed. ## The sentence to say in an interview "Parallel does all its work stop-the-world with many threads. That maximises the proportion of time the app runs, because nothing is stolen from the app between collections and the barrier overhead is minimal — but a full collection compacts the entire heap in one stop, so pause length scales with live data."
- How would you actually measure whether the Parallel collector is giving you the throughput you want?Sum the pause durations from the GC log over a window and divide by the wall-clock length of that window; that is your GC overhead, and 1 minus it is throughput. `-Xlog:gc` plus a log parser, or the GC time counters exposed by the JVM's memory-management beans and jstat, both give you this. Compare it against the implicit goal of `-XX:GCTimeRatio` (default 99, i.e. 1% overhead) and against the job's total runtime, which is the number that actually matters for a batch workload.
- Does the Parallel collector use write barriers at all, given that nothing runs concurrently?Yes, but only a card-marking barrier. Every reference store into an object dirties the corresponding card so that a young collection can find old-to-young references without scanning the whole old generation. It is an unconditional store to a byte array — a couple of instructions with no branch. That is far cheaper than the snapshot-at-the-beginning or load barriers a concurrent collector needs, which is a large part of why Parallel's raw throughput is hard to beat.
- If Parallel has the best throughput, why isn't it still the default collector?Because the common case shifted. Most JVMs today run long-lived services with latency objectives and heaps of several gigabytes, where a whole-heap compaction pause is unacceptable even if it is rare. The default was changed to the region-based collector, which trades a few percent of throughput for pauses that scale with a bounded region set instead of with the whole live heap. Parallel remains the better pick where finishing time, not tail latency, is the metric.
A road crew that closes the whole motorway at 3 a.m. and repaves it with fifty workers at once. Total disruption is minimal and the work finishes fastest, but while it happens nobody drives at all — unlike a crew that coneS off one lane and works all week alongside traffic.
saying these in an interview costs you the question
- Saying the Parallel collector runs concurrently with the application — every one of its phases is stop-the-world
- Thinking 'parallel' refers to the application being multi-threaded rather than to the GC threads
- Claiming -XX:MaxGCPauseMillis guarantees a maximum pause; it is a best-effort ergonomic hint
- Assuming its full collection is single-threaded — modern Parallel does the old-generation mark-compact with all GC threads
- Dismissing it as obsolete; on CPU-constrained containers and batch jobs it still posts the best throughput