skip to content

Garbage Collection

Garbage collection from the algorithms up: marking, copying, compaction, safepoints and the collectors you can choose. Asked in any round on latency or footprint, since collector choice visibly changes production.

on this pageshow

questions

page 2 of 2

How does HotSpot make a running application thread actually stop when a safepoint is requested — where are the polls placed in compiled code, and what happens to a thread that is executing native code at that moment?

level: seniorimportance: should knowfreq 32%

basics

~20 s

The compiler emits cheap polls — a load from a special polling page — at method returns and at loop back-edges. When a safepoint is requested the VM makes that page unreadable (or flips a thread-local poll word), so the next poll traps into the runtime and parks the thread. Threads in native code are already counted as safe and are blocked when they try to return.

open as a page

An application records a 300 ms pause while the JVM's collector reports only 8 ms of collection work for that event. What accounts for the difference, and what typically makes threads slow to reach a safepoint?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Total stopped time equals the time for every thread to reach a safepoint plus the time the operation runs plus resume. Here roughly 292 ms was time-to-safepoint: one or more threads did not poll promptly. Common causes are long-running counted loops, huge uninterruptible operations, descheduled threads on an oversubscribed host, and page faults on swapped-out stacks.

open as a page

ZGC stores garbage-collection metadata inside the 64-bit object reference itself — the technique known as colored pointers. What is kept in those bits, how does the JVM still dereference such a pointer correctly, and what constraints does the technique impose on the platform and on memory footprint?

level: seniorimportance: should knowfreq 40%

basics

~20 s

ZGC puts GC state — marked, remapped, and in the generational design generation and remembered-set bits — into unused bits of every 64-bit heap reference. Barriers interpret and strip the color before use. It requires 64-bit addressing, bounds the address space, and rules out compressed oops.

open as a page

ZGC compacts the heap by moving live objects while application threads keep reading and writing them. Describe how it chooses what to move, how it prevents threads from working on a stale copy, and at what point the old memory can be handed back for reuse.

level: seniorimportance: should knowfreq 35%

basics

~20 s

ZGC selects the pages holding the most garbage as the relocation set, then copies their live objects concurrently. Off-heap forwarding tables map old to new addresses, and load barriers redirect — or perform — each move via compare-and-set. Pages are freed immediately; stale references get remapped lazily.

open as a page

HotSpot's ZGC compiles a load barrier into application code. When does that barrier run, what does its fast path test, what can happen on the slow path, and what does it mean that the barrier is 'self-healing'?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Every read of a reference field from the heap runs a load barrier. The fast path tests the pointer's color against the current good color; a bad color takes a slow path that marks the object or looks up its new location, then writes the corrected pointer back into the field it was read from — self-healing.

open as a page

HotSpot's Concurrent Mark-Sweep collector was deprecated in JDK 9 and removed entirely in JDK 14. If you owned a latency-sensitive service still pinned to an old JDK because of it, how would you reason about the migration and what would you expect to change operationally?

level: principalimportance: should knowfreq 34%

basics

~20 s

Treat it as unavoidable: staying pinned costs security updates and language features for a collector nobody maintains. Move to the region-based collector (G1) as the default — it compacts as part of normal collection, removing the fragmentation failure mode. Drop CMS-specific flags rather than translating them, set a pause-time goal and heap size, then validate against production-shaped load.

open as a page

The Garbage-First collector accepts a pause-time goal via -XX:MaxGCPauseMillis, but the goal is explicitly soft. Explain the machinery behind it and how you would reason about a service whose latency requirement is stricter than what that goal delivers.

level: principalimportance: should knowfreq 38%

basics

~20 s

G1 keeps statistics on past pauses and predicts the cost of evacuating candidate regions, then sizes the young generation and the collection set so the predicted pause fits the goal. It is a prediction, not a guarantee: mispredictions, humongous allocation, evacuation failure, and full GCs all overshoot it.

open as a page

You are choosing a garbage collector for a JVM workload. Under what conditions would you deliberately pick the throughput-oriented Parallel collector (-XX:+UseParallelGC) instead of the modern default region-based collector, and how would you justify that call with evidence?

level: principalimportance: should knowfreq 38%

basics

~20 s

Choose it when the success metric is completion time or cost rather than tail latency, and when CPU is scarce: batch and ETL jobs, CI and build workers, short-lived tasks, and small containers with one or two cores where concurrent GC threads would compete with application threads. Justify with measured GC overhead and end-to-end runtime, not with a preference.

open as a page

Modern JVM collectors perform most of their work concurrently with the running application but still take short stop-the-world pauses. What work is placed in those pauses, and what does moving work out of them cost the system?

level: principalimportance: should knowfreq 34%

basics

~20 s

Pauses retain the work that needs global agreement or a stable per-thread view: capturing roots from thread stacks, switching collector phase and barrier state, and finishing marking. Moving the rest off the pause costs throughput through barriers on every reference access, CPU shared with the application, floating garbage, and heap headroom so the concurrent cycle can finish before the heap fills.

open as a page

You are picking a garbage collector for a latency-sensitive JVM service and are considering ZGC, the sub-millisecond colored-pointer collector. What do you gain, what do you pay for it, and what failure mode should you size capacity around?

level: principalimportance: should knowfreq 30%

basics

~20 s

You gain GC pauses under a millisecond regardless of heap or live-set size. You pay a barrier cost on reference loads, CPU for concurrent GC threads, extra heap headroom, and the loss of compressed references. The failure mode is allocation stalls, not long pauses — size CPU and heap for that.

open as a page

A service maintains large mutable object graphs and rewrites reference fields constantly. How would you reason about the throughput cost that a concurrent collector's write barriers impose on that workload, and what levers exist?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

Barrier cost scales with reference-store frequency, not heap size. Measure it by comparing collectors with different barrier designs on real traffic, watching application throughput rather than pause charts. Levers: reduce reference writes in hot code, prefer primitive or immutable structures, size regions and heap so cross-region traffic falls, and pick a collector whose barrier profile fits.

open as a page

You are choosing between a moving collection strategy (copying or compacting) and a non-moving one (mark-sweep) for a runtime. Lay out the trade-offs that decide it, including effects on the allocation path, fragmentation risk, and code that holds raw object addresses.

level: principalimportance: nice to knowfreq 34%

basics

~20 s

Moving buys bump-pointer allocation, no fragmentation, and better locality, paid for with relocation cost proportional to live data, precise reference information, and address instability that native code must tolerate. Non-moving keeps addresses stable and collection cheap per cycle, paid for with free-list allocation and fragmentation that can eventually block large allocations.

open as a page

You operate several hundred small JVM services on shared nodes, each with a heap under 512 MB. How would you decide whether to standardize the fleet on the single-threaded stop-the-world collector, and what would you measure to defend the decision?

level: principalimportance: nice to knowfreq 22%

basics

~20 s

Split the fleet by SLO class rather than deciding once. Measure per-service live set, full-GC pause distribution, container RSS, and CPU-seconds per request. Standardize the single-threaded collector as the default for low-SLO and short-lived services, with an explicit, reviewed override for latency-critical ones.

open as a page

What does adopting a concurrent-compaction garbage collector such as Shenandoah actually cost in CPU and memory, and how would you validate that a given service genuinely benefits from it?

level: principalimportance: nice to knowfreq 21%

basics

~20 s

It costs throughput (barriers on reference loads run on every request path), CPU (collector threads run alongside the application), and memory headroom (it must have free regions to evacuate into). Validate with real traffic: compare caller-side p99.9, CPU-seconds per request, RSS, and the absence of degenerated or full collections.

open as a page

showing 31–44 of 44