skip to content

Stop-the-World, Safepoints & GC Triggers

Why the collector needs every thread parked at a safepoint to see a consistent heap, what triggers minor versus full collections, and how modern collectors shrink the stop-the-world portion. Interviewers ask whenever latency spikes are on the table, because "the JVM paused" has to be explained mechanically.

on this pageshow

questions

6

What does a 'stop-the-world' pause mean inside a JVM, and why does a garbage collector need application threads suspended at all?

level: juniorimportance: must knowfreq 58%

answer

  1. all application threads halted; VM threads still working
  2. roots come from thread stacks — must be a snapshot
  3. moving objects needs references fixed atomically or via barriers
  4. modern collectors: concurrent phases + short STW phases
  5. not every pause is a GC pause

basics

~20 s

A stop-the-world pause is an interval where the JVM suspends every application thread so the runtime can do work that needs a stable view of the heap and of thread state — for example scanning thread stacks for references. Application code makes no progress for the duration.

solid answer

~50 s

During a stop-the-world pause the JVM brings **all** application threads to a halt, performs some VM-level work, then resumes them. No application code runs in between, so the pause shows up directly in request latency. The collector needs it because some of its work is only correct against a **consistent snapshot**. Finding the roots means reading every thread's stack and registers to discover which references are live; if a thread kept running it could move a reference from a place already scanned to a place already passed over, and a live object would be missed and freed. Moving (compacting) collectors also need to update every reference to a relocated object, which cannot be done safely while threads are dereferencing those references without extra machinery. Modern collectors shrink but do not eliminate pauses: they do marking, and sometimes relocation, concurrently with the application, and keep only short phases — root scanning, phase transitions — inside stop-the-world.

go deeper

for a junior

Define it plainly — all application threads stopped so the collector can work on a stable view — and give the root-scanning reason.

for a middle

Add why moving objects needs either a pause or barriers, and that collectors differ in how much work they place inside the pause.

for a senior

Connect pauses to tail latency and to operational consequences such as timeouts and health checks, and note that non-GC VM operations also stop the world.

for a principal

Frame it as a systems tradeoff: pause work versus concurrent work paid for with barrier overhead, CPU headroom and floating garbage, chosen against a latency objective.

## The definition A **stop-the-world (STW) pause** is a window in which the JVM has suspended every thread that runs application code, does some work that requires the world to hold still, and then releases them. From the application's point of view nothing executes: requests in flight simply stop advancing, and the pause is added to their latency. From the operating system's point of view the process is still alive and some VM threads are busy. The suspension is **cooperative**, not preemptive: the JVM does not yank threads at arbitrary instructions. Threads stop at defined points where the runtime knows exactly what each one is holding — but the essential idea at this level is simply that all application threads are stopped simultaneously and application progress halts. ## Why a collector needs the world still ### Finding the roots Garbage collection starts from **roots** — references held outside the heap that are definitely live: local variables and operands in every thread's frames, machine registers, static fields, JNI handles, and so on. To enumerate the roots in a thread, the runtime must read that thread's stack. A running thread mutates its stack constantly: it pushes and pops frames, it moves references between locals and the operand stack, it stores references into fields. If the collector scanned a live thread's stack, the thread could take the only reference to an object out of a location the collector has not yet reached and put it into one the collector already scanned. The collector would conclude the object is unreachable and reclaim memory that is still in use — a use-after-free inside a managed runtime, which is catastrophic and essentially undebuggable. Halting the threads makes each stack a fixed snapshot for the duration of the scan. ### Moving objects Collectors that **compact** — copying survivors into a fresh region to eliminate fragmentation — face a second problem. After an object moves, every reference to it must point to its new address. If application threads were running, one could read a reference between the copy and the update and dereference a stale address. A simple collector solves this by doing the whole move-and-fix-up inside a pause. Sophisticated ones use load or write barriers so that threads accessing a moved object are corrected on the fly, which is exactly how the low-pause collectors relocate objects concurrently. ## What varies is how much work sits in the pause Collectors differ enormously in how much they put inside STW: - **Fully stop-the-world collectors** do all marking and all copying/compaction inside the pause. Simple and efficient in CPU terms; the pause grows with the amount of live data. - **Mostly-concurrent collectors** trace the object graph while the application runs and keep only short pauses — typically an initial root scan and a termination/transition phase. - **Low-latency collectors** additionally relocate objects concurrently, using barriers, and aim for pauses that are short and largely independent of heap size. No mainstream collector is pause-free, because agreeing on a phase change and capturing thread roots requires some point at which threads are known to be in a defined state. ## GC is not the only reason the world stops A collection is the most common cause but not the only one. The runtime also needs a globally quiet moment for operations such as deoptimizing compiled code, redefining classes (as an agent or debugger does), taking a heap dump, and various diagnostic operations. So a pause observed in an application is not automatically a GC pause, and a log that shows GC accounting for less pause time than the application experienced is a normal and important observation, not a contradiction. ## What it means for an application - Pauses are **added latency at the tail**. Average throughput can look fine while the 99th percentile carries the pauses. - **Longer pauses are not proportional to garbage**; in a tracing collector the dominant cost is usually the *live* set that must be traced and possibly copied. - A pause suspends *application* threads. VM threads doing the collection are running, so CPU usage during a pause can be high, not zero. - Because pauses are added latency, timeouts, health checks and leader-election leases all need headroom greater than the worst pause you tolerate, or a long pause turns into a cascading failure. ## Summary Stop-the-world means every application thread halted so the runtime can act on a consistent view of thread state and the heap — above all to enumerate roots safely, and in simple collectors to move objects safely. Modern collectors move most of the work outside the pause but keep a few short phases inside it, and the runtime uses global pauses for non-GC operations too.

  • If a collector is described as 'concurrent', why does it still pause the application at all?
    Concurrent means most of the tracing, and in some collectors the relocation, happens while application threads run. But capturing each thread's roots and agreeing on a phase transition still needs each thread to be in a known state, so short pauses or per-thread handshakes remain. The goal is pauses that are short and roughly independent of heap size, not zero pauses.
  • During a stop-the-world pause, is the CPU idle?
    No. Application threads are suspended, but the collector's own threads are typically working hard, often on several cores. That is why a pause can coincide with a CPU spike and why giving the collector more parallel threads can shorten pauses at the cost of CPU that the application would otherwise use.

Taking inventory of a warehouse while forklifts keep moving pallets: you would count some pallets twice and miss others. Either everyone stands still while you count, or you install a system that logs every move as it happens — which is exactly the concurrent collector's barrier machinery.

saying these in an interview costs you the question

  • Saying stop-the-world means the whole process, including the collector, is frozen.
  • Claiming modern collectors are entirely pause-free.
  • Assuming pause length is proportional to the amount of garbage rather than to the live set and the work placed in the pause.
  • Treating every observed application pause as a GC pause without checking other VM operations.
  • Believing threads can be suspended at any arbitrary instruction with no cooperation from the running code.

context

open as a page

What events cause a JVM to start a young (minor) collection, and what causes it to fall back to a full collection of the whole heap?

level: middleimportance: must knowfreq 58%

basics

~20 s

A young collection is triggered when allocation cannot be satisfied from the young space — Eden is full. A full collection happens when the old generation cannot accept what must be promoted or has no room, when the class-metadata area hits its threshold, or when something explicitly requests one; in modern collectors it is mostly a fallback after concurrent work fails to keep up.

open as a page

What is a JVM safepoint, and why can an application thread only be suspended for a VM operation when it has reached one?

level: middleimportance: must knowfreq 50%

basics

~20 s

A safepoint is a point in a thread's execution where the runtime knows exactly where every object reference that thread holds is located, because the compiler recorded a map for that point. Threads can only be paused there; suspending at an arbitrary instruction would leave references the collector cannot find or update.

open as a page

How does HotSpot make a running application thread actually stop when a safepoint is requested — where are the polls placed in compiled code, and what happens to a thread that is executing native code at that moment?

level: seniorimportance: should knowfreq 32%

basics

~20 s

The compiler emits cheap polls — a load from a special polling page — at method returns and at loop back-edges. When a safepoint is requested the VM makes that page unreadable (or flips a thread-local poll word), so the next poll traps into the runtime and parks the thread. Threads in native code are already counted as safe and are blocked when they try to return.

open as a page

An application records a 300 ms pause while the JVM's collector reports only 8 ms of collection work for that event. What accounts for the difference, and what typically makes threads slow to reach a safepoint?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Total stopped time equals the time for every thread to reach a safepoint plus the time the operation runs plus resume. Here roughly 292 ms was time-to-safepoint: one or more threads did not poll promptly. Common causes are long-running counted loops, huge uninterruptible operations, descheduled threads on an oversubscribed host, and page faults on swapped-out stacks.

open as a page

Modern JVM collectors perform most of their work concurrently with the running application but still take short stop-the-world pauses. What work is placed in those pauses, and what does moving work out of them cost the system?

level: principalimportance: should knowfreq 34%

basics

~20 s

Pauses retain the work that needs global agreement or a stable per-thread view: capturing roots from thread stacks, switching collector phase and barrier state, and finishing marking. Moving the rest off the pause costs throughput through barriers on every reference access, CPU shared with the application, floating garbage, and heap headroom so the concurrent cycle can finish before the heap fills.

open as a page