skip to content

What does a 'stop-the-world' pause mean inside a JVM, and why does a garbage collector need application threads suspended at all?

level: juniorimportance: must knowfreq 58%

answer

  1. all application threads halted; VM threads still working
  2. roots come from thread stacks — must be a snapshot
  3. moving objects needs references fixed atomically or via barriers
  4. modern collectors: concurrent phases + short STW phases
  5. not every pause is a GC pause

basics

~20 s

A stop-the-world pause is an interval where the JVM suspends every application thread so the runtime can do work that needs a stable view of the heap and of thread state — for example scanning thread stacks for references. Application code makes no progress for the duration.

solid answer

~50 s

During a stop-the-world pause the JVM brings **all** application threads to a halt, performs some VM-level work, then resumes them. No application code runs in between, so the pause shows up directly in request latency. The collector needs it because some of its work is only correct against a **consistent snapshot**. Finding the roots means reading every thread's stack and registers to discover which references are live; if a thread kept running it could move a reference from a place already scanned to a place already passed over, and a live object would be missed and freed. Moving (compacting) collectors also need to update every reference to a relocated object, which cannot be done safely while threads are dereferencing those references without extra machinery. Modern collectors shrink but do not eliminate pauses: they do marking, and sometimes relocation, concurrently with the application, and keep only short phases — root scanning, phase transitions — inside stop-the-world.

go deeper

for a junior

Define it plainly — all application threads stopped so the collector can work on a stable view — and give the root-scanning reason.

for a middle

Add why moving objects needs either a pause or barriers, and that collectors differ in how much work they place inside the pause.

for a senior

Connect pauses to tail latency and to operational consequences such as timeouts and health checks, and note that non-GC VM operations also stop the world.

for a principal

Frame it as a systems tradeoff: pause work versus concurrent work paid for with barrier overhead, CPU headroom and floating garbage, chosen against a latency objective.

## The definition A **stop-the-world (STW) pause** is a window in which the JVM has suspended every thread that runs application code, does some work that requires the world to hold still, and then releases them. From the application's point of view nothing executes: requests in flight simply stop advancing, and the pause is added to their latency. From the operating system's point of view the process is still alive and some VM threads are busy. The suspension is **cooperative**, not preemptive: the JVM does not yank threads at arbitrary instructions. Threads stop at defined points where the runtime knows exactly what each one is holding — but the essential idea at this level is simply that all application threads are stopped simultaneously and application progress halts. ## Why a collector needs the world still ### Finding the roots Garbage collection starts from **roots** — references held outside the heap that are definitely live: local variables and operands in every thread's frames, machine registers, static fields, JNI handles, and so on. To enumerate the roots in a thread, the runtime must read that thread's stack. A running thread mutates its stack constantly: it pushes and pops frames, it moves references between locals and the operand stack, it stores references into fields. If the collector scanned a live thread's stack, the thread could take the only reference to an object out of a location the collector has not yet reached and put it into one the collector already scanned. The collector would conclude the object is unreachable and reclaim memory that is still in use — a use-after-free inside a managed runtime, which is catastrophic and essentially undebuggable. Halting the threads makes each stack a fixed snapshot for the duration of the scan. ### Moving objects Collectors that **compact** — copying survivors into a fresh region to eliminate fragmentation — face a second problem. After an object moves, every reference to it must point to its new address. If application threads were running, one could read a reference between the copy and the update and dereference a stale address. A simple collector solves this by doing the whole move-and-fix-up inside a pause. Sophisticated ones use load or write barriers so that threads accessing a moved object are corrected on the fly, which is exactly how the low-pause collectors relocate objects concurrently. ## What varies is how much work sits in the pause Collectors differ enormously in how much they put inside STW: - **Fully stop-the-world collectors** do all marking and all copying/compaction inside the pause. Simple and efficient in CPU terms; the pause grows with the amount of live data. - **Mostly-concurrent collectors** trace the object graph while the application runs and keep only short pauses — typically an initial root scan and a termination/transition phase. - **Low-latency collectors** additionally relocate objects concurrently, using barriers, and aim for pauses that are short and largely independent of heap size. No mainstream collector is pause-free, because agreeing on a phase change and capturing thread roots requires some point at which threads are known to be in a defined state. ## GC is not the only reason the world stops A collection is the most common cause but not the only one. The runtime also needs a globally quiet moment for operations such as deoptimizing compiled code, redefining classes (as an agent or debugger does), taking a heap dump, and various diagnostic operations. So a pause observed in an application is not automatically a GC pause, and a log that shows GC accounting for less pause time than the application experienced is a normal and important observation, not a contradiction. ## What it means for an application - Pauses are **added latency at the tail**. Average throughput can look fine while the 99th percentile carries the pauses. - **Longer pauses are not proportional to garbage**; in a tracing collector the dominant cost is usually the *live* set that must be traced and possibly copied. - A pause suspends *application* threads. VM threads doing the collection are running, so CPU usage during a pause can be high, not zero. - Because pauses are added latency, timeouts, health checks and leader-election leases all need headroom greater than the worst pause you tolerate, or a long pause turns into a cascading failure. ## Summary Stop-the-world means every application thread halted so the runtime can act on a consistent view of thread state and the heap — above all to enumerate roots safely, and in simple collectors to move objects safely. Modern collectors move most of the work outside the pause but keep a few short phases inside it, and the runtime uses global pauses for non-GC operations too.

  • If a collector is described as 'concurrent', why does it still pause the application at all?
    Concurrent means most of the tracing, and in some collectors the relocation, happens while application threads run. But capturing each thread's roots and agreeing on a phase transition still needs each thread to be in a known state, so short pauses or per-thread handshakes remain. The goal is pauses that are short and roughly independent of heap size, not zero pauses.
  • During a stop-the-world pause, is the CPU idle?
    No. Application threads are suspended, but the collector's own threads are typically working hard, often on several cores. That is why a pause can coincide with a CPU spike and why giving the collector more parallel threads can shorten pauses at the cost of CPU that the application would otherwise use.

Taking inventory of a warehouse while forklifts keep moving pallets: you would count some pallets twice and miss others. Either everyone stands still while you count, or you install a system that logs every move as it happens — which is exactly the concurrent collector's barrier machinery.

saying these in an interview costs you the question

  • Saying stop-the-world means the whole process, including the collector, is frozen.
  • Claiming modern collectors are entirely pause-free.
  • Assuming pause length is proportional to the amount of garbage rather than to the live set and the work placed in the pause.
  • Treating every observed application pause as a GC pause without checking other VM operations.
  • Believing threads can be suspended at any arbitrary instruction with no cooperation from the running code.

context