skip to content

How does HotSpot make a running application thread actually stop when a safepoint is requested — where are the polls placed in compiled code, and what happens to a thread that is executing native code at that moment?

level: seniorimportance: should knowfreq 32%

answer

  1. poll = load from a polling page; arm it by protecting the page
  2. thread-local poll word → handshakes, per-thread stops
  3. placed at method returns + loop back-edges
  4. counted int loop omitted the poll → long time to safepoint
  5. native/blocked threads already safe; blocked at the transition back

basics

~20 s

The compiler emits cheap polls — a load from a special polling page — at method returns and at loop back-edges. When a safepoint is requested the VM makes that page unreadable (or flips a thread-local poll word), so the next poll traps into the runtime and parks the thread. Threads in native code are already counted as safe and are blocked when they try to return.

solid answer

~60 s

Suspension is cooperative, so generated code must ask. HotSpot emits a **poll**: an instruction that reads a designated polling page. In the fast path the read succeeds and costs almost nothing. To request a safepoint, the VM protects that page; the next poll faults, the JVM's signal handler recognises the address, and the thread parks itself at the safepoint. Modern HotSpot uses a **thread-local** poll word so it can also stop individual threads (handshakes) rather than only the whole world. **Placement** is chosen so a thread reaches a poll quickly without paying for polls everywhere: at method **returns** (or entries), and at the **back-edge of loops**. The interpreter checks on its dispatch path. The historically important gap is the **counted loop**: a loop with an `int` induction variable and a known trip count had its poll omitted as an optimization, so a long-running counted loop could delay the safepoint for a very long time. HotSpot now **strip-mines** such loops — running them in short inner chunks inside an outer loop that does contain a poll. A thread in native code needs no poll: its state is already safe, and it is held at the transition on return.

code

java · 7 lines
java
// int induction variable + compiler-known trip count = 'counted loop'
for (int i = 0; i < arr.length; i++) {
    total += expensiveButInlinable(arr[i]);
}
// Without loop strip mining, a safepoint requested mid-loop waits for the
// whole loop to finish. A long induction variable (long i = ...) historically
// kept its poll and did not show the problem.

go deeper

for a junior

Know that generated code periodically checks whether a safepoint was requested, and that the JVM cannot simply freeze a thread anywhere.

for a middle

Describe the polling page mechanism and the placement at returns and loop back-edges, and that native threads are already safe.

for a senior

Explain the counted-loop hole and strip mining, thread-local polls enabling handshakes, the transition block on return from native, and how to read safepoint logs.

for a principal

Discuss the design tension — near-zero fast-path cost versus bounded arrival latency — and its consequences for latency-critical services, profiler safepoint bias, and where the collector's pause floor really comes from.

## Cooperative suspension needs a cheap question Since a thread can only stop where its references are describable, the runtime cannot force it — the generated code must periodically ask "should I stop?". Two requirements pull against each other: the check must be *almost free* on the fast path, since it executes constantly, and threads must reach one *quickly*, since the whole application waits for the slowest. ## The polling trick HotSpot's answer is to turn the check into a memory read. The compiler emits an instruction that loads from a designated **polling page**. Normally the page is readable, the load succeeds, and nothing happens — no branch, no comparison, negligible cost and no effect on the surrounding optimization. To request a safepoint, the VM makes the page inaccessible. The next poll executed by any thread takes a fault; the JVM's signal handler recognises the faulting address as the polling page, identifies the safepoint the thread reached, uses the recorded reference map for that exact program point, and parks the thread in a blocked state until the operation completes. Modern HotSpot refines this with a **thread-local poll**: each thread has its own poll word/page pointer, so the VM can arm the poll for one thread only. That is what makes **thread-local handshakes** possible — performing an operation in one thread's context without stopping the world — and it is a building block for collectors that scan thread stacks one thread at a time to keep pauses short and independent of heap size. The **interpreter** does not need generated polls in the same way: its dispatch and branch handling check the safepoint state directly. ## Where polls are placed Emitting a poll at every instruction would be intolerable, so the compiler places them where they bound the time to the next one: - **Method return** (equivalently, on entry in some configurations). Any call chain therefore hits polls regularly, because straight-line code without loops or calls is inherently short. - **Loop back-edges.** A loop can run arbitrarily long, so the back-edge is the natural place: the poll executes once per iteration. Between these, execution is bounded by construction — a straight-line stretch of code has a finite instruction count, so the thread will reach a poll soon. ## The counted-loop hole and loop strip mining The classic exception, and the thing a senior candidate is expected to know: HotSpot's optimizer historically **omitted the back-edge poll from counted loops** — loops with an `int` induction variable and a compiler-known trip count — because the loop was assumed to terminate promptly and the poll interfered with vectorisation and other loop optimizations. The assumption fails badly when the loop body is expensive: ```java for (int i = 0; i < 1_000_000; i++) { heavyWork(data[i]); // if this is inlined and long-running, no poll runs } ``` If a safepoint is requested while such a loop is executing, every other thread waits until the loop *finishes*. This produced the notorious symptom of a several-hundred-millisecond "GC pause" whose collection work was a couple of milliseconds. The fix is **loop strip mining**: the compiler rewrites the counted loop as an outer loop over short inner chunks, and places the poll on the outer loop's back-edge. The inner loop keeps its optimizations; the thread still checks in every few thousand iterations. In modern HotSpot this is enabled by default with the low-pause collectors (`-XX:+UseCountedLoopSafepoints`, with a strip length flag). A `long`-typed induction variable does not qualify as counted in the same way and has historically kept its poll. ## Threads that are not running Java code Not every thread needs to hit a poll: - A thread **executing native code** via JNI is in a distinct runtime state. It cannot access Java objects without going back through a transition, so its Java frames are already stable and describable — it is counted as *at a safepoint* the moment the request is made. The safepoint therefore does not wait for it. When the native call returns, the transition checks the safepoint state and **blocks the thread there** if the operation is still in progress. - A thread **blocked** on a monitor, in `park`, or waiting on I/O within the runtime is similarly in a known state and counts as safe. This is why a JNI call that takes a minute does not delay a collection — but also why a JNI critical section, which pins array memory and must not have the array moved, *can* interfere with the operation itself. ## Observing it Unified logging exposes the accounting directly: `-Xlog:safepoint` reports, per safepoint, how long threads took to reach it and how long the operation ran. That split is the diagnostic tool, and it is what separates "the collector is slow" from "one thread would not stop". ## Summary Polls are near-free page reads emitted at returns and loop back-edges; arming them is a memory-protection flip (or a thread-local word flip for handshakes). The known hazard is a long-running counted loop whose poll was optimized away — addressed by strip mining. Threads in native or blocked states are safe by construction and are held at the transition on the way back.

  • Why is a polling page read preferred over an explicit 'if (safepointRequested)' branch in generated code?
    A plain load has no branch to predict, needs no comparison, and does not perturb the surrounding scheduling or register allocation, so its fast-path cost is close to zero. Arming it is a single memory-protection change that affects all threads at once, and the resulting fault carries the exact program counter, which lets the runtime look up the reference map for precisely that point.
  • How would you confirm that a long pause was caused by slow arrival at the safepoint rather than by the operation itself?
    Enable safepoint logging, which reports the time spent reaching the safepoint separately from the time spent at it. A large 'reaching' figure with a small 'at safepoint' figure points at a thread that would not poll — a counted loop, a long uninterruptible operation, or a descheduled or page-faulting thread — rather than at the collector's work.

saying these in an interview costs you the question

  • Saying the VM sends a signal that stops a thread wherever it happens to be executing.
  • Claiming polls are emitted at every instruction or every basic block.
  • Believing a thread running a long JNI call delays the safepoint — it is already counted as safe.
  • Asserting that counted loops still have no safepoint poll in modern HotSpot, ignoring loop strip mining.
  • Describing the poll as an expensive check whose overhead is a meaningful tuning target.

context