An application records a 300 ms pause while the JVM's collector reports only 8 ms of collection work for that event. What accounts for the difference, and what typically makes threads slow to reach a safepoint?
answer
- stopped = reaching + at safepoint + resume
- waits for the LAST thread — it is a max, not an average
- counted loop / bulk arraycopy / not scheduled / page fault / thread count
- -Xlog:safepoint splits the two numbers
- GC dashboards show only 'at safepoint'
basics
~20 sTotal stopped time equals the time for every thread to reach a safepoint plus the time the operation runs plus resume. Here roughly 292 ms was time-to-safepoint: one or more threads did not poll promptly. Common causes are long-running counted loops, huge uninterruptible operations, descheduled threads on an oversubscribed host, and page faults on swapped-out stacks.
solid answer
~60 sA pause has two halves, and only one is the collector's: ``` stopped time = time-to-safepoint (waiting for the slowest thread) + operation time + resume ``` A collector reporting 8 ms of work inside a 300 ms pause means ~292 ms was spent waiting for threads to arrive. Because the operation cannot start until the **last** thread parks, one slow thread stalls everyone. Typical causes: - **A counted loop without a poll** — historically an `int`-indexed loop with an expensive inlined body ran to completion before checking. Loop strip mining addresses this. - **Long uninterruptible operations** — a huge `System.arraycopy`, `Arrays.fill`, or a bulk intrinsic that contains no poll. - **CPU oversubscription** — a thread that is runnable but not scheduled cannot poll. Common in containers with tight CPU limits or noisy neighbours. - **Page faults** — a thread whose stack pages were swapped out must fault them in before it can proceed. - **Many threads**, each needing to arrive and be resumed. Diagnose with `-Xlog:safepoint`, which prints the two figures separately.
code
text · 5 lines[141.902s][info][safepoint] Safepoint "G1CollectForAllocation", Time since last: 812 ms,
Reaching safepoint: 291744010 ns, At safepoint: 8102773 ns, Total: 299846783 ns
// 291.7 ms reaching + 8.1 ms of actual collection = ~300 ms of application stall.
// Tuning the collector would address 8 ms of a 300 ms problem.go deeper
Know that the collector's reported time is not the whole pause and that threads must first be brought to a stop.
State the reaching + at-safepoint decomposition, know that the operation waits for the slowest thread, and know that safepoint logging shows both numbers.
Run the triage: read the log, separate operation-slow from arrival-slow, and enumerate causes from counted loops and bulk array operations to CPU throttling and swapping, with a fix for each.
Treat stopped time rather than GC time as the latency SLO input, provision CPU headroom as a pause-time parameter, and set observability so that reaching-safepoint is monitored rather than discovered during an incident.
## Two clocks, and the one people forget When the JVM performs a stop-the-world operation, the application experiences: ``` total stopped time = time to safepoint (from the request until every thread has parked) + time at safepoint (the operation itself — e.g. the collection) + resume time (releasing threads) ``` GC logs report the middle term. The application — through request latency, a heartbeat that missed, or `-XX:+PrintGCApplicationStoppedTime`-style accounting — experiences the sum. A 300 ms pause containing 8 ms of collection is therefore not a contradiction and not a broken log; it is a **time-to-safepoint (TTSP)** problem, and it is diagnosed and fixed completely differently from a slow collector. The structural reason TTSP dominates so dramatically when it goes wrong: the operation cannot begin until the **last** thread has arrived. It is a max, not an average. Nine hundred threads arriving in 20 µs and one arriving in 290 ms produces a 290 ms wait. ## What makes a thread slow to arrive ### 1. Compiled code with no poll on the hot path Polls sit at method returns and loop back-edges. The historical hole is the **counted loop**: an `int` induction variable with a compiler-known trip count had its back-edge poll optimized away. If the body is expensive but inlinable — so there is no call return to poll at either — the thread checks in only when the loop ends. ```java for (int i = 0; i < 50_000_000; i++) { acc = mix(acc, data[i]); // inlined; no return, no back-edge poll } ``` Modern HotSpot mitigates this with **loop strip mining**, which wraps the counted loop in an outer loop carrying a poll. If you are on an older configuration, or the flag is off, the symptom returns. ### 2. Long uninterruptible runtime operations Some runtime routines contain no poll at all: copying or filling a very large array, certain compression or encoding intrinsics, or a large `Object.clone()`. The thread is inside a single runtime call for milliseconds. Nothing is wrong with the code — the operation is simply atomic from the safepoint machinery's point of view. ### 3. The thread is not running A thread that is **runnable but not scheduled** cannot execute a poll. This is the cause most often missed, and it is increasingly common: - Containers with a low CPU quota, where the cgroup throttles the process for tens of milliseconds at a time. - Hosts with more busy threads than cores, or a noisy neighbour VM. - The collector's own parallel threads competing with application threads for the same cores. The pause is then not really the JVM's fault; the host cannot give the JVM the CPU it needs to stop cleanly. ### 4. Memory faults on the way to the poll If the machine is swapping, or pages were reclaimed under memory pressure, a thread may need to fault in its stack pages or the code page containing the poll. Disk-backed faults are milliseconds each. Any JVM on a swapping host will show erratic TTSP; this is a leading argument for disabling swap for latency-sensitive services. ### 5. Sheer thread count Every thread must be brought to a stop and later resumed. With very large numbers of platform threads the fixed cost per thread starts to matter, and the probability that at least one is in an awkward state rises. ### 6. Storms of frequent VM operations Sometimes TTSP is fine per event but the *rate* of safepoints is the problem — many small operations each with their own arrival and resume cost. Safepoint logging shows the interval between safepoints alongside the timings, which makes this shape obvious. ## How to diagnose Enable safepoint logging and read the two numbers: ``` -Xlog:safepoint ``` Each entry reports time since the last safepoint, time spent *reaching* the safepoint, and time spent *at* it. The triage is mechanical: - **High "at safepoint", low "reaching"** → the operation itself is slow. That is a collector/heap sizing question. - **High "reaching", low "at safepoint"** → TTSP. Now decide between the causes above: check host CPU saturation and throttling first (cheapest to rule out), then swapping, then look for long counted loops or bulk array operations in the code that runs at those moments. - **Low both, but very frequent** → safepoint rate; find what is requesting them. More detailed diagnostics can attribute the delay to a specific thread, and a profiler that is not safepoint-biased is essential here, since a safepoint-based profiler is blind precisely where the problem is. ## Remedies - Ensure loop strip mining is active on your configuration. - Break very large bulk array operations into chunks so polls occur between them. - Give the process enough CPU headroom that threads are actually scheduled; treat container CPU limits as a latency parameter, not just a cost one. - Disable swap, or ensure the heap and stacks stay resident. - Reduce needless safepoint-requiring operations (frequent thread dumps, diagnostic polling, repeated class redefinition by agents). ## Why it matters beyond the number TTSP is invisible to almost every GC dashboard, so teams tune collectors for months against a problem the collector does not have. Being able to say "stopped time is not GC time; show me the reaching-safepoint figure" is the whole value of this question.
- The safepoint log shows a large 'reaching safepoint' figure and the host is at 100% CPU with more busy threads than cores. What is the most likely explanation?Threads that are runnable but not scheduled cannot execute a safepoint poll, so arrival is delayed by the scheduler rather than by anything in the JVM. Container CPU quotas make this worse, because the cgroup can throttle the whole process for tens of milliseconds at a time. The fix is CPU headroom — fewer application threads, fewer collector threads, or a higher quota — not collector tuning.
- Why can a safepoint-based sampling profiler fail to show you the code responsible for a long time-to-safepoint?Such a profiler obtains stack traces at safepoints, so it can only sample code that has reached one. The code that is slow to reach a safepoint is precisely the code it cannot sample, and its samples are attributed to the nearest poll instead. Diagnosing TTSP requires a profiler that samples without requiring a safepoint.
saying these in an interview costs you the question
- Treating reported GC time as the application's stopped time.
- Concluding a collector is slow when the reaching-safepoint figure carries the pause.
- Ignoring host-level causes — CPU throttling, oversubscription, swapping — and tuning heap sizes instead.
- Assuming the pause is the average arrival time rather than the slowest thread's arrival.
- Trusting a safepoint-biased profiler to locate the offending loop.