skip to content

How do garbage collection and safepoint pauses distort a homegrown Java microbenchmark, and how should you account for them?

level: seniorimportance: should knowfreq 55%

answer

  1. STW GC + safepoints freeze all threads, injecting pauses
  2. Safepoints also for deopt / biased-lock revoke / thread dump
  3. Mean hides spikes; report percentiles + max
  4. Don't blindly trim outliers — the op's own allocation is real cost
  5. Use JMH -prof gc for allocation rate; fork fresh JVMs

basics

~20 s

Garbage collection can pause your code in the middle of timing, adding spikes that aren't part of the operation you're measuring. The JVM also pauses all threads at 'safepoints' for housekeeping. If you report only an average, these pauses get hidden or smeared; you should look at the distribution and control allocation.

solid answer

~50 s

The JVM periodically stops application threads: garbage collection reclaims dead objects, and many GCs include 'stop-the-world' phases; the JVM also brings all threads to a safepoint for tasks like deoptimization, biased-lock revocation, or thread dumps. Both inject pauses that have nothing to do with the operation under test. In a naive benchmark these show up as occasional huge spikes; if you report only the mean, a few multi-millisecond GC pauses can inflate the average, while if you trim outliers you may hide real allocation cost the operation genuinely incurs. The right approach: minimize incidental allocation so GC reflects only the work being measured, run long enough that GC behavior is representative, report the full distribution (percentiles, max), not just the mean, and isolate per-iteration timing from GC by looking at percentiles. JMH helps by forking fresh JVMs, supporting -prof gc to attribute allocation rate, and reporting percentiles, so you can distinguish 'this operation allocates' from 'a stray GC perturbed one iteration.'

go deeper

for a junior

Knows GC can pause the program and add random spikes to timings, so a single average can be misleading.

for a middle

Explains stop-the-world GC, that pauses can land mid-iteration, and that you should look at percentiles and control allocation rather than trust the mean.

for a senior

Distinguishes safepoints from GC, knows safepoints serve deopt/lock-revocation/dumps, separates the operation's legitimate allocation cost from stray perturbation, and uses -prof gc + percentiles.

for a principal

Connects pause behavior to collector choice and workload SLOs, recognizes coordinated omission, and decides when bytes/op or percentile latency is the right metric versus mean throughput.

## Two sources of involuntary pauses A Java program does not have uninterrupted control of the CPU. The JVM steps in at two related mechanisms: ### Garbage collection (GC) Java manages memory automatically: objects you allocate live on the **heap**, and a **garbage collector** periodically reclaims those that are no longer reachable. Most collectors have at least some **stop-the-world (STW)** phases — moments where *all* application threads are frozen so the collector can safely scan or move objects. Even "mostly concurrent" collectors (G1, ZGC, Shenandoah) have short STW pauses. A GC pause can be microseconds to many milliseconds depending on collector and heap. ### Safepoints A **safepoint** is a point in execution where the JVM knows the precise state of every thread (where all references are, etc.), so it can safely do global work: GC, **deoptimization** (undoing a speculative JIT optimization), revoking biased locks, taking a thread dump, or applying certain `Thread` operations. To reach a safepoint the JVM asks all threads to pause at the next safe location — this is **time-to-safepoint** plus the **safepoint operation** itself. STW GC *uses* safepoints, but safepoints also happen for non-GC reasons. ## Why this distorts a microbenchmark Your benchmark wants to measure the cost of *one operation*. A GC or safepoint pause can land **in the middle of a timed iteration**, adding, say, 3 ms to an operation that really costs 50 ns. Now: - **If you report the mean**, a handful of multi-millisecond pauses across millions of iterations can dominate or visibly inflate the average — your "50 ns" operation reports as 200 ns for reasons unrelated to the code. - **If you naively discard outliers**, you might *also* be discarding the operation's *legitimate* allocation cost. If the operation itself allocates, the GC it triggers is part of its true cost — throwing it away under-reports reality. So both directions of naive handling lie: include the noise and you over-report; strip it blindly and you under-report. ## Coordinated omission lurks here too If the benchmark loop *waits* for each iteration to finish before starting the next, a long GC pause stalls the whole loop — but you only record *one* slow sample, when in a real system many requests would have queued up behind that pause and all suffered. This is **coordinated omission**: the pause's true impact on latency percentiles is massively under-counted. (It's a distinct pitfall, but GC pauses are a classic trigger.) ## How to account for it correctly 1. **Control allocation.** Minimize *incidental* allocation (autoboxing, temporary objects in the harness) so the GC you observe reflects the *operation's* real allocation, not benchmark scaffolding. 2. **Measure allocation explicitly.** Use a profiler (JMH `-prof gc`) to report allocation rate (bytes/op). An operation's allocation footprint is often more stable and informative than wall-clock for comparing implementations, because it's deterministic. 3. **Run long and representative.** Time long enough that GC happens at a realistic rate; a too-short run might see *zero* GC and under-report, or *one* and over-report. 4. **Report the distribution, not just the mean.** Percentiles (p50/p99/p999) and max separate "typical operation cost" from "perturbation." A clean p50 with a spiky p999 tells you GC/safepoints hit some iterations. 5. **Reduce safepoint noise.** Avoid triggering deopts mid-measurement (warm up the exact paths), and be aware of safepoint-bias in sampling profilers. ## Why a harness exists JMH forks **fresh JVMs** per run (so prior state doesn't leak in), provides `-prof gc` to attribute allocation, reports **percentiles** out of the box, and structures measurement so a single GC spike is visible rather than silently smeared into a mean. Reproducing all of that in hand-rolled timing is error-prone — which is exactly why you don't. ## The takeaway GC and safepoint pauses are *real* and part of how Java runs — you can't pretend they don't exist. The skill is distinguishing **"this operation's own allocation/GC cost"** (keep it) from **"a stray global pause perturbed this sample"** (account for it via distribution, not by faking a clean average).

  • Why is measuring allocation rate (bytes/op) often more useful than wall-clock time for comparing two implementations?
    Allocation is largely deterministic and not subject to JIT timing noise, frequency scaling, or stray GC pauses, so bytes/op is a stable, reproducible signal. Lower allocation usually means less GC pressure and better scalability under load, even when wall-clock differences are within noise.
  • What is a safepoint that is NOT caused by garbage collection?
    Deoptimization of JIT-compiled code, revoking biased locks, taking a thread/heap dump, certain Thread.stop/suspend-style operations, and JVMTI/debugger requests all bring threads to a safepoint. These can perturb a benchmark even on a GC that isn't currently collecting.

saying these in an interview costs you the question

  • Assuming modern concurrent collectors (ZGC/Shenandoah) have zero pauses — they have short STW phases and still hit safepoints.
  • Reporting only the mean and treating it as 'the' operation cost.
  • Blindly deleting all outliers, which can erase the operation's real allocation/GC cost.
  • Equating safepoints with GC — safepoints also serve deopt, lock revocation, dumps.

context