skip to content

What benchmarking pitfalls does JMH guard against, and how do dead-code elimination, constant folding, and Blackhole/@State relate to warmup-stage validity?

level: principalimportance: should knowfreq 40%

answer

  1. Warmup → C2 → it can delete/precompute your work
  2. DCE: return value or Blackhole.consume
  3. Constant folding: @State non-final fields, not constants
  4. @State scope controls sharing & false sharing
  5. Stable number can still measure an empty loop

basics

~20 s

If your benchmark's result isn't used, the JIT can delete it (dead-code elimination) or precompute it (constant folding) — so you'd measure nothing. JMH fixes this: return values, consume them with Blackhole, and keep inputs in @State so they aren't treated as constants.

solid answer

~50 s

Once warmup pushes the JVM to steady state, the C2 compiler aggressively optimizes — and a naive benchmark can be optimized into nothing. Two classic traps: **dead-code elimination (DCE)**, where a computed result that's never used is deleted, so you time an empty loop; and **constant folding**, where the JIT precomputes an expression whose inputs are compile-time constants, so the work happens once at compile time, not per call. JMH's defenses: **return the result** from the @Benchmark method (JMH implicitly sinks returns into a Blackhole), or pass intermediates to a **Blackhole** to mark them as 'used'; keep inputs in **@State** objects and read them as instance fields so the JIT treats them as opaque, not constants. There's also **loop-unrolling** distortion (don't hand-roll loops inside the method) and false sharing (`@State` scope choices). The subtlety: these distortions appear *after* warmup, so a benchmark that 'warmed up fine' can still measure garbage — warmup makes the optimizer *more* dangerous, not less.

code

java · 21 lines
java
@State(Scope.Thread)
public class SafeBench {
    // Non-final field in @State -> defeats constant folding
    private double x = 42.0;

    // Returning the result -> defeats dead-code elimination (JMH sinks it)
    @Benchmark
    public double goodSingle() {
        return Math.sqrt(x) + Math.log(x);
    }

    // Multiple results -> consume each via Blackhole
    @Benchmark
    public void goodMulti(Blackhole bh) {
        bh.consume(Math.sqrt(x));
        bh.consume(Math.log(x));
    }

    // BAD: constant input folds; result unused -> measures nothing
    // @Benchmark public void bad() { Math.sqrt(64.0); }
}

go deeper

for a junior

Knows you must use the result (return it or consume it) or the JIT might delete the code.

for a middle

Can name dead-code elimination and constant folding and apply the basic fixes (return value, Blackhole, @State inputs).

for a senior

Chooses the right @State scope, avoids manual loops, and uses Blackhole for multi-result benchmarks; understands these issues arise at steady state.

for a principal

Reasons about inlining/monomorphic-profile illusions and false sharing, verifies generated code with profilers (perfasm/gc), and reviews whether the warmup profile matches production before trusting any conclusion.

## Why warmup makes optimization both necessary and dangerous Warmup exists to reach **steady state**, where the C2 JIT has fully optimized the hot path. But 'fully optimized' cuts both ways: the very optimizations you *want* (so you measure realistic native code) are the same ones that can **erase your benchmark** if you wrote it naively. So the more thoroughly you warm up, the more important it is that your benchmark is written to survive optimization. JMH provides specific countermeasures. ## Pitfall 1 — Dead-Code Elimination (DCE) The JIT removes computations whose results are never observed (have no side effects). If your benchmark computes `Math.log(x)` and discards it, C2 sees the result is unused and **deletes the call** — after warmup you're timing an empty method. **Defenses:** - **Return the value.** A value returned from a `@Benchmark` method is implicitly consumed by JMH (it sinks it into a Blackhole), so the JIT must keep the computation. - **Blackhole.** Add a `Blackhole` parameter and call `bh.consume(value)`. A Blackhole is a JMH object engineered to 'use' a value in a way the JIT cannot see through and cannot eliminate, without itself costing meaningful time. Use it when you produce *multiple* results per invocation (you can only return one). ## Pitfall 2 — Constant Folding If all inputs to an expression are **compile-time constants**, the JIT computes the answer **once, at compile time**, and just returns the precomputed value forever. `"hello".length()` folds to `5`; `Math.sqrt(64)` folds to `8.0`. Your loop then measures returning a constant, not the computation. **Defense:** put inputs in a **`@State` object** and read them as **non-final instance fields**. Because the JIT can't prove an instance field is constant across calls, it must actually perform the computation each time. (Marking the field `final` can re-enable folding — avoid that for benchmark inputs.) ## @State — what it is and why it matters here `@State` marks a class whose instance holds the benchmark's mutable inputs/fixtures. Its **scope** controls sharing: - `Scope.Thread` — each thread gets its own instance (default-safe; avoids contention and false sharing). - `Scope.Benchmark` — one shared instance across all threads (use to measure contention deliberately). - `Scope.Group` — per thread-group, for `@Group` asymmetric benchmarks. Beyond folding-prevention, the scope choice affects cache effects like **false sharing** (two threads writing adjacent fields on the same cache line, ping-ponging it). Choosing the wrong scope can make a single-threaded benchmark look artificially slow or a contended one artificially fast. ## Pitfall 3 — Loop optimizations inside the method If you hand-write a loop inside the @Benchmark method to 'do more work per call', the JIT may **unroll** it, hoist invariants, or vectorize it in ways that don't reflect the real per-call cost — and it can even fold/eliminate the whole loop. Rule: let JMH drive the repetition (iterations/time); keep the @Benchmark body the single unit of work. For very cheap operations, use `@OperationsPerInvocation` or `batchSize` rather than a manual loop. ## Pitfall 4 — Inlining / monomorphic profile illusions Warmup's profiling can make a benchmark's call site **monomorphic** (only ever one concrete type), letting C2 inline and devirtualize perfectly — which may *not* be true in production where the call site is megamorphic. A benchmark can therefore over-state performance simply because its warmup taught the JIT an unrealistically clean profile. Mitigation: feed representative type/branch mixes during warmup. ## The deeper point for a principal All of these distortions surface **after warmup**, at steady state — exactly the regime you intended to measure. So 'the benchmark warmed up and stabilized' is necessary but **not sufficient** for validity. A stable, low-error, beautifully warmed-up number can still be measuring an empty loop (DCE), a constant (folding), or an unrealistically optimistic inlining profile. The reviewer's job is to confirm the benchmark *survives* the optimizer: results are sunk (return/Blackhole), inputs are opaque (@State, non-final), repetition is JMH-driven, and the warmup profile matches production. Tools like `-prof gc`, `-prof perfasm` (inspect the generated assembly), and `-prof dtlab` help verify that real work is actually happening. ## Interview-ready summary Warmup gets you to steady state, but at steady state C2 will delete unused results (DCE) and precompute constant inputs (constant folding). JMH counters with returned values / Blackhole (defeat DCE) and `@State` non-final fields (defeat folding), plus correct `@State` scope to control sharing/false-sharing and JMH-driven repetition instead of manual loops. Validity = the benchmark survives the very optimizations warmup enables.

  • Why does putting the input in a @State field (instead of a literal) prevent constant folding?
    The JIT can only fold expressions whose inputs it can prove are compile-time constants. A non-final instance field read at runtime is opaque to the compiler — it can't assume the value, so it must perform the computation every invocation. A literal or final-constant input, by contrast, lets C2 precompute the result once.
  • A benchmark warmed up cleanly and shows a tiny error. Why might it still be invalid?
    Because the distortions that invalidate a benchmark — dead-code elimination, constant folding, unrealistically clean inlining/monomorphic profiles — all happen *at steady state*, after warmup. A stable, low-error number can be precisely measuring an empty loop or a constant. You must verify the benchmark survives optimization (sunk results, opaque inputs, representative profile), e.g. via -prof perfasm.

saying these in an interview costs you the question

  • Believing a clean warmup guarantees a valid benchmark — the worst distortions appear at steady state.
  • Computing a value and not returning/consuming it (invites dead-code elimination).
  • Using literal or final-constant inputs (invites constant folding).
  • Hand-rolling loops inside the @Benchmark body and trusting the per-call cost.
  • Ignoring @State scope, leading to false sharing or unintended contention.

context