How can the JIT compiler make a Java microbenchmark report a near-zero time for work that genuinely runs, and how do you prevent it?
answer
- Compiler preserves observable behavior, not your timing intent
- Unused result -> dead-code elimination -> empty loop
- Constant input -> constant folding / loop-invariant hoisting
- Fix: consume result (Blackhole/return) + non-constant @State input
- Tell: sub-nanosecond time or identical impls = work deleted
basics
~20 sIf the result of your code is never used, the JIT can delete the whole computation as dead code. If inputs are constants, it can compute the answer at compile time. Either way the loop disappears and you measure nothing. Fix it by consuming the result (e.g. return it / feed it to a sink) and using non-constant inputs.
solid answer
~50 sAn optimizing compiler (C2) only has to preserve observable behavior. If a benchmark computes a value but never uses it, that's dead code and gets eliminated — the loop body vanishes and you time an empty loop (often 'nanoseconds' for arbitrary work). If the inputs are compile-time constants, constant folding evaluates the expression once at compile time, or loop-invariant code is hoisted out of the loop. Both make the benchmark lie. To prevent it: (1) consume every result so it becomes observable — return it from the benchmark method or pass it to a sink that the compiler can't see through; in JMH that's the Blackhole or returning the value; (2) make inputs non-constant — read them from @State fields the JIT must treat as runtime values, not literals; (3) avoid accidentally creating loop-invariant work. This is the single biggest reason hand-rolled benchmarks over-report speed, and the core reason JMH provides Blackhole and @State.
code
java · 17 lines// WRONG: result discarded -> dead-code elimination deletes the work,
// and the constant 42 is folded/hoisted.
long start = System.nanoTime();
for (int i = 0; i < N; i++) {
expensiveHash(42); // return value ignored, input constant
}
long tookNs = System.nanoTime() - start; // ~0, meaningless
// RIGHT (JMH): input comes from @State, result is consumed by Blackhole.
@State(Scope.Thread)
public static class St { int x = ThreadLocalRandom.current().nextInt(); }
@Benchmark
public int hash(St st) {
return expensiveHash(st.x); // returned value is consumed by the harness
}
// or: public void hash(St st, Blackhole bh) { bh.consume(expensiveHash(st.x)); }go deeper
Understands that if you don't use the result, the computation can be deleted, so 'use the result somehow' and don't feed it constants.
Names dead-code elimination and constant folding, explains why an unused local is still dead, and uses Blackhole/return + @State inputs to fix it.
Adds loop-invariant code motion, dead-store elimination of the sink, and recognizes tells (sub-ns times, identical impls); reasons about what 'observable behavior' the compiler must keep.
Can explain how Blackhole is engineered to be uneliminable yet near-zero-cost, the limits of @State in defeating folding, and why outsourcing this to a harness is the only reliable approach at scale.
## The compiler's contract An optimizing compiler is allowed to transform your program in any way that does not change its **observable behavior** — the outputs, side effects, and exceptions a program is *defined* to produce. It is *not* obligated to actually run computations whose results no one observes. In a real program this is a feature. In a benchmark it's a trap, because the *only* thing you care about — how long the work takes — is **not** observable behavior the compiler must preserve. ## Pitfall 1: Dead-Code Elimination (DCE) **Dead code** is code whose results are never used. Consider: ```java for (int i = 0; i < N; i++) { Math.sqrt(i); // result thrown away } ``` The return value of `Math.sqrt` is discarded, and `Math.sqrt` has no side effects. The C2 JIT proves the loop body is observably useless and **deletes it entirely** — you end up timing an *empty* loop, then possibly deleting that too. The benchmark reports a near-zero time for an operation that, run for real, costs many nanoseconds. The lie is in the *direction that flatters you*: it says the code is infinitely fast. ## Pitfall 2: Constant Folding **Constant folding** evaluates expressions whose inputs are known at compile time. If you write: ```java int x = 42; for (int i = 0; i < N; i++) { result = expensiveHash(x); // x is always 42 } ``` The compiler can compute `expensiveHash(42)` *once* (or even at compile time) and reuse it, because the input never varies — this is **loop-invariant code motion** (hoisting work that doesn't depend on the loop variable out of the loop). Again: the per-iteration cost you wanted to measure evaporates. ## Why these are so dangerous Both optimizations are *correct* and *desirable* in production — that's the point. They only "lie" because the benchmark's intent (measure the cost of doing the work N times) isn't expressed in a way the compiler must respect. And they're invisible: the code looks like it's doing work, the benchmark runs, it prints a number. The number is just meaningless. ## The fixes **1. Make the result observable (defeat DCE).** Ensure each computed value is *used* in a way the compiler can't prove is useless: - Return it from the benchmark method (the harness consumes returned values). - Feed it to a **sink** the compiler can't see through. JMH provides `Blackhole.consume(value)`, which is engineered so the JIT cannot eliminate it but it also doesn't add meaningful overhead or dead-store-eliminate. - Accumulate into a field that is later read. **2. Make inputs non-constant (defeat constant folding / hoisting).** Read inputs from state the JIT must treat as runtime-variable — in JMH, `@State` object fields. Don't hardcode literals into the hot expression; the compiler treats fields loaded each iteration as unknown values, so it can't fold them. **3. Watch for accidental loop-invariance.** If the work genuinely doesn't depend on the loop index, the compiler will hoist it. Vary the input per iteration (or per invocation) so each iteration is real work. ## Why a harness exists Getting a sink that the JIT can't optimize through — *without* the sink itself dominating the measurement or being dead-store-eliminated — is genuinely hard; it took the JMH authors deep JIT knowledge to build `Blackhole`. Likewise, `@State` exists precisely so inputs aren't seen as constants. This is the central reason you use JMH instead of hand-rolled timing: it makes "the work actually happened and the compiler couldn't cheat" the default. ## A concrete tell If a benchmark reports that an obviously non-trivial operation takes a fraction of a nanosecond, or that two clearly different implementations have *identical* times, suspect DCE or constant folding — the compiler likely deleted the work.
- Why can't you just assign the result to a local variable to stop DCE?A local that is never read afterward is itself dead — the store gets dead-store-eliminated and the computation feeding it is removed. The result must reach something observable: a returned value the harness consumes, a Blackhole, or a field that is genuinely read later.
- How does JMH's Blackhole avoid both deleting the value AND distorting the measurement?It's carefully written so the JIT can't prove the consumed value is unused (defeating DCE) while keeping its own per-call cost tiny and constant, and it avoids dead-store elimination of the sink itself. It uses tricks like conditionally touching a volatile/state so the compiler must keep the value live without adding real arithmetic to your hot path.
saying these in an interview costs you the question
- Believing assigning to an unused local variable prevents elimination — it doesn't; the store is dead too.
- Thinking the JIT 'shouldn't' delete running code — it's allowed to, since timing isn't observable behavior.
- Hardcoding literal inputs into the measured expression, enabling constant folding.
- Trusting a benchmark that reports sub-nanosecond times for real work.