You inherit a JMH benchmark reporting 0.3 ns/op for a non-trivial computation. How do you systematically audit whether the JIT optimized the work away?
answer
- Sub-ns for real work = optimized away
- Audit 3 traps: folding, DCE, loops
- Compare to an empty/constant baseline
- perfasm / PrintAssembly = ground truth
- Re-run after fixes; expect plausible magnitude
basics
~20 sA sub-nanosecond result for real work is a red flag. Check that inputs come from non-final @State fields (not constants), that every output is returned or Blackhole-consumed, and that any loop varies its data. Then re-run, ideally inspecting the generated assembly to confirm the work runs.
solid answer
~50 s0.3 ns/op is below the cost of most real operations, so I'd treat the number as 'optimized away until proven otherwise'. I audit the three classic traps. Constant folding: are inputs literals or static final/final-constant fields? They must come from mutable @State instance fields so the JIT can't precompute them. Dead-code elimination: is every result returned or passed to Blackhole.consume? An unused result gets deleted. Loop effects: if there's a hand loop, are iterations identical (foldable) or unrolled, and is each result consumed? Beyond reading code, I confirm empirically: compare against a known baseline (an empty benchmark should be near the same number if work was eliminated), check JMH's own 'benchmark may have been optimized away' warning, and for certainty use -prof perfasm or print-assembly to see whether the computation's instructions actually appear in the compiled code. If they're absent, the JIT removed them. I also re-validate after fixes by confirming the score moves to a plausible magnitude.
go deeper
Recognizes that a sub-nanosecond result for real work is suspicious and knows the basic fixes (return/Blackhole, @State inputs).
Audits the three traps in code and compares against a baseline to catch elimination.
Adds empirical verification — JMH warnings, magnitude sanity checks — and reasons about which optimization is responsible before fixing.
Uses assembly/profiler inspection as ground truth, understands environment-dependence of JIT behavior, re-validates after fixes, and institutes team-wide benchmarking standards so performance decisions rest on trustworthy data.
## Why 0.3 ns/op is suspicious On modern hardware a single nanosecond is a few cycles. Sub-nanosecond per-op times are plausible only for the very cheapest operations (a trivial add). For anything non-trivial — a hash, a `Math` call, an allocation, a map lookup — a sub-nanosecond result almost always means the **JIT** (the JVM's runtime optimizing compiler) deleted or precomputed the work. Treat the number as guilty until proven innocent. ## The three traps to rule out **1. Constant folding (inputs).** If the inputs are compile-time constants — literals, or `static final` / final-with-constant fields — the JIT computes the expression once at compile time and the benchmark returns a precomputed value. *Fix/verify:* inputs must be read from **mutable instance fields of a `@State` object** so the value is a heap load the compiler cannot prove constant. **2. Dead-code elimination (outputs).** A result that is never observed is dead code and gets deleted. *Fix/verify:* every result must be **returned** (JMH consumes the return) or passed to **`Blackhole.consume()`**. Discarded locals, fields written but never read, or only-the-last-iteration-kept loops all leak work to DCE. **3. Loop effects.** A manual loop can be unrolled, have invariant work hoisted out, or — if iterations are identical — folded to one. *Fix/verify:* prefer no manual loop (JMH repeats for you); if a loop is needed, vary data from a `@State` array, consume each result, and use `@OperationsPerInvocation` to normalize. ## A systematic audit procedure 1. **Read the code against the three traps.** Where do inputs come from? Are they constant? Where do outputs go? Are they consumed? Is there a loop, and does it vary data and consume each result? 2. **Baseline comparison.** Add or run an empty/`return constant` benchmark. If the suspect benchmark matches that baseline, the real work was eliminated. 3. **Read JMH's warnings.** JMH prints heuristic warnings like 'the benchmark may have been optimized away' or about constant inputs — heed them. 4. **Sanity-check magnitude.** Estimate the expected cost from first principles (cache misses, allocations, syscalls). If the measured number is orders of magnitude smaller, suspect elimination. 5. **Inspect the generated assembly (the definitive check).** Run with `-prof perfasm` (Linux) or `-XX:+PrintAssembly` / the JMH `-prof dtraceasm`/`perfasm` profilers to see the compiled native code. If the instructions implementing the computation are absent, the JIT removed them; if present and hot, the measurement is real. 6. **Re-validate after fixes.** Apply the @State-input and consume-output corrections, re-run, and confirm the score moves to a physically plausible magnitude and that the assembly now contains the work. ## Why code review alone is not enough The JIT's behavior depends on inlining decisions, JVM version, and whether compiler blackholes are available, so a benchmark that *looks* correct can still be partly optimized in a specific environment. The assembly/profiler check is the ground truth: it tells you what actually executed, independent of your assumptions. A principal-level engineer treats benchmark numbers as claims that must be falsified, not facts. ## Governance angle Beyond the single benchmark, the takeaway is process: require @State inputs, consumed outputs, plausibility checks against a baseline, and occasional assembly verification as standard practice, so the whole team's performance numbers are trustworthy and decisions (which is faster, did this change help) rest on real measurements.
- The benchmark already returns its result and reads input from a @State field, yet it's still 0.3 ns/op. What next?Check the @State field isn't effectively constant (static final / final-with-constant), look for a hand loop folding identical iterations, then read the compiled assembly with perfasm/PrintAssembly. If the computation's instructions are missing, something upstream (e.g. an inlined pure call on a value the JIT proved constant) was still folded; if present, the operation may simply be that cheap.
- Why is comparing against an empty-benchmark baseline informative?JMH's harness has irreducible per-invocation overhead. If your 'real work' benchmark reports the same time as an empty one, the work contributed nothing measurable — strong evidence it was eliminated rather than genuinely free.
saying these in an interview costs you the question
- Trusting a sub-nanosecond result for non-trivial work without investigation
- Believing code review alone proves a benchmark is correct, without checking assembly or a baseline
- Disabling the JIT to 'fix' it, producing numbers unrepresentative of production
- Ignoring JMH's own optimized-away warnings