skip to content

JMH Setup & Benchmark Lifecycle

JMH fundamentals: @Benchmark methods, choosing a Mode, scoping @State with @Setup and @TearDown, and why each trial forks a fresh JVM to avoid profile pollution. Knowing that JMH exists and why is the first thing interviewers check about benchmarking.

part ofJavaoverview, primer and where to startread it →
on this pageshow

questions

5

What is JMH, and why should you use it instead of hand-rolled timing with System.nanoTime() loops for Java microbenchmarks?

level: juniorimportance: must knowfreq 55%

answer

  1. Official OpenJDK harness, same team as the JVM
  2. @Benchmark + annotation processor generates the harness
  3. Defends against: warmup/JIT, dead-code elimination, constant folding, OSR
  4. Forks a fresh JVM per trial
  5. Naive nanoTime loops measure the wrong thing

basics

~20 s

JMH is the Java Microbenchmark Harness, an official tool for measuring how fast small pieces of Java code run. It's better than manual nanoTime loops because the JVM warms up and optimizes code over time, and JMH handles that warmup and the JIT effects so your numbers are trustworthy.

solid answer

~50 s

JMH (Java Microbenchmark Harness) is the OpenJDK-blessed framework for writing reliable microbenchmarks on the JVM. Hand-rolled nanoTime loops are notoriously wrong because the JVM is a dynamic, adaptive runtime: the JIT compiler only optimizes hot code after thousands of iterations (warmup), it can dead-code-eliminate work whose result is unused, constant-fold inputs the optimizer can prove are fixed, and on-stack-replace a running loop mid-measurement. JMH addresses all of these: it runs warmup iterations before measuring, it forks a fresh JVM per trial to avoid profile pollution between benchmarks, it provides Blackhole and return-value consumption to stop dead-code elimination, and it reports statistics across many iterations and forks. You annotate a method with @Benchmark and JMH's annotation processor generates the harness. The result is numbers you can actually trust instead of measuring the JIT's ability to delete your code.

code

java · 18 lines
java
import org.openjdk.jmh.annotations.*;
import org.openjdk.jmh.infra.Blackhole;

public class WhyJmh {

    // WRONG: a naive loop -- the JIT may dead-code-eliminate or constant-fold this.
    // long t = System.nanoTime();
    // for (int i = 0; i < 1_000_000; i++) Math.log(i); // result thrown away
    // long ns = System.nanoTime() - t;                 // meaningless number

    // RIGHT: JMH consumes the result so the work cannot be deleted.
    @Benchmark
    public double log(Blackhole bh) {
        double sum = 0;
        for (int i = 1; i < 1_000; i++) sum += Math.log(i);
        return sum; // returned -> consumed by JMH
    }
}

go deeper

for a junior

Knows JMH is the standard Java benchmarking tool and that manual nanoTime loops are unreliable because the JVM warms up.

for a middle

Can name the specific distortions (warmup/JIT, dead-code elimination, constant folding) and how JMH counters each.

for a senior

Explains forking, Blackhole/return-value consumption, and warmup vs measurement iterations, and articulates when JMH is the wrong tool (end-to-end load).

for a principal

Frames microbenchmarking results in context — knows micro numbers rarely predict whole-system performance, insists on profiling the real workload, and treats JMH as one input among allocation/GC/cache effects.

## What a microbenchmark is A *microbenchmark* measures the speed of a very small piece of code — a single method, a loop body, one operation — rather than a whole program. The goal is to answer questions like "is approach A faster than approach B for this one operation?" ## Why this is hard on the JVM The **JVM (Java Virtual Machine)** is not a simple interpreter; it is a dynamic, adaptive runtime. Several of its behaviors make naive timing wrong: - **JIT compilation & warmup.** Java code first runs *interpreted* (slow), and only after a method or loop becomes "hot" (executed thousands of times) does the **JIT (Just-In-Time) compiler** compile it to optimized machine code. The first measurements are therefore measuring the interpreter, not the optimized code you care about. The period of running until performance stabilizes is called **warmup**. - **Dead-code elimination (DCE).** The optimizer removes computations whose results are never used. If your benchmark computes `Math.log(x)` and throws the result away, the JIT may delete the call entirely — you'd measure nothing. - **Constant folding.** If an input is a compile-time or provably-constant value, the optimizer can precompute the answer once, so your loop measures nothing. - **On-stack replacement (OSR).** The JVM can swap a still-running interpreted loop for a compiled version mid-flight, producing measurements from artificially-shaped code that never occurs in real call sites. - **GC and background compilation noise** add variance. A hand-written `long t = System.nanoTime(); for (...) { work(); } long elapsed = System.nanoTime() - t;` ignores all of this and routinely produces numbers that are off by orders of magnitude or measure the wrong thing. ## What JMH is **JMH = Java Microbenchmark Harness.** It is the benchmarking framework maintained by the same OpenJDK engineers who build the JVM, precisely because they understand these pitfalls. You write a benchmark by annotating a method with **`@Benchmark`**; JMH's **annotation processor** generates boilerplate "harness" code at build time that runs your method correctly. ## How JMH defends correctness - **Warmup iterations** run your code (results discarded) until the JIT has compiled and stabilized it; only then does it run **measurement iterations**. - **Forking:** each trial runs in a *fresh JVM process* so one benchmark's compilation profile can't pollute another's. - **Blackhole / return-value consumption:** returning a value from `@Benchmark` (or passing it to a `Blackhole`) tells JMH to *consume* it so the JIT cannot dead-code-eliminate the work. - **State objects** (`@State`) hold inputs in a way the optimizer can't constant-fold. - **Statistics:** JMH aggregates many iterations across many forks and reports a score with error/confidence, not a single lucky number. ## When NOT to use it JMH is for *micro* benchmarks (sub-millisecond to millisecond operations). For end-to-end/application throughput you'd use load-testing tools instead. But for "which of these two implementations of a hot method is faster," JMH is the standard answer.

  • Name two ways the JIT can make a naive benchmark report a misleadingly fast result.
    Dead-code elimination (it deletes work whose result is unused) and constant folding (it precomputes results for provably-constant inputs), so the loop ends up measuring nothing.
  • Who maintains JMH and why does that matter?
    The OpenJDK / JVM engineers maintain it; it matters because they know exactly which JIT and runtime behaviors corrupt naive benchmarks and bake the defenses into the harness.

saying these in an interview costs you the question

  • Claiming a simple nanoTime loop is 'good enough' for micro-level comparisons
  • Not knowing the JVM has a warmup/JIT phase that distorts early measurements
  • Thinking JMH is for end-to-end load testing rather than micro-level method timing

context

open as a page

What are JMH's benchmark Modes (Throughput, AverageTime, SampleTime, SingleShotTime), and how does @OutputTimeUnit relate to them?

level: middleimportance: should knowfreq 45%

basics

~20 s

A Mode tells JMH how to express the result. Throughput counts operations per unit of time (higher is better); AverageTime is time per operation (lower is better); SampleTime samples individual call times to build a distribution including percentiles; SingleShotTime measures one run with no warmup, for cold-start cost. @OutputTimeUnit picks the unit (ms, us, ns) the score is printed in.

open as a page

Walk through setting up and running a minimal JMH benchmark project: dependency/annotation-processor setup, the generated harness, and how a benchmark gets executed.

level: middleimportance: should knowfreq 40%

basics

~20 s

Add the JMH core library plus its annotation processor as dependencies. Write a method annotated with @Benchmark. At build time the annotation processor generates harness code and a runnable Main. Then you run the benchmarks by executing that generated jar (or via a Runner in code), and JMH does warmup, measurement, and forking automatically.

open as a page

Why does JMH fork a fresh JVM for each trial, and what is 'profile pollution' that forking prevents?

level: seniorimportance: should knowfreq 38%

basics

~20 s

JMH runs each benchmark in a brand-new JVM process (a fork). This is because the JVM remembers and adapts based on what it ran before, so running two benchmarks in the same JVM lets the first one's optimization decisions affect the second's results. A fresh JVM per trial keeps each benchmark's measurement honest and independent.

open as a page

Explain @State scopes (Benchmark, Thread, Group) and the @Setup/@TearDown lifecycle with their Levels in JMH.

level: seniorimportance: should knowfreq 42%

basics

~20 s

A @State class holds the data your benchmark uses. Its scope says who shares one instance: Benchmark = all threads share one, Thread = each thread gets its own, Group = one per group of cooperating threads. @Setup methods prepare state before the benchmark and @TearDown cleans up after; each can run at Trial, Iteration, or Invocation level depending on how often you need it.

open as a page