skip to content

JMH Warmup & Iterations

Warmup iterations exist because the JIT needs time to reach steady state through C1, C2 and on-stack replacement; measurement iterations are the ones reported. Interviewers ask what warmup is for and how to read the score plus error interval.

part ofJavaoverview, primer and where to startread it →
on this pageshow

questions

5

What is the difference between @Warmup and @Measurement iterations in JMH, and how do you configure their count and time?

level: juniorimportance: must knowfreq 60%

answer

  1. Iteration = a time window, not one call
  2. Warmup discarded; measurement reported
  3. Knobs: iterations + time/timeUnit
  4. CLI: -wi/-i (counts), -w/-r (time)
  5. All of it repeats per fork

basics

~20 s

@Warmup runs the code to get the JVM up to speed and throws those numbers away. @Measurement runs it again and keeps those numbers. You set how many rounds (iterations) and how long each round lasts (time).

solid answer

~40 s

Both annotations describe a sequence of iterations, where an 'iteration' is a fixed time window in which JMH repeatedly invokes your @Benchmark method and counts operations. @Warmup iterations are run first and their measurements are discarded — they only exist to push the JVM toward steady state. @Measurement iterations follow and their measurements are what JMH aggregates into the reported score and error. You configure each with `iterations` (how many rounds) and `time`/`timeUnit` (how long each round lasts), e.g. `@Warmup(iterations = 5, time = 2, timeUnit = SECONDS)`. You can set them as class/method annotations or override at the command line (`-wi`, `-i`, `-w`, `-r`). Rule of thumb: enough warmup that per-iteration scores stabilize, and enough measurement iterations to get a tight, trustworthy error interval.

code

java · 12 lines
java
@Warmup(iterations = 5, time = 1, timeUnit = TimeUnit.SECONDS)
@Measurement(iterations = 10, time = 1, timeUnit = TimeUnit.SECONDS)
@Fork(2)
public class Bench {
    @Benchmark
    public long sum() {
        long s = 0;
        for (int i = 0; i < 1000; i++) s += i;
        return s;
    }
}
// CLI override: java -jar benchmarks.jar -wi 8 -i 15 -w 2s -r 2s -f 3

go deeper

for a junior

Can state that warmup numbers are thrown away and measurement numbers are kept, and can set iterations and time on the annotations.

for a middle

Understands an iteration is a time window, knows the CLI overrides, and can reason about how counts affect noise and run duration.

for a senior

Picks warmup/measurement budgets from observed iteration stability and target error width, and accounts for per-fork multiplication.

for a principal

Sets org-wide conventions for benchmark configuration, balances run-time cost vs statistical confidence, and reviews benchmarks for adequate warmup/measurement before trusting conclusions.

## The vocabulary first **JMH** (Java Microbenchmark Harness) measures tiny code snippets. Inside JMH: - A **@Benchmark method** is the code under test. - An **iteration** is *not* one call of that method — it is a **time window** (e.g. 1 second). During that window JMH calls your method as many times as it can and counts how many operations completed (or how long each took, depending on the mode). One iteration produces one data point (e.g. '32,140,000 ops/sec'). - A **run** is a whole set of iterations. ## @Warmup vs @Measurement A benchmark run is split into two ordered phases: 1. **@Warmup** — the first set of iterations. JMH runs them but **discards** their measurements. Their sole purpose is to let the JVM reach **steady state** (hot code JIT-compiled, caches warm). Think of them as practice rounds. 2. **@Measurement** — the second set of iterations. JMH runs them and **records** every data point, then aggregates them into the final reported **score** plus an **error / confidence interval**. The key distinction: *discarded vs reported*. Same machinery, different fate for the numbers. ## Configuring them Each annotation takes the same knobs: - `iterations` — how many rounds in that phase. - `time` + `timeUnit` — how long **each** round runs (e.g. `time = 2, timeUnit = SECONDS` → each iteration is ~2 seconds). - `batchSize` — (less common) how many method calls count as one logical operation, for ops too cheap to time individually. Example: ```java @Warmup(iterations = 5, time = 1, timeUnit = TimeUnit.SECONDS) @Measurement(iterations = 10, time = 1, timeUnit = TimeUnit.SECONDS) ``` That's ~5 seconds of throwaway warmup, then ~10 seconds of recorded measurement, **per fork**. You can place these as annotations on the class (defaults for all benchmarks) or on a method (override for one), and you can override everything at the command line: `-wi` (warmup iterations), `-i` (measurement iterations), `-w` (warmup time), `-r` (measurement time). ## Why both numbers matter - **Too few warmup iterations** → measurement includes JIT compilation cost → slow, noisy, misleading results. - **Too few measurement iterations** → a wide error interval → you can't tell two benchmarks apart. - **Too many of either** → slower benchmark runs with diminishing returns. The practical approach: watch the per-iteration scores JMH prints during warmup. Once they stop trending and just wobble around a stable value, that's your warmup budget. Then use enough measurement iterations (often 5–20) that the reported error is small relative to the score. ## Multiplied by forks Whatever you configure happens **per fork** (a separate JVM process). With `@Fork(3)`, the warmup+measurement cycle repeats in 3 fresh JVMs and JMH aggregates across all of them.

  • Does 'iterations = 5' mean the method is called 5 times?
    No. An iteration is a time window (set by `time`/`timeUnit`). Within each of the 5 iterations, JMH calls the method as many times as fit in that window and counts the operations. So 5 iterations of 1 second each means 5 measured time windows, not 5 method calls.

saying these in an interview costs you the question

  • Thinking 'iterations' counts method invocations rather than time windows.
  • Believing warmup and measurement use different code or different method calls — they run the same @Benchmark, only the data's fate differs.
  • Forgetting that the configured iterations repeat per fork.

context

open as a page

Why does a JMH benchmark need a warmup phase, and what happens to the JVM during it?

level: middleimportance: must knowfreq 70%

basics

~10 s

The JVM speeds up code as it runs by compiling it. Warmup runs the code a while first so it is already fast, so you measure the fast version, not the slow startup.

open as a page

Why does JMH run benchmarks in multiple forks, and what does @Fork control?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Each fork is a fresh JVM. JMH runs your benchmark in several separate JVMs and averages the results, because one JVM can get 'lucky' or 'unlucky' in how it compiles your code. Multiple forks smooth that out.

open as a page

How do you read a JMH result line — the score, the ± error, and the units — and decide whether two benchmarks actually differ?

level: seniorimportance: should knowfreq 50%

basics

~20 s

The score is the average result (e.g. nanoseconds per operation or ops per second). The ± number is the margin of error — how much the score could wobble. If two benchmarks' score-plus-or-minus ranges overlap, you can't claim one is faster.

open as a page

What benchmarking pitfalls does JMH guard against, and how do dead-code elimination, constant folding, and Blackhole/@State relate to warmup-stage validity?

level: principalimportance: should knowfreq 40%

basics

~20 s

If your benchmark's result isn't used, the JIT can delete it (dead-code elimination) or precompute it (constant folding) — so you'd measure nothing. JMH fixes this: return values, consume them with Blackhole, and keep inputs in @State so they aren't treated as constants.

open as a page