Why does a JMH benchmark need a warmup phase, and what happens to the JVM during it?
answer
- Interpreted → C1 → C2 (tiered)
- Warmup iterations discarded, measurement recorded
- OSR swaps a running loop mid-flight
- Goal = steady state, not cold start
- Watch per-iteration stability to size warmup
basics
~10 sThe JVM speeds up code as it runs by compiling it. Warmup runs the code a while first so it is already fast, so you measure the fast version, not the slow startup.
solid answer
~40 sThe HotSpot JVM starts by interpreting bytecode, then the JIT (Just-In-Time) compiler converts hot methods to native machine code. There are two compilers: C1 (fast to compile, less optimized) and C2 (slower, heavily optimized). A method typically runs interpreted, then C1, then C2 as invocation/loop counters cross thresholds. Until that settles, timings are dominated by interpretation and compilation work, not the code you care about. JMH's @Warmup runs many iterations whose results are thrown away, giving the JVM time to reach 'steady state' where the hot path is fully C2-compiled. Only then does @Measurement begin, so the reported score reflects the optimized code. Skipping warmup yields numbers that are slower and noisier, and can mislead you into 'optimizing' code whose real cost was just the JIT warming up.
code
java · 12 lines@BenchmarkMode(Mode.AverageTime)
@OutputTimeUnit(TimeUnit.NANOSECONDS)
@Warmup(iterations = 5, time = 1, timeUnit = TimeUnit.SECONDS) // discarded
@Measurement(iterations = 5, time = 1, timeUnit = TimeUnit.SECONDS) // reported
@Fork(2)
@State(Scope.Thread)
public class HashBenchmark {
@Benchmark
public int hash(Blackhole bh) {
return "steady-state".hashCode();
}
}go deeper
Knows that Java code gets faster as it runs and that warmup runs the code first so you measure the fast version.
Explains tiered JIT (interpreter → C1 → C2), that warmup iterations are discarded vs measurement recorded, and that the goal is steady state.
Adds OSR, sizing warmup by observing iteration stability, and recognizes when cold-start (no warmup) is the correct thing to measure.
Reasons about how warmup choices change what conclusion the benchmark supports, deopt/reopt churn, profile-guided C2 decisions, and aligning the benchmark regime (steady vs cold) with the real production workload.
## What a benchmark is trying to measure A microbenchmark measures how long a small piece of code takes to run. **JMH** (Java Microbenchmark Harness) is the standard tool for this on the JVM. The hard part is that the JVM does not run your code at a constant speed — it gets *faster over time* — so a naive 'start a timer, run the code, stop the timer' gives a misleading number. ## Why the JVM gets faster: interpretation and JIT compilation Java source compiles to **bytecode**, a portable instruction set. When the program starts, the JVM **interprets** that bytecode: it reads each instruction and executes it, which is flexible but slow. As the program runs, the JVM watches which methods and loops run a lot ('hot' code). The **JIT (Just-In-Time) compiler** then translates that hot bytecode into **native machine code** — instructions the CPU runs directly — which is far faster. HotSpot (the standard JVM) has **two JIT compilers** arranged in 'tiered compilation': - **C1** (the 'client' compiler): compiles quickly but applies fewer optimizations. It also inserts profiling counters. - **C2** (the 'server' compiler): compiles slowly but produces highly optimized code, using the profile gathered earlier (e.g. which branches are usually taken, which types actually occur). A typical method's life: **interpreted → C1-compiled → C2-compiled**, as invocation counters and loop back-edge counters cross internal thresholds. So the *same* method can run at three very different speeds during one program. ## On-stack replacement (OSR) Normally a newly compiled method is only used the *next* time it is called. But a benchmark often has a long-running loop that is already executing when the loop becomes hot. **On-stack replacement (OSR)** lets the JVM swap the running interpreted loop for a compiled version *mid-flight*, without waiting for the method to be re-entered. OSR-compiled code can have slightly different performance characteristics than a normally-compiled method, which is one more reason early loop iterations are unrepresentative. ## Steady state '**Steady state**' is the point where all this has settled: the hot path is fully C2-compiled, the code caches are populated, the CPU branch predictors and caches are warm, and timings stop drifting. That is the regime you almost always want to measure, because it is what a long-running server actually experiences. ## What warmup does in JMH JMH splits a run into two phases controlled by annotations (or command-line flags): - **@Warmup** — runs N iterations whose results are **discarded**. Their only job is to push the JVM to steady state. - **@Measurement** — runs N iterations whose results are **recorded and reported**. Example: `@Warmup(iterations = 5, time = 1, timeUnit = SECONDS)` means 5 warmup iterations of ~1 second each (≈5s of throwaway work), and `@Measurement(iterations = 5, time = 1)` records the next 5. If you skip or under-size warmup, the first measured iterations include interpretation + compilation cost. Symptoms: the first iteration is dramatically slower than later ones, high variance, and you might 'optimize' code that was only slow because the JIT hadn't kicked in yet. ## How much warmup is enough There's no universal number. The signal is *stability*: watch the per-iteration scores JMH prints; once consecutive warmup iterations stop trending and start oscillating around a stable value, you've reached steady state. Code with large method bodies, lots of branches, or deopt/reopt churn needs more warmup. For some workloads (cold-start, lambda functions, CLI tools) the *cold*, un-warmed-up behavior is actually what you care about — there warmup is the wrong default and you'd measure single-shot instead. ## Putting it together for an answer Warmup exists because JVM performance is *time-dependent*: interpreted → C1 → C2, with OSR for hot loops. Discarded warmup iterations drive the code to steady state so the recorded @Measurement iterations reflect fully optimized native code rather than one-off compilation cost.
- What is on-stack replacement (OSR) and why does it matter for benchmarks?OSR lets the JVM replace a currently-running interpreted loop with compiled code mid-execution, instead of waiting for the method to be called again. It matters because benchmark loops are often already running when they get hot; OSR-compiled code can differ from normally-compiled code, so early loop iterations are unrepresentative.
- When would you deliberately NOT warm up?When cold-start behavior is what you care about — e.g. CLI tools, serverless/Lambda functions, or class-loading/first-call latency. There you'd use single-shot mode (Mode.SingleShotTime) with no warmup, because steady state never happens in production.
Like a sprinter: the JVM jogs (interprets), then warms up (C1), then runs at race pace (C2). You time the race, not the stretching.
saying these in an interview costs you the question
- Saying warmup 'lets the cache fill' as the only reason — JIT compilation (interpreted→C1→C2) is the main driver, not just CPU/data caches.
- Claiming warmup iterations are averaged into the result — they are discarded.
- Believing one warmup iteration is always enough — steady state depends on the code and must be observed.
- Confusing warmup with JVM startup; warmup happens per-benchmark within an already-running JVM.