skip to content

Why must a Java microbenchmark include a warmup phase, and what happens if you measure the very first iterations of a hot loop?

level: juniorimportance: must knowfreq 70%

answer

  1. Interpreter → C1 → C2 tiered JIT; warmup = reaching steady state
  2. First iterations pay class-load + init + cold caches
  3. Skipping warmup = 10-100x too slow and noisy
  4. Discard warmup timings, then measure
  5. Steady-state is wrong target for short-lived processes

basics

~20 s

The JVM starts running code interpreted and only compiles hot code to fast machine code after it runs many times. If you measure the first runs, you measure slow startup, not the real speed. Warmup runs the code first so the fast version is in place before timing.

solid answer

~50 s

Java runs bytecode through a tiered JIT: code starts interpreted, then C1 and finally C2 compile methods that get hot, after thousands of invocations. The first iterations of a loop also pay class-loading and one-time initialization costs. Measuring those gives numbers dominated by interpretation and compilation, not steady-state performance — often 10-100x slower than the real figure, and noisy because compilation happens mid-measurement. A warmup phase runs the benchmarked code (with the same inputs/branches you'll measure) long enough for the JIT to settle on the final compiled version, deoptimizations to stop, and the code/data caches to fill. Only then do you time. JMH formalizes this with explicit warmup iterations. The flip side: if your real workload is short-lived (a CLI, a Lambda cold start), steady-state numbers can overstate performance — then single-shot/AOT-aware measurement matters more.

go deeper

for a junior

Knows the JVM runs slower at first and speeds up, so you must warm up before timing and throw those early numbers away.

for a middle

Can name the tiered JIT (interpreter → C1 → C2), explain class-loading/init costs, and that compilation mid-measurement causes noise; warms up with representative inputs.

for a senior

Discusses deoptimization, profile-guided C2 assumptions, warming up the exact branches/types measured, and reads a warmup curve to decide when steady state is reached rather than guessing.

for a principal

Frames warmup as a function of the workload: chooses steady-state vs single-shot measurement based on whether production is long-running or short-lived, and connects it to AOT/CDS/startup-tuning trade-offs.

## What a microbenchmark is A *microbenchmark* measures the time of a tiny piece of code — one method, one loop, one operation — in isolation, to compare implementations (e.g. `ArrayList` vs `LinkedList` iteration). The goal is the *steady-state* cost: what the operation costs once everything has stabilized. ## Why Java is special: the JVM doesn't run your code the same way twice When you run Java, the compiler (`javac`) produces *bytecode* — a portable intermediate instruction set, not native CPU instructions. At runtime the **JVM** executes that bytecode. It does so in stages, called **tiered compilation**: 1. **Interpretation** — at first the JVM *interprets* bytecode one instruction at a time. This is correct but slow (roughly an order of magnitude slower than native). 2. **C1 (client) JIT** — the *Just-In-Time* compiler. Once a method has been called enough times (it crosses an *invocation/back-edge counter threshold*), the JVM compiles it to native machine code with light optimizations, while also inserting *profiling* counters. 3. **C2 (server) JIT** — for the hottest methods, a heavier optimizing compiler kicks in, using the profile gathered by C1 (which branches are taken, which types actually occur) to produce aggressively optimized native code. So the *same* loop can run as: interpreted → C1-compiled → C2-compiled, getting faster at each step. The transition between these is called **warmup**. ## One-time costs that pollute the first iterations Beyond JIT, the first runs pay costs you pay only once: - **Class loading**: the first time a class is touched it is loaded, verified, and *linked*. - **Class initialization**: static initializers (`<clinit>`) run on first use. - **Lazy resolution** of constant-pool entries and call sites. - **Cold CPU caches and branch predictor**: instruction/data caches and the branch predictor are empty, so early iterations stall more. ## What goes wrong if you skip warmup If you wrap a loop in `System.nanoTime()` and time the first N iterations, you measure mostly interpretation + compilation + class loading. Typical symptoms: - The measured time is **10-100x too slow**. - It is **noisy**: compilation can happen *in the middle* of your measured window, so the same benchmark gives wildly different numbers across runs. - **Deoptimization**: C2 makes speculative assumptions (e.g. "this call site only ever sees type `Foo`"). If a later input violates that, the JVM *deoptimizes* back to the interpreter and recompiles — a sudden slowdown mid-measurement. ## What warmup actually does A warmup phase runs the *same* code, with representative inputs and the *same branches/types* you'll measure, repeatedly until the JIT has reached its final compiled form and stops recompiling. You **discard** those timings and only then measure the *steady state*. Warmup must exercise the real code paths: if you warm up with one type and measure another, you trigger a deopt exactly when timing starts. ## Why a harness exists Deciding "has it warmed up?" reliably is hard. The OpenJDK **JMH** (Java Microbenchmark Harness) runs explicit warmup iterations, then measurement iterations in separate forked JVMs, reports per-iteration variance, and provides hooks (`Blackhole`, `@State`) so the JIT can't optimize the work away. Hand-rolled timing almost always gets warmup (and several other pitfalls) wrong. ## The important caveat Steady-state is the *right* target only if your production workload is long-running. For **short-lived** processes — a CLI tool, a serverless cold start, a build step — the code may *never* reach C2, so the interpreted/C1 numbers are the realistic ones. There, warmed-up steady-state numbers *overstate* real-world performance, and you'd measure single-shot time (or use AOT/CDS) instead.

  • How would you know warmup is 'enough' without guessing a fixed iteration count?
    Run until per-iteration times stabilize across consecutive warmup iterations (variance drops below a threshold), or use a harness like JMH that reports per-iteration numbers so you can see the curve flatten — and run multiple forks to confirm it's reproducible, not a one-off.
  • When is measuring the un-warmed code actually the correct thing to do?
    For short-lived workloads — CLI tools, serverless cold starts, build steps — where the JIT never reaches C2 in production. Then single-shot or startup-time measurement (and AOT/CDS) reflects reality, and steady-state numbers would mislead.

saying these in an interview costs you the question

  • Claiming 'the JVM compiles everything to native up front' — it interprets first, then JIT-compiles hot code lazily.
  • Thinking a fixed 'warm up for 5 iterations' is always enough — thresholds and recompilation depend on the code.
  • Believing warmup just fills CPU caches — the dominant effect is JIT tiered compilation.
  • Assuming steady-state is always the goal — it's wrong for short-lived processes.

context