skip to content

Microbenchmark Pitfalls

All the ways a hand-rolled timing loop lies: no warmup, JIT dead-code elimination, GC and safepoint pauses, coordinated omission, and CPU frequency scaling. The point of the question is to have you conclude that this is exactly why JMH exists.

part ofJavaoverview, primer and where to startread it →
on this pageshow

questions

5

Why must a Java microbenchmark include a warmup phase, and what happens if you measure the very first iterations of a hot loop?

level: juniorimportance: must knowfreq 70%

answer

  1. Interpreter → C1 → C2 tiered JIT; warmup = reaching steady state
  2. First iterations pay class-load + init + cold caches
  3. Skipping warmup = 10-100x too slow and noisy
  4. Discard warmup timings, then measure
  5. Steady-state is wrong target for short-lived processes

basics

~20 s

The JVM starts running code interpreted and only compiles hot code to fast machine code after it runs many times. If you measure the first runs, you measure slow startup, not the real speed. Warmup runs the code first so the fast version is in place before timing.

solid answer

~50 s

Java runs bytecode through a tiered JIT: code starts interpreted, then C1 and finally C2 compile methods that get hot, after thousands of invocations. The first iterations of a loop also pay class-loading and one-time initialization costs. Measuring those gives numbers dominated by interpretation and compilation, not steady-state performance — often 10-100x slower than the real figure, and noisy because compilation happens mid-measurement. A warmup phase runs the benchmarked code (with the same inputs/branches you'll measure) long enough for the JIT to settle on the final compiled version, deoptimizations to stop, and the code/data caches to fill. Only then do you time. JMH formalizes this with explicit warmup iterations. The flip side: if your real workload is short-lived (a CLI, a Lambda cold start), steady-state numbers can overstate performance — then single-shot/AOT-aware measurement matters more.

go deeper

for a junior

Knows the JVM runs slower at first and speeds up, so you must warm up before timing and throw those early numbers away.

for a middle

Can name the tiered JIT (interpreter → C1 → C2), explain class-loading/init costs, and that compilation mid-measurement causes noise; warms up with representative inputs.

for a senior

Discusses deoptimization, profile-guided C2 assumptions, warming up the exact branches/types measured, and reads a warmup curve to decide when steady state is reached rather than guessing.

for a principal

Frames warmup as a function of the workload: chooses steady-state vs single-shot measurement based on whether production is long-running or short-lived, and connects it to AOT/CDS/startup-tuning trade-offs.

## What a microbenchmark is A *microbenchmark* measures the time of a tiny piece of code — one method, one loop, one operation — in isolation, to compare implementations (e.g. `ArrayList` vs `LinkedList` iteration). The goal is the *steady-state* cost: what the operation costs once everything has stabilized. ## Why Java is special: the JVM doesn't run your code the same way twice When you run Java, the compiler (`javac`) produces *bytecode* — a portable intermediate instruction set, not native CPU instructions. At runtime the **JVM** executes that bytecode. It does so in stages, called **tiered compilation**: 1. **Interpretation** — at first the JVM *interprets* bytecode one instruction at a time. This is correct but slow (roughly an order of magnitude slower than native). 2. **C1 (client) JIT** — the *Just-In-Time* compiler. Once a method has been called enough times (it crosses an *invocation/back-edge counter threshold*), the JVM compiles it to native machine code with light optimizations, while also inserting *profiling* counters. 3. **C2 (server) JIT** — for the hottest methods, a heavier optimizing compiler kicks in, using the profile gathered by C1 (which branches are taken, which types actually occur) to produce aggressively optimized native code. So the *same* loop can run as: interpreted → C1-compiled → C2-compiled, getting faster at each step. The transition between these is called **warmup**. ## One-time costs that pollute the first iterations Beyond JIT, the first runs pay costs you pay only once: - **Class loading**: the first time a class is touched it is loaded, verified, and *linked*. - **Class initialization**: static initializers (`<clinit>`) run on first use. - **Lazy resolution** of constant-pool entries and call sites. - **Cold CPU caches and branch predictor**: instruction/data caches and the branch predictor are empty, so early iterations stall more. ## What goes wrong if you skip warmup If you wrap a loop in `System.nanoTime()` and time the first N iterations, you measure mostly interpretation + compilation + class loading. Typical symptoms: - The measured time is **10-100x too slow**. - It is **noisy**: compilation can happen *in the middle* of your measured window, so the same benchmark gives wildly different numbers across runs. - **Deoptimization**: C2 makes speculative assumptions (e.g. "this call site only ever sees type `Foo`"). If a later input violates that, the JVM *deoptimizes* back to the interpreter and recompiles — a sudden slowdown mid-measurement. ## What warmup actually does A warmup phase runs the *same* code, with representative inputs and the *same branches/types* you'll measure, repeatedly until the JIT has reached its final compiled form and stops recompiling. You **discard** those timings and only then measure the *steady state*. Warmup must exercise the real code paths: if you warm up with one type and measure another, you trigger a deopt exactly when timing starts. ## Why a harness exists Deciding "has it warmed up?" reliably is hard. The OpenJDK **JMH** (Java Microbenchmark Harness) runs explicit warmup iterations, then measurement iterations in separate forked JVMs, reports per-iteration variance, and provides hooks (`Blackhole`, `@State`) so the JIT can't optimize the work away. Hand-rolled timing almost always gets warmup (and several other pitfalls) wrong. ## The important caveat Steady-state is the *right* target only if your production workload is long-running. For **short-lived** processes — a CLI tool, a serverless cold start, a build step — the code may *never* reach C2, so the interpreted/C1 numbers are the realistic ones. There, warmed-up steady-state numbers *overstate* real-world performance, and you'd measure single-shot time (or use AOT/CDS) instead.

  • How would you know warmup is 'enough' without guessing a fixed iteration count?
    Run until per-iteration times stabilize across consecutive warmup iterations (variance drops below a threshold), or use a harness like JMH that reports per-iteration numbers so you can see the curve flatten — and run multiple forks to confirm it's reproducible, not a one-off.
  • When is measuring the un-warmed code actually the correct thing to do?
    For short-lived workloads — CLI tools, serverless cold starts, build steps — where the JIT never reaches C2 in production. Then single-shot or startup-time measurement (and AOT/CDS) reflects reality, and steady-state numbers would mislead.

saying these in an interview costs you the question

  • Claiming 'the JVM compiles everything to native up front' — it interprets first, then JIT-compiles hot code lazily.
  • Thinking a fixed 'warm up for 5 iterations' is always enough — thresholds and recompilation depend on the code.
  • Believing warmup just fills CPU caches — the dominant effect is JIT tiered compilation.
  • Assuming steady-state is always the goal — it's wrong for short-lived processes.

context

open as a page

How can the JIT compiler make a Java microbenchmark report a near-zero time for work that genuinely runs, and how do you prevent it?

level: middleimportance: must knowfreq 75%

basics

~20 s

If the result of your code is never used, the JIT can delete the whole computation as dead code. If inputs are constants, it can compute the answer at compile time. Either way the loop disappears and you measure nothing. Fix it by consuming the result (e.g. return it / feed it to a sink) and using non-constant inputs.

open as a page

Why does a dedicated harness like JMH exist instead of hand-rolling System.nanoTime() around a loop? Summarize the pitfalls it addresses.

level: middleimportance: must knowfreq 80%

basics

~20 s

Hand-rolled timing usually lies: it forgets warmup, lets the JIT delete or constant-fold the work, gets perturbed by GC pauses, and uses timers wrong. JMH is built to handle all of that for you — warmup, preventing dead-code elimination, forking JVMs, and reporting proper statistics — so the numbers actually mean something.

open as a page

How do garbage collection and safepoint pauses distort a homegrown Java microbenchmark, and how should you account for them?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Garbage collection can pause your code in the middle of timing, adding spikes that aren't part of the operation you're measuring. The JVM also pauses all threads at 'safepoints' for housekeeping. If you report only an average, these pauses get hidden or smeared; you should look at the distribution and control allocation.

open as a page

What is coordinated omission in latency benchmarking, and why does a naive 'measure each request after the previous one finishes' loop drastically under-report tail latency?

level: principalimportance: should knowfreq 45%

basics

~20 s

If your benchmark sends the next request only after the previous one finishes, then when one request stalls, you simply don't send the requests that should have arrived during the stall. You record one slow sample instead of many. So your worst-case (tail) latency looks far better than what real users would see.

open as a page