skip to content

Profiling & Benchmarking

Getting numbers you can trust: microbenchmark harnesses, the two families of profiler, and JVM-level event recording. Interviewers ask about measurement because most performance claims fall apart at the methodology stage.

part ofJavaoverview, primer and where to startread it →
on this pageshow

explore

questions

30

What is JMH, and why should you use it instead of hand-rolled timing with System.nanoTime() loops for Java microbenchmarks?

level: juniorimportance: must knowfreq 55%

answer

  1. Official OpenJDK harness, same team as the JVM
  2. @Benchmark + annotation processor generates the harness
  3. Defends against: warmup/JIT, dead-code elimination, constant folding, OSR
  4. Forks a fresh JVM per trial
  5. Naive nanoTime loops measure the wrong thing

basics

~20 s

JMH is the Java Microbenchmark Harness, an official tool for measuring how fast small pieces of Java code run. It's better than manual nanoTime loops because the JVM warms up and optimizes code over time, and JMH handles that warmup and the JIT effects so your numbers are trustworthy.

solid answer

~50 s

JMH (Java Microbenchmark Harness) is the OpenJDK-blessed framework for writing reliable microbenchmarks on the JVM. Hand-rolled nanoTime loops are notoriously wrong because the JVM is a dynamic, adaptive runtime: the JIT compiler only optimizes hot code after thousands of iterations (warmup), it can dead-code-eliminate work whose result is unused, constant-fold inputs the optimizer can prove are fixed, and on-stack-replace a running loop mid-measurement. JMH addresses all of these: it runs warmup iterations before measuring, it forks a fresh JVM per trial to avoid profile pollution between benchmarks, it provides Blackhole and return-value consumption to stop dead-code elimination, and it reports statistics across many iterations and forks. You annotate a method with @Benchmark and JMH's annotation processor generates the harness. The result is numbers you can actually trust instead of measuring the JIT's ability to delete your code.

code

java · 18 lines
java
import org.openjdk.jmh.annotations.*;
import org.openjdk.jmh.infra.Blackhole;

public class WhyJmh {

    // WRONG: a naive loop -- the JIT may dead-code-eliminate or constant-fold this.
    // long t = System.nanoTime();
    // for (int i = 0; i < 1_000_000; i++) Math.log(i); // result thrown away
    // long ns = System.nanoTime() - t;                 // meaningless number

    // RIGHT: JMH consumes the result so the work cannot be deleted.
    @Benchmark
    public double log(Blackhole bh) {
        double sum = 0;
        for (int i = 1; i < 1_000; i++) sum += Math.log(i);
        return sum; // returned -> consumed by JMH
    }
}

go deeper

for a junior

Knows JMH is the standard Java benchmarking tool and that manual nanoTime loops are unreliable because the JVM warms up.

for a middle

Can name the specific distortions (warmup/JIT, dead-code elimination, constant folding) and how JMH counters each.

for a senior

Explains forking, Blackhole/return-value consumption, and warmup vs measurement iterations, and articulates when JMH is the wrong tool (end-to-end load).

for a principal

Frames microbenchmarking results in context — knows micro numbers rarely predict whole-system performance, insists on profiling the real workload, and treats JMH as one input among allocation/GC/cache effects.

## What a microbenchmark is A *microbenchmark* measures the speed of a very small piece of code — a single method, a loop body, one operation — rather than a whole program. The goal is to answer questions like "is approach A faster than approach B for this one operation?" ## Why this is hard on the JVM The **JVM (Java Virtual Machine)** is not a simple interpreter; it is a dynamic, adaptive runtime. Several of its behaviors make naive timing wrong: - **JIT compilation & warmup.** Java code first runs *interpreted* (slow), and only after a method or loop becomes "hot" (executed thousands of times) does the **JIT (Just-In-Time) compiler** compile it to optimized machine code. The first measurements are therefore measuring the interpreter, not the optimized code you care about. The period of running until performance stabilizes is called **warmup**. - **Dead-code elimination (DCE).** The optimizer removes computations whose results are never used. If your benchmark computes `Math.log(x)` and throws the result away, the JIT may delete the call entirely — you'd measure nothing. - **Constant folding.** If an input is a compile-time or provably-constant value, the optimizer can precompute the answer once, so your loop measures nothing. - **On-stack replacement (OSR).** The JVM can swap a still-running interpreted loop for a compiled version mid-flight, producing measurements from artificially-shaped code that never occurs in real call sites. - **GC and background compilation noise** add variance. A hand-written `long t = System.nanoTime(); for (...) { work(); } long elapsed = System.nanoTime() - t;` ignores all of this and routinely produces numbers that are off by orders of magnitude or measure the wrong thing. ## What JMH is **JMH = Java Microbenchmark Harness.** It is the benchmarking framework maintained by the same OpenJDK engineers who build the JVM, precisely because they understand these pitfalls. You write a benchmark by annotating a method with **`@Benchmark`**; JMH's **annotation processor** generates boilerplate "harness" code at build time that runs your method correctly. ## How JMH defends correctness - **Warmup iterations** run your code (results discarded) until the JIT has compiled and stabilized it; only then does it run **measurement iterations**. - **Forking:** each trial runs in a *fresh JVM process* so one benchmark's compilation profile can't pollute another's. - **Blackhole / return-value consumption:** returning a value from `@Benchmark` (or passing it to a `Blackhole`) tells JMH to *consume* it so the JIT cannot dead-code-eliminate the work. - **State objects** (`@State`) hold inputs in a way the optimizer can't constant-fold. - **Statistics:** JMH aggregates many iterations across many forks and reports a score with error/confidence, not a single lucky number. ## When NOT to use it JMH is for *micro* benchmarks (sub-millisecond to millisecond operations). For end-to-end/application throughput you'd use load-testing tools instead. But for "which of these two implementations of a hot method is faster," JMH is the standard answer.

  • Name two ways the JIT can make a naive benchmark report a misleadingly fast result.
    Dead-code elimination (it deletes work whose result is unused) and constant folding (it precomputes results for provably-constant inputs), so the loop ends up measuring nothing.
  • Who maintains JMH and why does that matter?
    The OpenJDK / JVM engineers maintain it; it matters because they know exactly which JIT and runtime behaviors corrupt naive benchmarks and bake the defenses into the harness.

saying these in an interview costs you the question

  • Claiming a simple nanoTime loop is 'good enough' for micro-level comparisons
  • Not knowing the JVM has a warmup/JIT phase that distorts early measurements
  • Thinking JMH is for end-to-end load testing rather than micro-level method timing

context

open as a page

What is dead-code elimination in the context of a JMH benchmark, and why can it make a microbenchmark report a misleadingly fast result?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Dead-code elimination is when the JVM deletes code whose result is never used. In a benchmark, if you compute something but never use the value, the JIT may delete the work, so the benchmark times nothing and reports an impossibly fast result.

open as a page

What is the difference between @Warmup and @Measurement iterations in JMH, and how do you configure their count and time?

level: juniorimportance: must knowfreq 60%

basics

~20 s

@Warmup runs the code to get the JVM up to speed and throws those numbers away. @Measurement runs it again and keeps those numbers. You set how many rounds (iterations) and how long each round lasts (time).

open as a page

Why must a Java microbenchmark include a warmup phase, and what happens if you measure the very first iterations of a hot loop?

level: juniorimportance: must knowfreq 70%

basics

~20 s

The JVM starts running code interpreted and only compiles hot code to fast machine code after it runs many times. If you measure the first runs, you measure slow startup, not the real speed. Warmup runs the code first so the fast version is in place before timing.

open as a page

What is the difference between a sampling profiler and an instrumentation (tracing) profiler in Java?

level: juniorimportance: must knowfreq 70%

basics

~20 s

A sampling profiler peeks at what the program is doing every so often and adds up where it spends most time. An instrumentation profiler adds extra code to every method to count and time every single call exactly, which is more accurate but much slower.

open as a page

How do you start a JFR recording — both at JVM launch and on an already-running process — and how do you produce a .jfr file from it?

level: middleimportance: must knowfreq 60%

basics

~20 s

At launch, add the JVM flag -XX:StartFlightRecording with options like duration and filename. On a running process, use the jcmd tool: 'jcmd <pid> JFR.start', then 'JFR.dump' to write a .jfr file and 'JFR.stop' to end it.

open as a page

When should you return a value from a @Benchmark method versus using a Blackhole, and how do you sink multiple results correctly?

level: middleimportance: must knowfreq 68%

basics

~20 s

Return the value when your benchmark produces a single result. Use a Blackhole when you produce more than one result, or a result inside a loop, calling bh.consume() on each so none of them gets deleted by the JIT.

open as a page

Why does a JMH benchmark need a warmup phase, and what happens to the JVM during it?

level: middleimportance: must knowfreq 70%

basics

~10 s

The JVM speeds up code as it runs by compiling it. Warmup runs the code a while first so it is already fast, so you measure the fast version, not the slow startup.

open as a page

How can the JIT compiler make a Java microbenchmark report a near-zero time for work that genuinely runs, and how do you prevent it?

level: middleimportance: must knowfreq 75%

basics

~20 s

If the result of your code is never used, the JIT can delete the whole computation as dead code. If inputs are constants, it can compute the answer at compile time. Either way the loop disappears and you measure nothing. Fix it by consuming the result (e.g. return it / feed it to a sink) and using non-constant inputs.

open as a page

Why does a dedicated harness like JMH exist instead of hand-rolling System.nanoTime() around a loop? Summarize the pitfalls it addresses.

level: middleimportance: must knowfreq 80%

basics

~20 s

Hand-rolled timing usually lies: it forgets warmup, lets the JIT delete or constant-fold the work, gets perturbed by GC pauses, and uses timers wrong. JMH is built to handle all of that for you — warmup, preventing dead-code elimination, forking JVMs, and reporting proper statistics — so the numbers actually mean something.

open as a page

You suspect a Java service is slow under production-like load. Walk through how you'd choose between a sampling and an instrumentation profiler and which tools you'd reach for.

level: middleimportance: must knowfreq 55%

basics

~20 s

Start with a low-overhead sampling profiler (like async-profiler or JFR) under realistic load to see where time actually goes, because it barely slows the app. Only switch to an instrumentation profiler if you need exact call counts or to trace one specific path, knowing it's slower and can distort results.

open as a page

What is constant folding in a benchmark, and how does the JMH @State pattern prevent the JIT from precomputing your inputs?

level: seniorimportance: must knowfreq 60%

basics

~20 s

Constant folding is when the JIT precomputes an expression whose inputs are compile-time constants, so the benchmark never actually runs the work. Putting inputs in a non-final @State field that the JIT can't treat as constant forces the real computation to happen every time.

open as a page

What is Java Flight Recorder (JFR), and what makes it suitable for use in production rather than only in a test environment?

level: juniorimportance: should knowfreq 55%

basics

~20 s

JFR is a profiling and event-recording tool built into the JVM. It records what your application does (allocations, GC, locks, exceptions, method samples) with very low overhead—around 1% or less—so you can safely leave it on in production.

open as a page

Once you have a .jfr file, how do you analyze it, and what is JDK Mission Control's role? How might custom application events fit in?

level: middleimportance: should knowfreq 40%

basics

~20 s

You open the .jfr file in JDK Mission Control (JMC), a desktop app that shows automated analysis and views like CPU usage, allocations, GC, locks, and exceptions. You can also read .jfr files from the command line with 'jfr print' or programmatically with the jdk.jfr.consumer API, and define your own custom events for app-specific telemetry.

open as a page

What are JMH's benchmark Modes (Throughput, AverageTime, SampleTime, SingleShotTime), and how does @OutputTimeUnit relate to them?

level: middleimportance: should knowfreq 45%

basics

~20 s

A Mode tells JMH how to express the result. Throughput counts operations per unit of time (higher is better); AverageTime is time per operation (lower is better); SampleTime samples individual call times to build a distribution including percentiles; SingleShotTime measures one run with no warmup, for cold-start cost. @OutputTimeUnit picks the unit (ms, us, ns) the score is printed in.

open as a page

Walk through setting up and running a minimal JMH benchmark project: dependency/annotation-processor setup, the generated harness, and how a benchmark gets executed.

level: middleimportance: should knowfreq 40%

basics

~20 s

Add the JMH core library plus its annotation processor as dependencies. Write a method annotated with @Benchmark. At build time the annotation processor generates harness code and a runnable Main. Then you run the benchmarks by executing that generated jar (or via a Runner in code), and JMH does warmup, measurement, and forking automatically.

open as a page

Why does a sampling profiler have roughly constant overhead while an instrumentation profiler's overhead scales with how often methods are called?

level: middleimportance: should knowfreq 50%

basics

~20 s

Sampling only does work once per timer tick (e.g. every 10 ms), no matter how busy the code is, so its cost is fixed. Instrumentation does extra work on every single method call, so the more calls there are, the more it costs.

open as a page

What is the difference between a continuous (always-on) JFR recording and a profiling recording, and how does this shape an incident-investigation strategy in production?

level: seniorimportance: should knowfreq 42%

basics

~20 s

A continuous recording runs all the time with low-overhead settings and a bounded ring buffer, so you can dump the recent history when something goes wrong. A profiling recording is a short, time-boxed session with richer events and slightly higher cost, used for a focused investigation.

open as a page

Describe JFR's event model: what kinds of events does it capture, and how does the sampling-and-buffering design keep the cost low enough for production?

level: seniorimportance: should knowfreq 48%

basics

~20 s

JFR records typed 'events' for things like object allocation, garbage collection, lock contention, thrown exceptions, file/socket I/O, and periodic method (CPU) samples. Events go into per-thread buffers and CPU profiling uses sampling, so the overhead stays around 1%.

open as a page

Why does JMH fork a fresh JVM for each trial, and what is 'profile pollution' that forking prevents?

level: seniorimportance: should knowfreq 38%

basics

~20 s

JMH runs each benchmark in a brand-new JVM process (a fork). This is because the JVM remembers and adapts based on what it ran before, so running two benchmarks in the same JVM lets the first one's optimization decisions affect the second's results. A fresh JVM per trial keeps each benchmark's measurement honest and independent.

open as a page

Explain @State scopes (Benchmark, Thread, Group) and the @Setup/@TearDown lifecycle with their Levels in JMH.

level: seniorimportance: should knowfreq 42%

basics

~20 s

A @State class holds the data your benchmark uses. Its scope says who shares one instance: Benchmark = all threads share one, Thread = each thread gets its own, Group = one per group of cooperating threads. @Setup methods prepare state before the benchmark and @TearDown cleans up after; each can run at Trial, Iteration, or Invocation level depending on how often you need it.

open as a page

Why are manual loops inside a @Benchmark body discouraged, and what JIT optimizations can corrupt the measurement?

level: seniorimportance: should knowfreq 52%

basics

~20 s

A loop inside a benchmark can be unrolled or partly optimized by the JIT, and repeated identical work can be folded so only one iteration really runs. Prefer letting JMH do the repetition, and if you must loop, consume each iteration's result and vary the inputs.

open as a page

Why does JMH run benchmarks in multiple forks, and what does @Fork control?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Each fork is a fresh JVM. JMH runs your benchmark in several separate JVMs and averages the results, because one JVM can get 'lucky' or 'unlucky' in how it compiles your code. Multiple forks smooth that out.

open as a page

How do you read a JMH result line — the score, the ± error, and the units — and decide whether two benchmarks actually differ?

level: seniorimportance: should knowfreq 50%

basics

~20 s

The score is the average result (e.g. nanoseconds per operation or ops per second). The ± number is the margin of error — how much the score could wobble. If two benchmarks' score-plus-or-minus ranges overlap, you can't claim one is faster.

open as a page

How do garbage collection and safepoint pauses distort a homegrown Java microbenchmark, and how should you account for them?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Garbage collection can pause your code in the middle of timing, adding spikes that aren't part of the operation you're measuring. The JVM also pauses all threads at 'safepoints' for housekeeping. If you report only an average, these pauses get hidden or smeared; you should look at the distribution and control allocation.

open as a page

What is safepoint bias in JVM sampling profilers, and how do modern profilers like async-profiler avoid it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Older JVM samplers can only grab a thread's stack when the thread reaches a 'safepoint' — a safe pause spot. The JIT removes safepoints from tight loops, so the sampler records the wrong place and blames safepoint-friendly methods. async-profiler avoids this by capturing the stack instantly via OS signals, no safepoint needed.

open as a page

You inherit a JMH benchmark reporting 0.3 ns/op for a non-trivial computation. How do you systematically audit whether the JIT optimized the work away?

level: principalimportance: should knowfreq 40%

basics

~20 s

A sub-nanosecond result for real work is a red flag. Check that inputs come from non-final @State fields (not constants), that every output is returned or Blackhole-consumed, and that any loop varies its data. Then re-run, ideally inspecting the generated assembly to confirm the work runs.

open as a page

What benchmarking pitfalls does JMH guard against, and how do dead-code elimination, constant folding, and Blackhole/@State relate to warmup-stage validity?

level: principalimportance: should knowfreq 40%

basics

~20 s

If your benchmark's result isn't used, the JIT can delete it (dead-code elimination) or precompute it (constant folding) — so you'd measure nothing. JMH fixes this: return values, consume them with Blackhole, and keep inputs in @State so they aren't treated as constants.

open as a page

What is coordinated omission in latency benchmarking, and why does a naive 'measure each request after the previous one finishes' loop drastically under-report tail latency?

level: principalimportance: should knowfreq 45%

basics

~20 s

If your benchmark sends the next request only after the previous one finishes, then when one request stalls, you simply don't send the requests that should have arrived during the stall. You record one slow sample instead of many. So your worst-case (tail) latency looks far better than what real users would see.

open as a page

Beyond raw overhead, how can an instrumentation profiler change a program's behavior so much that its results misdirect optimization, and what would you do about it at scale?

level: principalimportance: nice to knowfreq 25%

basics

~20 s

Adding probes to every method can stop the JIT from optimizing (especially inlining) small methods, so the profiled program runs differently from production and points you at the wrong 'hot' code. To avoid being misled, prefer low-overhead sampling, scope any instrumentation narrowly, and always validate findings against a sampled, optimized run.

open as a page