Why does JMH run benchmarks in multiple forks, and what does @Fork control?
answer
- Fork = a fresh JVM process
- Variance lives between JVMs, not just within
- Each fork warms up then measures
- @Fork(0) = in-harness, debug only
- More forks → honest (wider) error interval
basics
~20 sEach fork is a fresh JVM. JMH runs your benchmark in several separate JVMs and averages the results, because one JVM can get 'lucky' or 'unlucky' in how it compiles your code. Multiple forks smooth that out.
solid answer
~50 s@Fork controls how many separate JVM processes JMH launches to run the benchmark, each doing its own warmup + measurement. The reason is run-to-run variance: a single JVM can land in a non-deterministic state — a particular JIT compilation plan, inlining decision, profile, or memory layout — that biases the result. Some of that variance is *between* JVM instances, not within one, so averaging many iterations inside one JVM won't reveal it. Multiple forks expose and average over that inter-JVM variance, giving a more honest score and a wider, more truthful error. `@Fork(1)` is acceptable for quick local checks; serious comparisons use `@Fork(3)` or more. `@Fork(0)` runs in the harness JVM (handy for debugging, but not for real numbers). You can also pass extra JVM args per fork via `@Fork(jvmArgs = ...)` to test under different flags.
code
java · 11 lines@Fork(value = 3, jvmArgs = {"-Xms1g", "-Xmx1g"}) // 3 fresh JVMs, fixed heap
@Warmup(iterations = 5, time = 1, timeUnit = TimeUnit.SECONDS)
@Measurement(iterations = 5, time = 1, timeUnit = TimeUnit.SECONDS)
public class ForkedBench {
@Benchmark
public int work() {
return Integer.bitCount(0xDEADBEEF);
}
}
// Debug only (do NOT report numbers from this):
// @Fork(0)go deeper
Knows a fork is a fresh JVM and that running several gives a more reliable average.
Explains @Fork controls the number of JVM processes and that each fork warms up then measures.
Articulates why — inter-JVM variance from non-deterministic JIT/inlining/profiling — and uses error intervals across forks to judge whether differences are real.
Sets fork-count policy for trustworthy comparisons, uses per-fork jvmArgs to isolate flag effects, and treats overlapping confidence intervals across forks as 'no measured difference'.
## The problem forks solve: per-JVM non-determinism You might assume that if you run enough iterations in one JVM you'll get a stable, trustworthy number. Often you won't — because a meaningful chunk of benchmark variance is **between JVM runs**, not within a single run. Why? The JVM makes many **non-deterministic, run-specific decisions** while it executes: - **Which JIT compilation plan it lands on.** The C2 compiler uses *profiling data* gathered early in the run to decide which branches to optimize for, which method calls to **inline**, which types are likely. Slight timing differences in early execution can lead two JVMs to make *different* optimization choices, so each one's steady-state code is genuinely different. - **Inlining decisions and code layout.** Whether a hot call gets inlined, and where compiled code lands in memory, affects instruction-cache behavior. - **GC and heap layout, address-space randomization, background compiler threads.** All introduce per-process variation. The consequence: JVM #1 might consistently report 100 ops/µs and JVM #2 might consistently report 115 ops/µs, each *stable within itself*. If you only ever run one JVM, you'd report whichever you happened to get and call it precise — but it's biased. Averaging a million iterations inside that one JVM does **not** fix it, because the variance lives between JVMs. ## What @Fork does A **fork** is a brand-new, separate JVM **process** that JMH spawns to run the benchmark from scratch (its own warmup, its own measurement). `@Fork(n)` runs `n` such processes sequentially and **aggregates the measurement data across all of them**. This way the reported score and error reflect inter-JVM variance, not just intra-JVM noise. - `@Fork(1)` — one fresh JVM. Better than nothing; fine for quick local iteration. - `@Fork(3)` (or 5, 10) — recommended for results you intend to trust or publish; the more you compare small differences, the more forks you want. - `@Fork(0)` — runs in the **harness's own JVM** (no separate process). Useful for attaching a debugger or profiler, but the harness JVM is already 'dirtied' by JMH's own code, so the numbers are unreliable. Never report from `@Fork(0)`. ## warmupForks and warmup-per-fork Warmup happens **inside each fork** — every fresh JVM repeats the discarded warmup iterations before its measured ones. (There is also a separate `warmups` concept — dedicated whole forks whose entire output is discarded — used to stabilize cross-fork state, but the common case is: each fork warms up then measures.) ## Passing JVM args per fork `@Fork(value = 3, jvmArgs = {"-Xms2g", "-Xmx2g"})` (or `jvmArgsAppend`) lets you control the child JVM's flags — e.g. fix the heap size, disable tiered compilation, or turn on `-XX:+PrintCompilation`. This is also how you run the *same* benchmark under two different flag sets to compare them. ## How forks interact with the error interval JMH's reported **error** (confidence interval) is computed from the spread of measurement data across iterations **and forks**. More forks generally widen the interval to its *honest* size by capturing inter-JVM variance that a single fork hides. A suspiciously tiny error from `@Fork(1)` is a classic trap: it looks precise but is really just blind to between-JVM variation. ## Practical guidance - Quick local exploration: `@Fork(1)`, modest iterations. - A/B comparison or anything you'll cite: `@Fork(3)`+ and compare the error intervals, not just the central scores. If intervals overlap, the difference may not be real. - Debugging/profiling: `@Fork(0)` to stay in-process, but don't trust those numbers. ## One-line summary for an interview @Fork runs the benchmark in N fresh JVM processes and averages across them, because a real and large part of benchmark variance is *between* JVM instances (different JIT/inlining/profile decisions), which no amount of in-JVM iteration can reveal.
- Why can't you just average more iterations in a single JVM instead of using multiple forks?Because a large part of benchmark variance is *between* JVM instances — different JIT compilation plans, inlining decisions, and profiles. Within one JVM those decisions are fixed once steady state is reached, so extra iterations just measure the same biased state more precisely. Only separate forks resample that between-JVM variance.
- When is @Fork(0) appropriate?Only for debugging or attaching a profiler, because it runs in the harness's own JVM with no separate process. The harness JVM is already polluted by JMH's own code, so the resulting numbers must not be reported.
Tasting one batch of cookies a hundred times tells you about that batch; baking three batches and tasting each tells you about the recipe. Forks are extra batches (JVMs).
saying these in an interview costs you the question
- Claiming one fork with many iterations is statistically equivalent to many forks — it ignores inter-JVM variance.
- Reporting numbers from @Fork(0).
- Thinking a fork is a thread; it is a separate JVM process.
- Assuming a small error interval from @Fork(1) means the result is trustworthy.