skip to content

Why does JMH fork a fresh JVM for each trial, and what is 'profile pollution' that forking prevents?

level: seniorimportance: should knowfreq 38%

answer

  1. Fork = fresh JVM process per trial
  2. JVM is adaptive: profiles branches/types, JIT inlines & speculates
  3. Profile pollution = benchmark A biases B's inherited profile (megamorphic, deopt)
  4. Fresh fork => clean profile, order-independent
  5. Multiple forks expose run-to-run/between-JVM variance; @Fork(0)=debug only

basics

~20 s

JMH runs each benchmark in a brand-new JVM process (a fork). This is because the JVM remembers and adapts based on what it ran before, so running two benchmarks in the same JVM lets the first one's optimization decisions affect the second's results. A fresh JVM per trial keeps each benchmark's measurement honest and independent.

solid answer

~60 s

Forking means JMH launches a separate JVM process for each trial instead of running everything in one VM. The reason is that the JVM is adaptive: the JIT makes optimization decisions based on the runtime profile it has gathered — which branches were taken, which types appeared at a call site, which methods got inlined. If two benchmarks share one JVM, the first warms up the runtime in a particular shape, and the second inherits that biased profile — for example a megamorphic call site (many implementing types seen) or a deoptimization triggered by the first benchmark. That cross-contamination is 'profile pollution,' and it makes results non-reproducible and order-dependent. A fresh fork gives each benchmark a clean profile. JMH also runs multiple forks (@Fork(n)) and aggregates across them, because even a single clean JVM can land in a 'lucky' or 'unlucky' compilation/layout that differs run to run; multiple forks expose that run-to-run variance instead of trusting one process. @Fork(0) disables forking (runs in-process) and is for debugging only, never for trustworthy numbers.

go deeper

for a junior

Knows JMH starts a fresh JVM per benchmark so runs don't interfere, and that you generally shouldn't disable forking.

for a middle

Can define profile pollution and explain that the JVM adapts based on what it has seen, so sharing a JVM contaminates results.

for a senior

Explains the JIT mechanisms behind pollution (megamorphic call sites, deopt, inlining) and why multiple forks are needed for run-to-run variance, plus the debug-only role of @Fork(0).

for a principal

Reasons about reproducibility and statistical validity end to end — fork count vs iteration count trade-offs, between- vs within-fork variance, and how to defend a benchmark's conclusions against an adaptive runtime.

## What 'forking' means here In JMH, a **fork** is a **separate JVM operating-system process** started just to run one benchmark **trial**. `@Fork(5)` means JMH will start the benchmark in 5 fresh JVMs in turn and combine the results. This is unrelated to `fork()` in the C/Unix sense beyond the shared idea of "new process." ## Why a fresh JVM matters: the JVM is adaptive The JVM does not execute Java the same way every time. As code runs, it **profiles** itself: it records statistics like *which branch of an if was usually taken*, *which concrete types showed up at a virtual call site*, *how hot each method is*. The **JIT compiler** then uses that profile to make aggressive optimizations: - **Inlining** hot, small methods into their callers. - **Monomorphic/bimorphic call-site optimization** — if a call site has only seen one or two concrete types, the JIT compiles a fast direct/inlined path with a type guard. If it later sees *many* types, the site becomes **megamorphic** and that fast path is lost. - **Branch prediction / speculative optimization** based on observed branch frequencies. - **Deoptimization** — if a speculative assumption is later violated, the JIT throws away the compiled code and falls back, which is slow and changes subsequent behavior. ## Profile pollution **Profile pollution** is when running benchmark A in a JVM **biases the runtime profile that benchmark B then inherits** in the *same* JVM. Concrete examples: - A runs a method against types `X` and `Y`; a shared call site becomes **bimorphic/megamorphic**. B reuses that site and is now slower than it would be in isolation — or faster/slower depending on direction. - A triggers a **deoptimization** or class loading that changes the compilation state B starts from. - A's execution warms up shared library code (e.g. `HashMap`) into a shape tuned for A's key distribution, advantaging or penalizing B. The result: B's number **depends on what ran before it**. Benchmarks become **order-dependent and non-reproducible** — the cardinal sin of measurement. ## Forking fixes it Giving each benchmark its **own fresh JVM** means it starts from a **clean, empty profile**. No previous benchmark could have polluted its call sites or compilation state. This is why JMH **forks by default** (`@Fork` ≥ 1). ## Why *multiple* forks, not just one Even a single clean JVM can land in a **'lucky' or 'unlucky'** state: address-space layout, exact compilation order, GC timing, and CPU/cache effects vary run to run, and the JIT can compile a method into a genuinely faster or slower shape on different runs. Running **several forks** and aggregating exposes this **run-to-run (between-JVM) variance** and prevents you from trusting one process that happened to be fast or slow. JMH reports error/confidence partly from this fork-level spread. ## @Fork(0) Setting `@Fork(0)` runs the benchmark **in the same JVM as the harness/runner** (no new process). It's faster to start and lets you attach a debugger, so it's handy for **debugging the benchmark code** — but the numbers are polluted by the harness and any prior runs, so **never use it for results you'll trust or publish.**

  • Give a concrete example of profile pollution between two benchmarks in one JVM.
    Benchmark A exercises a virtual call site with several concrete types, making it megamorphic; benchmark B then reuses that site and loses the monomorphic fast path, so B's score is distorted by A having run first.
  • If forking already gives a clean profile, why run more than one fork?
    Because a single JVM run can be 'lucky' or 'unlucky' — compilation order, memory layout, GC and cache effects vary between processes. Multiple forks reveal that between-JVM variance instead of trusting one possibly-atypical run.

saying these in an interview costs you the question

  • Thinking forking is about threads/parallelism rather than process isolation of the JIT profile
  • Using @Fork(0) for real measurements
  • Believing one fork is enough — ignoring between-JVM run-to-run variance
  • Not knowing the JVM's adaptive profiling/compilation is what creates the contamination

context