skip to content

How do you read a JMH result line — the score, the ± error, and the units — and decide whether two benchmarks actually differ?

level: seniorimportance: should knowfreq 50%

answer

  1. Score ± Error Units, under a Mode
  2. ns/op lower better; ops/s higher better
  3. Error ≈ 99.9% CI half-width
  4. No overlap of intervals → real difference
  5. Big error → more warmup/iterations/forks

basics

~20 s

The score is the average result (e.g. nanoseconds per operation or ops per second). The ± number is the margin of error — how much the score could wobble. If two benchmarks' score-plus-or-minus ranges overlap, you can't claim one is faster.

solid answer

~40 s

A JMH result line reports a **Mode** (what's being measured — average time, throughput, etc.), a **Score** (the central estimate), an **Error** shown as `± x` (typically a 99.9% confidence interval half-width), and a **Unit** (e.g. `ns/op` or `ops/s`). The score is aggregated across all measurement iterations and forks; the error captures the spread, so a small error means a tight, reproducible result. The crucial interpretation rule: to claim benchmark A is faster than B, their `score ± error` ranges must **not overlap** — overlapping intervals mean the measured difference is within noise and not established. Also mind the unit: in `AverageTime`/`ns/op`, **lower is better**; in `Throughput`/`ops/s`, **higher is better** — confusing the two flips your conclusion. A large error usually signals insufficient warmup, too few iterations/forks, or a noisy machine.

code

java · 9 lines
java
// Example output:
// Benchmark      Mode  Cnt   Score   Error  Units
// A.fast         avgt   15  10.0  ± 0.3   ns/op   -> [9.7, 10.3]
// B.slow         avgt   15  12.0  ± 0.4   ns/op   -> [11.6, 12.4]
// Intervals do NOT overlap  -> A is faster (real difference).
//
// A.fast         avgt   15  10.0  ± 1.5   ns/op   -> [8.5, 11.5]
// B.slow         avgt   15  11.0  ± 1.5   ns/op   -> [9.5, 12.5]
// Intervals OVERLAP        -> difference not established.

go deeper

for a junior

Identifies the score as the average and the ± as a margin of error, and reads the unit.

for a middle

Knows the error is a confidence interval and that direction depends on mode (ns/op vs ops/s).

for a senior

Applies the non-overlapping-intervals rule to decide if a difference is real and diagnoses a large error (warmup/iterations/forks/machine noise).

for a principal

Distinguishes mean vs tail (percentiles), reasons about how warmup and fork choices bias the error, and sets a team standard for what counts as a statistically supported performance claim.

## A sample result line ``` Benchmark Mode Cnt Score Error Units MyBench.hash avgt 15 12.481 ± 0.207 ns/op ``` Every column matters; let's define each. ## Mode — *what* is being measured JMH can measure several things, set by `@BenchmarkMode`: - **Throughput** (`thrpt`) — operations per unit time. **Higher is better.** Unit like `ops/s`. - **AverageTime** (`avgt`) — average time **per operation**. **Lower is better.** Unit like `ns/op`. - **SampleTime** — samples individual call latencies, lets you read percentiles (p50, p99…). - **SingleShotTime** — one cold call, no warmup; for measuring first-call/cold-start cost. The single biggest reading mistake is forgetting the direction: with `ns/op` lower is better, with `ops/s` higher is better. ## Cnt — how many data points The number of measurement iterations aggregated, summed across all forks (e.g. 3 forks × 5 iterations = 15). More data points generally tighten the error. ## Score — the central estimate The **mean** of all measurement-iteration results across all forks. It's your best single-number answer ('about 12.5 ns per call'). But a score *without* its error is meaningless — you can't tell if 12.5 vs 12.7 is signal or noise. ## Error (`± x`) — the confidence interval JMH prints `± x` where `x` is the **half-width of the confidence interval** (by default ~99.9%). Interpretation: JMH is highly confident the true mean lies within `score ± x`, i.e. roughly `[12.274, 12.688] ns/op` for the example. - **Small error** relative to score → tight, reproducible measurement; you can trust comparisons at that resolution. - **Large error** → the measurement is noisy. Common causes: not enough warmup (JIT still settling), too few iterations or forks, a busy/throttling machine, GC pauses, or inherently variable code (e.g. allocation-heavy). ## The decision rule: do A and B really differ? To claim A is faster than B, compare their **intervals**, not just their scores: - A = `10.0 ± 0.3`, B = `12.0 ± 0.4`: A's interval `[9.7, 10.3]` and B's `[11.6, 12.4]` **don't overlap** → the difference is real at that confidence. - A = `10.0 ± 1.5`, B = `11.0 ± 1.5`: intervals `[8.5, 11.5]` and `[9.5, 12.5]` **overlap** → you have *not* shown a difference; reduce noise (more warmup/iterations/forks, quieter machine) before concluding. This is the most important practical skill: **never compare bare scores; compare score±error ranges.** Overlap = inconclusive. ## Percentiles vs mean For latency-sensitive work, the **mean is not enough** — a low average can hide a bad tail. Use `SampleTime` mode and read p99/p99.9; two services with the same mean can have very different tail latency. ## How error connects to warmup and forks (full circle) - Insufficient **warmup** → early measurement iterations include JIT cost → inflated, skewed score and a fat error. - Too few **forks** → blind to inter-JVM variance → the error can look *deceptively small* (precise but biased). So a 'good' result is: adequate warmup (scores stabilized), enough measurement iterations and forks, and a small error *that honestly includes* between-JVM variance. ## Interview-ready summary Read it as `Score ± Error Units` under a given Mode. Score is the mean across iterations and forks; Error is the (≈99.9%) confidence half-width. Respect the unit's direction (ns/op lower-is-better, ops/s higher-is-better), and only claim a difference when the two `score ± error` intervals don't overlap.

  • Two benchmarks score 10.0 ± 1.2 and 10.8 ± 1.1 ns/op. Can you say the first is faster?
    No. The intervals are roughly [8.8, 11.2] and [9.7, 11.9], which overlap heavily. The measured difference is within noise; you'd need to reduce variance (more warmup, more iterations/forks, a quieter machine) before making the claim.
  • Why might the mean score be misleading for a latency-sensitive service?
    Because the mean hides the tail. Two workloads with identical means can have very different p99/p99.9 latencies. Use SampleTime mode to read percentiles when tail latency matters.

saying these in an interview costs you the question

  • Comparing bare scores and ignoring the ± error / overlapping intervals.
  • Treating 'lower' as always better — wrong for Throughput (ops/s), where higher is better.
  • Assuming a tiny error means a correct result, even when only @Fork(1) was used.
  • Relying on the mean for latency-sensitive code instead of percentiles.

context