What are JMH's benchmark Modes (Throughput, AverageTime, SampleTime, SingleShotTime), and how does @OutputTimeUnit relate to them?
answer
- Throughput = ops/time, higher better, default
- AverageTime = time/op, lower better
- SampleTime = distribution + percentiles (tail latency p99/p999)
- SingleShotTime = one run, no warmup, cold cost
- @OutputTimeUnit only rescales the printed unit
basics
~20 sA Mode tells JMH how to express the result. Throughput counts operations per unit of time (higher is better); AverageTime is time per operation (lower is better); SampleTime samples individual call times to build a distribution including percentiles; SingleShotTime measures one run with no warmup, for cold-start cost. @OutputTimeUnit picks the unit (ms, us, ns) the score is printed in.
solid answer
~50 sJMH's Mode controls what the reported score means. Throughput measures operations per time unit (ops/s) — higher is better — and is the default. AverageTime is the inverse: average time per operation (e.g. ns/op) — lower is better. SampleTime randomly samples the duration of individual invocations so you get a distribution with percentiles (p50, p99, p999), which matters when you care about tail latency rather than the mean. SingleShotTime runs the method exactly once with no warmup, measuring cold, un-JITted, single-invocation cost — useful for startup or first-call overhead. You can request multiple modes at once or Mode.All. @OutputTimeUnit (a java.util.concurrent.TimeUnit such as MILLISECONDS, MICROSECONDS, NANOSECONDS) only changes the unit the score is printed in — it is purely cosmetic/scaling and does not change what is measured. Both Mode and @OutputTimeUnit can be set per-benchmark via annotations or globally via the runner Options.
code
java · 12 linesimport org.openjdk.jmh.annotations.*;
import java.util.concurrent.TimeUnit;
@BenchmarkMode({ Mode.AverageTime, Mode.SampleTime }) // two views at once
@OutputTimeUnit(TimeUnit.NANOSECONDS) // print scores in ns
public class ModeDemo {
@Benchmark
public int hash() {
return "benchmark".hashCode(); // returned -> consumed, not dead-code-eliminated
}
}go deeper
Knows Throughput means ops/sec and AverageTime means time-per-op, and that a Mode is set with an annotation.
Distinguishes all four modes including SampleTime (percentiles) and SingleShotTime (cold/no-warmup), and knows @OutputTimeUnit only rescales output.
Chooses the right mode for the question — tail-latency vs throughput vs startup — and can combine modes or use Mode.All, and overrides via OptionsBuilder.
Connects mode choice to the performance question being answered (SLA percentiles vs steady-state), and is wary of reporting a single mean when the distribution is what governs production behavior.
## The idea of a Mode When you time code you can express the answer two opposite ways: *how much work per unit time* or *how much time per unit work*. JMH formalizes this (and two more shapes) as **`Mode`**, set with `@BenchmarkMode(...)`. ## The four modes - **`Throughput`** (default). Reports **operations per unit of time**, e.g. `ops/s`. JMH runs your `@Benchmark` method as many times as it can in a fixed time window and counts. **Higher is better.** Best for "how many of these can the machine do per second." - **`AverageTime`**. The inverse view: **average time per operation**, e.g. `ns/op` or `us/op`. **Lower is better.** Same underlying measurement as Throughput, just reported as time/op. Best when you think in latencies. - **`SampleTime`**. Instead of dividing total time by count, JMH *samples* the duration of individual invocations and builds a **distribution**. You get percentiles: **p50 (median), p90, p99, p99.9, max**. This is the mode for **tail latency** — when an average hides occasional slow calls (GC pause, cache miss) that matter to an SLA. Sampling adds a little overhead, so very fast operations are less suited to it. - **`SingleShotTime`**. Runs the method **exactly once**, with **no warmup**. It measures **cold-start / first-invocation** cost: interpreter execution, class loading, one-time initialization. Useful for startup-sensitive code; useless for steady-state throughput because there is no warmup and a single sample. You can combine modes: `@BenchmarkMode({Mode.Throughput, Mode.AverageTime})` or `Mode.All` to emit every mode in one run. ## `@OutputTimeUnit` `@OutputTimeUnit(TimeUnit.MICROSECONDS)` (where `TimeUnit` is `java.util.concurrent.TimeUnit`) only chooses the **unit the score is printed in**. It rescales the number — e.g. ns/op into us/op, or ops/s into ops/ms — but **does not change what is measured or how**. Picking a unit that puts the score in a readable range (not `0.0000003` and not `3000000`) just makes output legible. It is a presentation/scaling knob, full stop. ## Where they're set Both are annotations on the benchmark class or method (`@BenchmarkMode`, `@OutputTimeUnit`) and can also be overridden globally through the runner's `OptionsBuilder` (`.mode(...)`, `.timeUnit(...)`), with the runner option winning. A common idiom: class-level `@BenchmarkMode(Mode.AverageTime)` + `@OutputTimeUnit(TimeUnit.NANOSECONDS)` applies to every `@Benchmark` in the class unless a method overrides it. ## Choosing a mode - Comparing two implementations for steady-state speed → Throughput or AverageTime. - Caring about worst-case / SLA latency → SampleTime (read p99/p999). - Measuring first-call / startup cost → SingleShotTime.
- Which mode would you pick to investigate occasional slow calls that hurt your p99 latency?SampleTime — it samples individual invocation durations and reports a distribution with percentiles, so you can read p99/p999 instead of just the mean.
- If you set @OutputTimeUnit to MILLISECONDS but the operation takes ~50 ns, what happens?The measurement is unchanged; the score is just printed in ms, so you'd see a tiny, hard-to-read number like 0.00005 ms/op. Choose a unit that keeps the score readable.
saying these in an interview costs you the question
- Saying @OutputTimeUnit changes what is measured (it only rescales the printout)
- Using SingleShotTime for steady-state throughput (it has no warmup and one sample)
- Confusing Throughput (higher=better) with AverageTime (lower=better) when reading results
- Using AverageTime/Throughput when the question is about tail latency (need SampleTime)