Why are manual loops inside a @Benchmark body discouraged, and what JIT optimizations can corrupt the measurement?
answer
- JMH already loops for you
- Unroll / hoist / fold identical iterations
- Vary data from @State per iteration
- Consume each iteration; declare @OperationsPerInvocation
- Loops multiply the JIT's optimization chances
basics
~20 sA loop inside a benchmark can be unrolled or partly optimized by the JIT, and repeated identical work can be folded so only one iteration really runs. Prefer letting JMH do the repetition, and if you must loop, consume each iteration's result and vary the inputs.
solid answer
~50 sJMH already repeats your @Benchmark method many times, so adding your own loop over identical work invites trouble. The JIT can unroll the loop, hoist invariant computations out, and — if every iteration computes the same thing — fold the repeated work so effectively one iteration runs while you divide by N and report an artificially low per-op cost. Loop unrolling also amortizes per-iteration overhead in ways that misrepresent the real cost. If you genuinely need a loop (e.g. to amortize a fixed setup, or to process a data structure), the rules are: consume every iteration's result with a Blackhole so none is eliminated, make iterations depend on data that varies (read from @State arrays) so they can't be folded into one, and be aware JMH will still report per-method time, not per-iteration. The cleanest approach is usually to size the work to one method call and let JMH handle repetition, optionally using @OperationsPerInvocation when a fixed internal count is unavoidable.
code
java · 28 linesimport org.openjdk.jmh.annotations.*;
import org.openjdk.jmh.infra.Blackhole;
public class LoopBench {
// Identical iterations + discarded result: JIT may fold to one and/or delete.
@Benchmark
public void risky() {
double s = 0;
for (int i = 0; i < 1000; i++) s += Math.sqrt(2.0);
// s is never consumed -> dead code; same input -> foldable.
}
@State(Scope.Thread)
public static class Data {
double[] xs = new double[1000];
@Setup public void fill() {
for (int i = 0; i < xs.length; i++) xs[i] = i + 1; // varied, non-constant
}
}
// Varied inputs + per-iteration consume + declared op count.
@Benchmark
@OperationsPerInvocation(1000)
public void better(Data d, Blackhole bh) {
for (double x : d.xs) bh.consume(Math.sqrt(x));
}
}go deeper
Knows JMH repeats the benchmark for you and that adding your own loop can cause trouble.
Avoids unnecessary loops, and when looping, consumes each result and reads varied inputs from @State.
Names the specific JIT effects (unrolling, hoisting, folding identical iterations) and uses @OperationsPerInvocation to normalize a deliberate internal count.
Designs benchmarks that isolate exactly the cost of interest, reasons about how unrolling/hoisting interact with the measured workload, and sets team conventions for when loops are acceptable.
## Why loops are even a question JMH already calls your `@Benchmark` method in a tight harness loop and times the whole run, dividing by the number of invocations. So a *hand-written* loop inside the method is redundant repetition — and worse, it hands the **JIT** (the runtime optimizing compiler) several ways to distort the result. ## The optimizations that bite **1. Loop unrolling.** The JIT can replicate the loop body so fewer loop-control instructions run per unit of work. That is great for real programs, but it changes the per-iteration cost you think you are measuring — the overhead you wanted to capture is amortized away. **2. Loop-invariant code motion (hoisting).** If part of the loop body does not change between iterations, the JIT lifts it out of the loop so it runs once. If your 'work' is invariant, it is now executed once, not N times. **3. Folding identical iterations.** If every iteration computes the *same* value from the *same* inputs, the compiler may compute it once and reuse it. You then divide a single computation's time by N and report a per-op cost that is N times too small. **4. Dead-code elimination across the loop.** If you only keep the last iteration's result (or none), the JIT deletes the others, just like the single-result DCE case. ## What to do instead **Default: no manual loop.** Let JMH do the repetition. Make the `@Benchmark` method do *one* unit of work, return/consume its result, read inputs from `@State`. JMH's invocation count is large and statistically managed. **If you must loop** (e.g. iterating a collection that is the thing under test, or amortizing an unavoidable fixed cost): - **Consume every iteration's result** with `Blackhole.consume()` so none is eliminated. - **Vary the data** per iteration by reading from a `@State` array/structure, so iterations are not identical and cannot be folded into one. - **Avoid invariant work** in the body that the JIT will hoist. - **Account for the count.** JMH times the whole method; if the method intentionally performs a fixed number of operations, annotate with `@OperationsPerInvocation(N)` so JMH normalizes the score to per-operation. ## Worked contrast ```java // RISKY: identical work each iteration -> may be folded; result discarded -> DCE. @Benchmark public void bad() { double s = 0; for (int i = 0; i < 1000; i++) s += Math.sqrt(2.0); // same input every time } // BETTER: vary inputs from @State, consume each result, declare the count. @State(Scope.Thread) public static class Data { double[] xs = /* filled with varied values */; } @Benchmark @OperationsPerInvocation(1000) public void good(Data d, Blackhole bh) { for (double x : d.xs) bh.consume(Math.sqrt(x)); } ``` The 'better' version still has a loop, but each iteration uses a different value (no folding), each result is consumed (no DCE), and `@OperationsPerInvocation` makes the reported score per-element. ## Mental model Every hazard here is the same family as DCE and constant folding: the JIT is allowed to remove or simplify work whose repetition or result is provably unnecessary. A loop multiplies the opportunities. Minimize loops, vary data, consume results, and declare counts.
- You loop 1000 times over Math.sqrt(2.0) and consume the sum. Inputs vary? No. What can still go wrong?Every iteration computes the same value from the same constant, so the JIT can fold them to a single computation and the per-op time is ~1000x too small. You must vary the input (e.g. read from a @State array) so iterations differ.
- When is a manual loop legitimately the right choice?When the structure under test is itself a loop/collection (you are measuring iteration), or when a fixed setup cost must be amortized over many ops. Then consume each result, vary the data, and use @OperationsPerInvocation to normalize the score.
saying these in an interview costs you the question
- Adding a loop of identical work and dividing by N, trusting the per-op number
- Discarding loop results or keeping only the last one
- Looping over the same constant input every iteration
- Forgetting @OperationsPerInvocation when the method does a fixed internal count, then misreading the score