HotSpot's C2 compiler has recognised a hot int-counted loop over an array and decided to unroll it. Describe the shape the compiled code actually takes — how the iteration space gets split up, what property the surviving hot loop's trip count must have, where SIMD (vector) instructions come from, and how safepoint polling is kept bounded inside such a loop. Then name the factors that cap how large an unroll factor the compiler will pick.
answer
- pre-loop / main loop / post-loop
- main trip count = multiple of unroll factor, no range checks
- pre-loop handles alignment + range-check split
- superword packs unrolled copies into SIMD
- strip mining: ~1000-iteration chunks, poll on outer back-edge
basics
~20 sC2 splits the iteration space into pre-loop, main loop and post-loop. The main loop is the unrolled one: its trip count is a multiple of the unroll factor, its bounds checks are gone, and the superword pass fuses its copies into SIMD instructions. Strip mining wraps it so safepoints still get polled.
solid answer
~60 sC2 does not emit one loop with a duplicated body; it emits **three loops**. - **Pre-loop** — runs a dynamically chosen handful of iterations. It absorbs alignment for vector memory access and, together with C2's range-check predication, is adjusted so the main loop is provably in bounds. - **Main loop** — the optimized one. Its trip count is arranged to be an exact multiple of the unroll factor, so no remainder test lives in the hot path, and it carries no array bounds checks. C2 unrolls it by repeated doubling (2, 4, 8, 16…). - **Post-loop** — mops up the remaining iterations. SIMD is not the unroller's doing: the **superword** pass runs over the unrolled main loop and packs the now-adjacent, independent, identically-shaped operations into vector instructions. That is usually where the speedup actually comes from. Because an int-counted loop has no safepoint poll on its back-edge, C2 applies **loop strip mining**: the optimized inner loop is executed in chunks (`LoopStripMiningIter`, default 1000) by an outer loop that polls between chunks. The factor is capped by the node-count budget (`LoopUnrollLimit`), non-inlined calls in the body, loop-carried dependences, vector register width, and small or unknown profiled trip counts.
code
text · 20 lines// source: for (int i = 0; i < n; i++) a[i] += b[i];
// compiled shape (schematic):
i = 0;
// pre-loop: fix alignment, set up the range-check-free window
while (i < preLimit) { a[i] += b[i]; i++; }
// main loop: unrolled x8, no bounds checks, superword-vectorised,
// (mainLimit - i) is an exact multiple of 8, run in strip-mined chunks
outer: while (i < mainLimit) {
stripEnd = min(i + 1000, mainLimit);
while (i < stripEnd) { // <- vector add of 8 lanes, no poll
vadd(a+i, b+i);
i += 8;
}
safepoint_poll(); // <- poll lives on the OUTER back-edge
}
// post-loop: leftover iterations
while (i < n) { a[i] += b[i]; i++; }go deeper
Know the vocabulary: the JIT can turn one hot loop into a pre-loop, a fast main loop, and a post-loop, and the fast one may use SIMD instructions. Being able to name the three parts is enough at this level.
Explain why the split exists — the main loop's trip count must be a multiple of the unroll factor and its bounds checks are gone, so nothing administrative sits in the hot path — and that superword turns the adjacent copies into vector instructions.
Add the operational layer: counted loops carry no back-edge safepoint poll, strip mining chunks the inner loop so time-to-safepoint stays bounded, and be able to list the real caps (node budget, non-inlined calls, dependences, vector width, small trip counts) and how you would confirm the shape from disassembly.
Frame it as a budget question: unrolling competes with inlining for compile-time and code-cache resources, code bloat costs instruction-cache locality across the whole service, and strip-mining chunk size trades throughput against tail latency. Decide when data-layout or algorithm changes (removing dependences, contiguous access) are worth more than any JIT tuning flag, and treat unroll flags as diagnostic instruments rather than production settings.
## The compiled shape is three loops, not one When C2 decides a counted loop is hot enough to optimize aggressively, the machine code it emits does not look like the source loop with a fatter body. The iteration space is **split into three consecutive loops**: 1. a **pre-loop**, which runs a small number of iterations decided at run time; 2. a **main loop**, which is the one that gets unrolled, range-check-free and vectorised, and which executes the overwhelming majority of the iterations for a large array; 3. a **post-loop**, which finishes whatever iterations remain. A "counted loop" here means the shape C2 can reason about: an `int` induction variable, a constant stride, an invariant limit, and a single back-edge. If the loop is not in that shape, none of the following happens at all. ## Why the split buys the main loop its properties **Trip count divisibility.** A body unrolled by N can only be entered when at least N iterations remain; otherwise the copies would run past the end. The naive fix is a remainder check inside the loop, which would put a branch right back into the hot path. Instead C2 arranges the *main* loop's trip count to be an exact multiple of N and pushes the awkward leftovers to the post-loop. The hot path therefore contains one compare and one back-branch per N elements and nothing else administrative. **Bounds checks.** C2 does not remove array range checks from the main loop by proving each access individually; it *shapes the loop so the proof is trivial*. Through range-check predication and iteration splitting, the pre-loop's exit condition and the main loop's limit are adjusted so every access in the main loop is provably within bounds, and the checks are hoisted out as guards. If a guard fails at run time, the code deoptimizes rather than faulting. **Alignment.** Vector loads and stores prefer (and on some hardware historically required) suitably aligned addresses. The pre-loop's second job is to advance the base offset until the main loop's memory accesses start on a vector-friendly boundary. So the main loop body ends up as pure straight-line replicated work: no remainder test, no bounds checks, no alignment fix-ups. ## Where the vector instructions come from Unrolling by itself only produces N adjacent scalar copies. The transformation that turns them into SIMD is a separate pass, **superword-level parallelism (SuperWord / auto-vectorisation)**. It scans the unrolled main loop for groups of operations that are (a) independent of one another, (b) the same operation on the same element type, and (c) on contiguous memory, and packs each group into one vector instruction operating on several lanes at once — e.g. eight `int` adds becoming a single 256-bit add on AVX2 hardware. This is why the unroll factor and the vector width are related: C2 will typically aim at an unroll factor that fills the widest usable vector register for the element type, since unrolling far past that adds code size without adding lanes. It is also why anything that breaks pattern (a) — a loop-carried dependence, a mixed set of operations, a strided or indirect access — can leave you with an unrolled but purely scalar main loop. ## Strip mining: keeping the optimized loop safepoint-friendly HotSpot deliberately emits **no safepoint poll on the back-edge of an int-counted loop**; a poll in the middle of the body would fragment the region the optimizer is working on and block unrolling and vectorisation. The cost is that a thread grinding through millions of iterations never reaches a safepoint, so every global operation that needs one — a stop-the-world GC phase, a deoptimization, a thread dump — stalls behind it. Time-to-safepoint, not GC pause time, becomes the latency outlier. **Loop strip mining** resolves this without giving anything up. C2 wraps the fully optimized inner loop in an outer loop that executes it in bounded chunks — `-XX:LoopStripMiningIter`, default 1000 iterations — and places the safepoint poll on the *outer* loop's back-edge. The inner loop keeps its unrolled, check-free, vectorised form; the outer loop caps how long a thread can go without polling. ## What bounds the unroll factor - **Node-count budget.** C2 unrolls by repeated doubling and stops when the unrolled body's intermediate-representation node count would exceed its limit (`-XX:LoopUnrollLimit`). Oversized bodies waste instruction cache and consume compile budget better spent on inlining. - **Non-inlined calls.** A call that was not inlined terminates the straight-line region, so replicating the body gains nothing schedulable and blocks vectorisation outright. - **Loop-carried dependences.** If iteration *i* reads what *i-1* wrote, the copies are not independent: superword bails out, and only the modest branch/counter savings remain. - **Vector width.** Beyond the point where one iteration of the main loop fills the widest available vector register per operation, extra unrolling mostly adds code size. - **Short or unknown trip counts.** Profile data showing a handful of iterations makes the pre/post-loop setup a net loss, so C2 declines. - **Loop shape.** Non-counted loops — iterator-driven, recomputed bounds, an induction variable mutated on some paths, `long` counters on older JDKs — never reach this pipeline. ## Observing it The honest way to confirm the shape is to disassemble the compiled method (`-XX:+UnlockDiagnosticVMOptions -XX:+PrintAssembly` with hsdis) and look for the three loop bodies plus vector opcodes in the middle one — not to infer it from a wall-clock number.
- Why does C2 bother with a pre-loop at all, instead of starting the unrolled main loop at index 0?The pre-loop absorbs the two things that would otherwise force per-iteration work inside the hot loop. It advances the index until the main loop's memory accesses are on a vector-friendly alignment boundary, and its exit condition is tuned so that every access in the main loop is provably in bounds, letting the range checks be hoisted into guards. Starting the unrolled loop at zero would mean either keeping the checks or emitting misaligned vector accesses.
- You benchmark a hot array loop and see the main loop unrolled in the disassembly but no vector instructions. What are the likely causes?Unrolling succeeded but the superword pass bailed out. The usual reasons are a loop-carried dependence (iteration i reading what i-1 wrote), non-contiguous or indirect access such as a[idx[i]], mixed or unsupported operation types in the body, a conditional inside the body that could not be turned into a branchless select, or an element type with no vector support on that CPU. A non-inlined call would normally have blocked the unroll too.
- If strip mining exists to restore safepoint polls, what actually goes wrong on a system running with -XX:-UseCountedLoopSafepoints?The optimized counted loop has no poll on its back-edge, so a thread inside it cannot reach a safepoint until the loop finishes. Any global operation that requires all threads to stop — a stop-the-world GC phase, a deoptimization, a thread dump — waits for the slowest such thread. You see long time-to-safepoint numbers in the safepoint log while GC pause time itself looks fine, which is a classic misdiagnosis.
saying these in an interview costs you the question
- Saying unrolling itself emits the SIMD instructions — vectorisation is a separate superword pass over the already-unrolled main loop
- Believing the compiled code is one loop with a remainder branch inside it, rather than three loops with the remainder pushed to the post-loop
- Claiming the main loop still performs array bounds checks that are merely 'cheap' — in the main loop they are gone, replaced by hoisted guards plus deoptimization
- Thinking the safepoint poll sits on the inner unrolled loop's back-edge; under strip mining it sits on the outer chunking loop's back-edge
- Assuming bigger unroll factors are always better, ignoring the node-count budget, instruction-cache cost, and vector-register width
- Asserting a loop with a non-inlined call in the body still gets unrolled and vectorised