Inside HotSpot, the C2 JIT compiler is allowed to move, merge and delete field accesses in Java code. Name the C2 optimizations that actually remove or relocate memory operations, explain what the Java Memory Model lets C2 assume about a plain (non-volatile) shared field between two synchronization actions, and explain why the same source code can behave differently once tiered compilation promotes it from the interpreter to C1 and then to C2.
answer
- DRF assumption → plain field = local between anchors
- hoist / redundant-load / sink / dead-store
- EA + scalar replacement = no memory op to fence
- inlining is the enabler, profile decides
- tier 0 interpreter hides it, OSR at tier 4 exposes it
basics
~20 sC2 hoists loop-invariant loads into registers, eliminates redundant loads, sinks or deletes stores, and via escape analysis scalar-replaces objects so their fields have no memory operations at all. The Java Memory Model lets it treat plain fields as thread-local between synchronization actions. Tiered compilation applies this only once code is hot, so racy code changes behaviour after warm-up.
solid answer
~60 sThe Java Memory Model guarantees sequential consistency only for **data-race-free** programs, so C2 compiles as if the program is race-free: between two synchronization actions in a thread — monitorenter/exit, a volatile or VarHandle acquire/release access, `Thread.start`/`join` — it assumes no other thread writes the plain fields it touches, and treats them like locals. That licence enables: **loop-invariant load hoisting / register promotion** (the classic spin-on-a-plain-flag hang), **redundant-load elimination** and store-to-load forwarding, **dead-store elimination and store sinking** (a repeated write is collapsed to one after the loop), aggressive **inlining** that exposes all of the above across method boundaries, and **escape analysis with scalar replacement**, where a non-escaping object is never allocated and its fields become SSA values in registers — there are no memory operations left to order, so no fence can affect them. Related: lock elision and lock coarsening. Tiered compilation means the interpreter and C1 perform real accesses in order, so the race is invisible until tier 4 (or an OSR compile of a hot loop) kicks in — and a deoptimization can hide it again.
code
java · 14 linesclass Spinner {
private boolean stop; // plain field, no synchronization
void run() {
while (!stop) { // loop-invariant read: C2 may hoist it
work();
}
}
void requestStop() { stop = true; }
}
// what C2 is permitted to compile it into:
// boolean t = stop; // one load, kept in a register
// if (!t) { for (;;) work(); } // never re-reads the fieldgo deeper
Know that the JIT may keep a field in a register or delete repeated reads, and that this is why a plain boolean flag can fail to stop a loop.
Name concrete transformations — loop-invariant hoisting, redundant-load elimination, dead-store elimination — and explain that the interpreter performs real accesses so the bug only shows after warm-up.
Tie the optimizations to the JMM's data-race-free assumption and to synchronization actions as anchors, include escape analysis/scalar replacement and its 'no memory operation to order' consequence, and describe reproducing the bug with tiering flags and OSR.
Frame it as a testing and verification problem: correctness cannot be established empirically at tier 0, so the team needs explicit ordering contracts (volatile/VarHandle/concurrent utilities), warm-up-aware stress harnesses, and review rules that treat 'adding a log line fixed it' as evidence of an unfixed race.
## What the JMM lets C2 assume The Java Memory Model's central promise is conditional: a program whose executions contain **no data races** behaves as if all its actions were interleaved in one sequentially consistent order. A data race is two accesses to the same non-volatile field, at least one a write, not ordered by happens-before. HotSpot's C2 compiler is written to exploit exactly that promise — it optimizes *as if* the program is data-race-free, because a program that isn't gets no strong guarantee anyway. Concretely: C2 does not consult happens-before at each individual field access. Instead, **synchronization actions anchor memory**. A monitorenter/monitorexit, a volatile or `VarHandle` acquire/release/opaque access, a `Thread.start()` or `join()` — these appear in C2's ideal graph as nodes that pin the memory state, and loads and stores cannot be freely moved across them. Everything *between* two such anchors is a synchronization-free region, and inside it C2 treats a plain shared field exactly like a local variable: read it once, keep it in a register, delete reads it can prove redundant, delete or move stores. The absence of synchronization is read as a promise from the programmer that nobody else is touching those fields. ## The optimizations that move or delete memory operations **Loop-invariant load hoisting / register promotion.** If a field read inside a loop cannot be changed by anything the loop does, C2 hoists it above the loop and keeps the value in a register. `while (!stop) { work(); }` becomes, in effect, `if (!stop) while (true) work();` — the canonical never-terminating spin on a plain `boolean`. **Redundant-load elimination and store-to-load forwarding.** Two reads of the same field with no intervening anchor collapse into one; a read that follows a write to the same field reuses the written value instead of reloading it. Either way, a write performed by another thread in between is simply never observed. **Dead-store elimination and store sinking.** A store overwritten later with no intervening synchronization is deleted. A store repeated in a loop can be sunk past the loop, so instead of a value another thread could watch changing, one write lands at the very end. **Inlining.** By itself inlining moves nothing, but it is the enabler: an accessor's field read only becomes loop-invariant *after* the accessor is inlined into the loop. C2 inlines using class-hierarchy analysis and the profile collected at tier 3, so whether a call site is monomorphic — which depends on what the application happened to execute earlier — decides whether the rest of these optimizations fire at all. **Escape analysis and scalar replacement.** If C2 proves an object does not escape its allocating method (or thread), the allocation is removed and its fields become plain SSA values living in registers. This is the sharpest case for the memory model: a scalar-replaced object has **no memory operations at all**, so there is nothing for a barrier to order and nothing another thread could observe even in principle. Related transformations on the same analysis are **lock elision** (`-XX:+EliminateLocks`) for locks on non-escaping objects and **lock coarsening**, which merges adjacent synchronized blocks on the same monitor — legal because it only widens a critical section. ## Why tiered compilation changes what you observe HotSpot executes the same bytecode at several tiers. **Tier 0**, the interpreter, performs every `getfield` and `putfield` as a genuine memory access in program order — none of the above happens, so a racy program looks perfectly correct. **Tiers 1–3** are C1: fast compilation, limited optimization, and at tier 3 additional profiling instrumentation. **Tier 4** is C2, where the full optimization set above applies. So the observable interleavings of a racy program are a function of how hot the code is. A method promoted after thousands of invocations starts behaving differently mid-run. Long-running loops matter even more: **on-stack replacement (OSR)** compiles a loop *while it is executing* and swaps the running frame over, which is precisely why a spin loop runs correctly for a second or two and then hangs forever. The reverse also happens — an **uncommon trap** deoptimizes back to the interpreter, the field is reloaded from memory, and the bug appears to fix itself. This is why racy code passes tests. Short unit tests never reach tier 4; a JMH benchmark with proper warm-up, or production after a few minutes, does. It also explains the folklore fixes: adding a `println`, a `synchronized` block, or an unrelated volatile read introduces an anchor or an unanalyzable call, defeating the hoisting. Nothing was made correct — the optimization was merely suppressed, and the next inlining decision can bring it back. ## How to say it in an interview Lead with the licence (DRF assumption, plain fields are locals between anchors), name three or four concrete C2 transformations including scalar replacement's "no memory operation exists" point, then explain tiering and OSR as the reason correctness is a function of warm-up. Close with the operational consequence: you cannot test your way to memory-model correctness at tier 0.
- A racy spin loop passes every unit test but hangs in production after a few minutes. What would you change about the test to reproduce it?Make the code hot enough to reach C2. Run the loop through many iterations or invocations with a real warm-up phase — a JMH harness or a long-running loop that triggers on-stack replacement — rather than a short test that never leaves the interpreter. Alternatively force compilation with -XX:-TieredCompilation or lower the compile thresholds, and check -XX:+PrintCompilation to confirm the method actually reached tier 4. The permanent fix is still an ordering edge on the flag, not a test setting.
- Adding a synchronized block or a logging call to the loop body makes the hang disappear. Is the code now correct?No. A synchronized block introduces a synchronization action that anchors memory state, and a call C2 cannot fully analyze or inline blocks the load from being hoisted — both suppress the optimization rather than establishing the happens-before edge you actually need. A later change to inlining decisions, a different profile, or a different JDK can re-enable the hoist. Correctness requires making the flag volatile or accessing it through a VarHandle with the ordering you want.
- If escape analysis scalar-replaces an object, can another thread ever see a torn or stale value of its fields?No — a scalar-replaced object was proven not to escape the allocating method or thread, so no other thread has a reference through which to observe it, and after the transformation its fields are registers with no heap memory operations at all. That is why the question of ordering does not arise: barriers order memory accesses, and there are none. If the object did escape, escape analysis would not apply the transformation in the first place.
C2 treats the gaps between synchronization actions like a private office: since you never announced anyone else was coming in, it feels free to leave the paperwork on the desk (in a register) instead of filing it in the shared cabinet after every edit.
saying these in an interview costs you the question
- Saying the JIT 'caches the field in the CPU cache' — caches are coherent; the value is being held in a register, or the load is deleted outright
- Believing code that behaves correctly during a short test is correct, rather than merely still at tier 0/C1
- Claiming a volatile read elsewhere in the method, or a println, 'fixes' the race, when it only suppresses one optimization
- Thinking escape analysis merely reorders the object's field writes, instead of removing the allocation and its memory operations entirely
- Assuming C2 checks happens-before at every field access, rather than assuming data-race freedom between synchronization anchors