skip to content

A Java field declared volatile is read many times inside a hot loop and written only rarely. Break down what that ordered access actually costs the running JVM, and say which part of the cost usually dominates in a real profile.

level: seniorimportance: nice to knowfreq 26%

answer

  1. Three costs: emitted fences, blocked JIT work, coherence traffic
  2. Fence cost is per-ISA: x86-64 pays on the write, AArch64 on the acquire load
  3. Ordered read = compiler barrier: no register promotion, no CSE, no vectorization
  4. Read once into a local: same guarantees, plain-code loop body
  5. Coherence miss shows as LLC misses on readers, not a stall on the store

basics

~20 s

Three separate costs: the fence instructions emitted (architecture-dependent), the JIT optimizations the ordered read forbids, and cache-coherence traffic on the field's line. For read-mostly fields the blocked JIT work usually dominates, and hoisting one read into a local recovers most of it.

solid answer

~60 s

I split it into three costs that behave differently. 1. **Emitted fences.** Ordering strength maps to real instructions, and the mapping is per-architecture: on x86-64 the read side of a volatile costs almost nothing at the hardware level while the write side carries the store-buffer drain; on AArch64 an acquire load is a real instruction and the release store is comparatively cheap. Java's access modes (plain, opaque, acquire/release, volatile, via VarHandle since Java 9) sit on that ladder. Exactly where HotSpot places those fences is its own topic. 2. **Lost JIT optimization, usually the dominant term for reads.** An ordered read is a compiler barrier too: C2 cannot promote the field to a register across iterations, cannot collapse redundant loads, and the loop typically will not vectorize. Ten reads per iteration means ten real loads. 3. **Coherence traffic.** Each rare write invalidates the line in every reader's cache, so reads turn into cache misses. That appears as memory stalls, not fence stalls. The cheap first move is reading the field once into a local, which changes nothing about the field's declared ordering.

code

java · 21 lines
java
private volatile boolean enabled;   // still volatile for every other thread

// Before: an ordered read per iteration.
// C2 cannot keep `enabled` in a register, cannot collapse the loads,
// and the array work will not vectorize.
void processSlow(int[] a) {
    for (int i = 0; i < a.length; i++) {
        if (enabled) a[i] = a[i] * 2 + 1;
    }
}

// After: one ordered read, snapshot semantics for the loop.
// The loop body now reads a local; register promotion and
// vectorization are available again.
void processFast(int[] a) {
    final boolean on = enabled;   // single acquire-strength read
    if (!on) return;
    for (int i = 0; i < a.length; i++) {
        a[i] = a[i] * 2 + 1;
    }
}

go deeper

for a junior

Know that a volatile field is re-read from memory every time it is accessed, and that copying it into a local variable before a loop is the standard cheap fix.

for a middle

Separate the hardware fence from the compiler restriction, and be able to name what the JIT loses: register promotion across iterations, redundant-load elimination, and vectorization.

for a senior

Rank the three costs for a read-mostly field and read a profile well enough to say which one is firing: fence cycles on the store, inflated load counts in compiled code, or LLC misses spread across readers.

for a principal

Frame the fix as structural rather than as ordering strength: read granularity, data layout and padding, sharding with aggregation on read, and write frequency, while accounting for the fact that the fence term differs per deployment architecture.

## "Ordered access is expensive" is three different costs When a Java field is declared `volatile` (or accessed through a `VarHandle` with an ordered access mode), engineers usually say "that costs a barrier" and stop. On a field that is read constantly and written rarely, the emitted barrier is frequently the smallest of three independent costs. Naming all three, and knowing which one a given profile is showing, is the whole skill here. ## Cost 1: the fences the JVM emits, which is a per-architecture number Java since 9 exposes a ladder of access modes through `VarHandle`: plain, opaque, acquire/release, and volatile, in increasing ordering strength. Each mode compiles to some combination of hardware ordering instructions, and the compiled cost is not a property of the mode alone but of the mode *and the target ISA*. On x86-64 the hardware already provides load-load, load-store and store-store ordering, so an ordered *read* needs no extra instruction; the cost sits on the *write* side, where the store buffer must be drained before subsequent loads may proceed. On AArch64 the split is different: a load-acquire is a genuine instruction with a genuine cost, while a release store is comparatively cheap. (Precisely which fences HotSpot emits around a volatile read and a volatile write on each architecture is covered separately; the point for costing is only that the number is asymmetric and does not transfer between architectures.) The practical consequence for a read-mostly field: whatever the write side costs, it is amortized over an enormous number of reads. Optimizing that term is usually optimizing the wrong one. ## Cost 2: the optimizations the JIT is no longer permitted to perform This is the term people miss, and it is usually the largest for hot reads. An ordered access constrains the *compiler* as well as the processor, and C2's loop optimizations are exactly the ones it constrains. **Register promotion is blocked.** For a plain field, C2 can prove nothing else in the loop writes it, load it once into a register, and use the register for every iteration. An ordered read must actually observe memory each time it executes, so the value cannot live in a register across iterations. A loop that reads the field ten times per iteration issues ten loads per iteration, forever. **Redundant-read elimination and common-subexpression elimination are blocked.** Two plain reads of the same field with no intervening write collapse into one. Two ordered reads may legally return different values (another thread may have written between them), so they cannot be collapsed, and expressions derived from them cannot be shared. **Loop transforms including vectorization are blocked.** Superword/vectorization requires the compiler to reorder and batch memory operations across iterations. An ordered access is a scheduling barrier inside the loop body, so unrolling and vectorization of the surrounding array work commonly fail to apply. The visible symptom is a loop that runs several times slower than an equivalent loop over plain data, with no fence stall anywhere in the profile. **The read-once-into-local pattern.** If the algorithm is happy with a single consistent snapshot per call or per loop, copying the field into a local variable before the loop removes every one of these losses. One ordered read happens, with full ordering semantics and full visibility for the values it publishes; the loop body then reads a local and optimizes exactly like plain code. Crucially, this does not weaken the field for anyone: other threads still see a `volatile` field with unchanged guarantees; only this thread's reading frequency changed. It is the highest-yield, lowest-risk move available, and it is why "the volatile is slow" so often turns out to be "the loop re-reads it." ## Cost 3: coherence traffic, and telling it apart from fence cost The third cost has nothing to do with ordering strength at all. Cache coherence works per cache line. When one core writes the line, every other core's copy is invalidated, and each of their next reads is a coherence miss served from a remote cache or memory. A field that is written even occasionally while dozens of cores read it can therefore be an order of magnitude more expensive than any fence. False sharing is the same mechanism with none of the sharing: an unrelated counter that happens to sit on the same 64-byte line produces identical invalidations. In a real Java profile the three look different. Fence cost shows as time attributed to the volatile store instruction itself. Blocked JIT optimization shows as *instruction count*: the same method with many more loads and no vector instructions, visible by comparing the compiled assembly or the profile's cycles-per-iteration against a plain-field variant. Coherence traffic shows as last-level cache misses and memory stalls on the load, spread across all the reading threads rather than concentrated on the writer. `@Contended` padding, per-thread or per-stripe sharding with aggregation on read, and simply writing less often (batching, sampling, write-on-change) address that cost; changing the access mode does not. ## Putting them together For the read-mostly hot field in the question, the ranking is usually: blocked JIT optimization first, coherence traffic second when the write rate or the reader count is high, emitted fences last. That ordering is why the productive interventions are structural, reading once into a local and fixing layout, rather than reaching down the access-mode ladder.

  • Does copying a volatile field into a local variable before a hot loop weaken the guarantees other threads rely on?
    No. The field remains volatile, and every other thread's reads and writes keep their full ordering and visibility semantics. The only change is that this thread performs one ordered read instead of many and then works from a snapshot. The trade-off is purely staleness within this thread: if the algorithm needs to observe a mid-loop change (a cancellation flag checked per chunk, say), you re-read at the granularity you actually need rather than every iteration.
  • Why can a benchmark of ordered-access cost on x86-64 mislead you about AArch64?
    The two architectures pay in different places. x86-64 already guarantees load-load, load-store and store-store ordering, so an ordered read needs no extra instruction and the cost concentrates on the write side's store-buffer drain. AArch64 is weaker, so a load-acquire is a real instruction with a real cost while the release store is comparatively cheap. A read-mostly workload therefore looks nearly free on x86-64 and measurably not free on AArch64, from the same bytecode.
  • A loop over a primitive array runs several times slower than expected, and the profile shows no fence stalls and no cache misses. What would you look for?
    Blocked compiler optimization rather than hardware cost. I would check whether an ordered access sits inside the loop body acting as a scheduling barrier, then compare the compiled code against a plain-field variant: the tell is a much higher load count per iteration and the absence of vector instructions, because superword optimization cannot batch memory operations across an ordered access.

A volatile read is like a rule that every glance at a shared whiteboard must be a fresh walk to the room. The walk itself is short; what hurts is that you are forbidden from writing the value on your own notepad, so you make the trip ten times a loop.

saying these in an interview costs you the question

  • Saying a volatile read costs a fence stall on x86-64, where the read side emits no extra ordering instruction
  • Assuming the JIT can hoist a volatile read out of a loop the way it hoists a plain field read
  • Blaming the fence for a slowdown that is actually cache-line invalidation or false sharing
  • Claiming that reading a volatile field into a local weakens its guarantees for other threads
  • Treating an ordered-access benchmark from one architecture as a valid number for another

context