skip to content

Where does HotSpot actually place memory barriers around a read and a write of a volatile field, and how does the machine code it emits differ between x86-64 and AArch64?

level: seniorimportance: should knowfreq 26%

answer

  1. read = acquire: LoadLoad + LoadStore after
  2. write = release + trailing StoreLoad
  3. x86-64: read is a plain mov, write adds lock addl
  4. AArch64: ldar / stlr, dmb ish only if needed
  5. barrier constrains the JIT even when no instruction is emitted

basics

~20 s

Conceptually: LoadLoad+LoadStore after a volatile read; StoreStore+LoadStore before a volatile write and StoreLoad after it. On x86-64 only the trailing StoreLoad costs an instruction, usually a locked stack operation. On AArch64 HotSpot uses ldar/stlr acquire/release instructions instead.

solid answer

~60 s

The conceptual placement HotSpot follows is: - **volatile read:** the load, then LoadLoad and LoadStore barriers - so nothing after the read may be hoisted above it. This is an acquire. - **volatile write:** StoreStore and LoadStore barriers, then the store, then a StoreLoad barrier. The leading pair makes it a release; the trailing StoreLoad is what gives Java volatile its total order over volatile accesses, so a volatile write followed by a volatile read of a *different* field cannot appear reordered. What reaches machine code depends on the target. On **x86-64** the hardware already provides LoadLoad, LoadStore and StoreStore, so a volatile read is an ordinary `mov` plus compiler-level ordering constraints, and a volatile write is a `mov` followed by a locked instruction - typically `lock addl $0,(%rsp)` - which drains the store buffer more cheaply than `mfence` on many microarchitectures. On **AArch64** HotSpot emits `ldar` for the read and `stlr` for the write: one-way acquire and release loads and stores, with `dmb ish` used where a full fence is genuinely required. In every case the JIT must also honour the ordering itself - the barrier is as much a compiler constraint as a hardware instruction.

code

text · 12 lines
text
conceptual:                 x86-64:              AArch64:

volatile read               volatile read        volatile read
  load                        mov                  ldar
  LoadLoad                    (free)
  LoadStore                   (free)

volatile write              volatile write       volatile write
  StoreStore                  (free)               stlr
  LoadStore                   (free)
  store                       mov
  StoreLoad                   lock addl $0,(%rsp)

go deeper

for a junior

Not expected in detail; know that volatile reads and writes constrain reordering and that a write costs more than a read.

for a middle

Give the conceptual barrier placement around volatile reads and writes and know that x86-64 needs a real instruction only on the write side.

for a senior

Name the actual instructions on both architectures, explain the trailing StoreLoad's purpose, and stress that the JIT constraint exists even when no instruction is emitted.

for a principal

Discuss the portability consequence - profiles differ between x86-64 and AArch64 - and how that affects benchmarking policy and the choice of ordering strength in shared library code.

## Two constraints, not one A barrier in the JVM is two things at once: a constraint on the JIT, which may not move memory operations across it, and, where the hardware needs it, an emitted instruction. Forgetting the compiler half leads people to conclude that volatile is free on x86-64 because a volatile read compiles to a plain `mov`. It is not free: the JIT must stop hoisting that read out of loops and stop eliminating it as redundant, which can be the more expensive half in a hot loop. ## The conceptual placement HotSpot's implementation notes describe the required barriers around volatile accesses in the four-barrier vocabulary: ``` // volatile read load x LoadLoad LoadStore // volatile write StoreStore LoadStore store x StoreLoad ``` The read side is a pure acquire: subsequent loads and stores cannot rise above it. The write side is a release (the leading pair) *plus* a trailing StoreLoad. That trailing barrier is what distinguishes Java volatile from a plain release store. Without it, a volatile write followed by a volatile read of a different field could appear reordered, and the model's requirement that volatile accesses have a single total order consistent across threads would break - which is exactly the guarantee Dekker-style algorithms and many lock-free algorithms depend on. ## What x86-64 actually gets x86-64 has a total-store-order model: it never reorders load-load, load-store or store-store. Therefore: - A **volatile read** emits no fence at all - just the load, with the JIT forbidden to move things across it. - A **volatile write** emits the store plus one instruction to satisfy StoreLoad. HotSpot typically uses `lock addl $0x0,(%rsp)` - a locked no-op arithmetic operation on the top of the stack - rather than `mfence`, because a locked operation has the required drain effect and has historically been faster on Intel and AMD parts. So on x86-64 the asymmetry is stark: volatile reads are nearly free, volatile writes carry a real stall. That asymmetry is why read-mostly volatile fields are cheap and write-heavy ones are not. ## What AArch64 gets AArch64 is weakly ordered and provides nothing for free, but it has dedicated one-way instructions: - `ldar` - load-acquire, used for a volatile read. - `stlr` - store-release, used for a volatile write. These express exactly the acquire and release constraints, and the architecture's rules for `stlr` followed by `ldar` also provide the ordering Java's volatile requires, so a separate full fence is usually unnecessary. Where a full fence is needed, `dmb ish` is emitted. The practical effect is that on AArch64 volatile *reads* cost more than on x86-64 (they are a distinct instruction with ordering constraints, not a plain load), while volatile writes are often cheaper than a full x86-style drain. Code tuned on x86-64 can therefore have a noticeably different profile on AArch64 hardware. ## Verifying it yourself HotSpot can print the generated assembly, which is the only way to settle arguments about what is emitted for a given field on a given machine. With the disassembler plug-in available, run with the diagnostic unlock flag plus print-assembly, optionally restricted to one method. The listing will show the load or store and any fence instruction next to it. Reading that output is a legitimate senior-level skill and a good follow-up if the interviewer pushes. ## Why the details matter operationally - The cost of the same source code differs by architecture, so a benchmark on one is not evidence for the other. - Volatile read cost is dominated by lost compiler optimizations, not by fences - a volatile field read in a tight loop cannot be cached in a register, which can cost more than any instruction. - The write-side StoreLoad is the reason a frequently written volatile counter scales poorly, quite apart from cache-line contention on the field itself. ## Interview framing Give the conceptual placement first, in the barrier vocabulary, and only then say what each architecture reduces it to. Emphasize that the JIT constraint exists on every architecture even when zero instructions are emitted - that is the part candidates most often miss.

  • On x86-64 a volatile read compiles to an ordinary mov with no fence. Does that mean volatile reads are free?
    No. The barrier is also a constraint on the JIT: it may not hoist the read out of a loop, cache it in a register, or eliminate it as redundant. In a tight loop those lost optimizations usually cost far more than any fence instruction would, so the read is cheap at the hardware level and not free at the compiler level.
  • Why does the volatile write need a trailing StoreLoad when a plain release store does not?
    Java's volatile guarantees that all volatile accesses appear in a single total order consistent with program order in each thread. Without the trailing StoreLoad, a volatile write followed by a volatile read of a different field could appear reordered, breaking that total order and with it algorithms such as Dekker's. A plain release store makes no such promise, which is why it is cheaper.

saying these in an interview costs you the question

  • Saying volatile emits a fence on both the read and write side on x86-64
  • Claiming the barrier only affects hardware and not the JIT
  • Assuming the same volatile field costs the same on x86-64 and AArch64
  • Treating specific instruction choices as required by the specification rather than implementation details

context