On a 64-bit HotSpot JVM running on x86-64, is a plain non-volatile `long` field ever actually torn? Explain the gap between what the hardware does and what the Java Language Specification permits, and when that gap still matters.
answer
- x86-64 aligned 8-byte mov = atomic in practice
- Spec still permits tearing (targets 32-bit words)
- 32-bit HotSpot: plain long split; volatile used cmpxchg8b/SSE2
- Unaligned off-heap 64-bit access is not atomic anywhere
- Visibility usually forces volatile regardless of tearing
basics
~20 sIn practice no: 64-bit HotSpot on x86-64 emits one naturally aligned 8-byte load or store, which the hardware performs atomically. The specification still permits tearing, so relying on the hardware is unportable, and you usually need volatile for visibility and ordering anyway.
solid answer
~50 sOn 64-bit HotSpot on x86-64, a `long` or `double` field is naturally aligned and accessed with a single 8-byte `mov`. x86-64 guarantees aligned accesses up to 8 bytes are atomic, so no tearing occurs — and you will never reproduce it in a test on that platform. The specification nonetheless leaves plain 64-bit accesses non-atomic (JLS §17.7), because it must also cover implementations whose natural word is 32 bits. On the now-removed 32-bit HotSpot x86 port, a plain `long` store really could be two 4-byte stores, while `volatile long` was implemented with an 8-byte SSE2 move or `cmpxchg8b` precisely to keep it atomic. 32-bit ARM had the analogous story with `ldrexd`/`strexd`. So the gap matters in three ways: **portability** (you are writing to the specification, not to one CPU), **alignment** (unaligned 64-bit access through `ByteBuffer` or off-heap memory is not atomic on any platform), and **the fact that atomicity was rarely the reason you wanted `volatile`** — visibility and ordering usually force the same annotation regardless.
code
java · 8 linesclass Pos {
long plain; // 64-bit HotSpot: single aligned 8-byte access; spec still allows tearing
volatile long shared; // atomic by specification everywhere, plus visibility/ordering
}
// Off-heap: alignment is yours to guarantee
ByteBuffer buf = ByteBuffer.allocateDirect(1024);
buf.putLong(offset, v); // atomic only if `offset` is 8-byte aligned; otherwise it can straddlego deeper
Know the rule from the specification — plain long/double may tear, volatile makes them atomic — without needing the hardware detail.
Add that 64-bit HotSpot on x86-64 uses a single aligned 8-byte access so tearing is not observable there, while the specification still permits it.
Explain the historical 32-bit implementation, the alignment caveat for off-heap access, and why visibility usually mandates the same annotation anyway.
Take the position explicitly: code targets the specification, not the current port set; encode the rule in review guidance and in the choice of access APIs rather than relying on deployment homogeneity.
## What the hardware actually does On x86-64, an aligned load or store of up to 8 bytes is performed atomically by the processor. HotSpot on a 64-bit build lays out object fields so that a `long` or `double` field is naturally aligned, and the JIT emits a single 8-byte `mov` for a plain access. Consequently a plain `long` field is, on that platform, never observed torn — no matter how hard you try to reproduce it with stress tests. The same practical outcome holds on AArch64 for aligned 64-bit accesses. This is why the tearing rule feels theoretical to most working engineers today. It is important to be able to say *why* it feels that way without concluding that the rule is obsolete. ## What the specification says, and why it still says it JLS §17.7 permits an implementation to treat a single non-volatile `long` or `double` read or write as two separate 32-bit actions. The reason is implementability on 32-bit platforms, where a 64-bit access has no free atomic instruction. Requiring atomicity unconditionally would have taxed every `long` field access on those machines to serve the rare shared case; instead the language made plain access cheap and required atomicity only for `volatile`, where the programmer has signalled that concurrency matters. Historically this was not theoretical. On the 32-bit HotSpot x86 port, plain `long` stores could be split into two 32-bit stores, and tearing was reproducible with a simple two-thread test alternating between `0L` and `-1L`: the reader would eventually print values that were neither. To honour the `volatile` guarantee on the same hardware, HotSpot used wider or locked instructions — 8-byte SSE2 moves, or `cmpxchg8b` — specifically so a `volatile long` could not tear. 32-bit ARM similarly needed paired exclusive instructions. Those ports have been retired (the 32-bit x86 port was deprecated and then removed in recent JDK releases), which is why the practical exposure on mainstream server deployments is now essentially nil. The rule remains in the specification because the specification defines the language, not one vendor's current port list, and other implementations and embedded targets exist. ## Where the gap still bites **Alignment is not automatic outside object fields.** The hardware atomicity guarantee is for *naturally aligned* accesses. Object fields are laid out aligned, but off-heap and buffer access is not necessarily so. A 64-bit read or write at an unaligned address — via `ByteBuffer.getLong(index)` at an index that is not 8-byte aligned, or via foreign-memory access at an arbitrary offset — can straddle a boundary and is not atomic even on x86-64. If you are building an off-heap data structure that multiple threads touch, alignment is your responsibility, and the memory-access APIs make you state the access mode you need. **Portability of the artifact, not the machine.** A JAR is not compiled for a CPU. Writing code whose correctness depends on "the JVM I happen to run on emits one instruction" is a latent bug, and it is invisible in review because the code looks fine. Marking a shared 64-bit field `volatile` costs essentially nothing on a field that is genuinely shared. **Atomicity was rarely the whole reason anyway.** The interesting realisation is that if a `long` field is shared across threads, tearing is usually the *least* of your problems: without `volatile` or a lock the reader may never see the update at all, and surrounding operations may be observed out of order. So the same annotation you would add for tearing is the one visibility already demanded. This is why the practical advice — "shared mutable 64-bit field: make it volatile or use `AtomicLong`" — survives even though tearing itself is unobservable on your laptop. **Compound updates are still unaffected.** Making a `long` `volatile` gives you atomic, ordered individual accesses. It does not make `position += n` safe. If the field is accumulated rather than assigned, you want `AtomicLong`, `LongAdder` or a `VarHandle` operation, regardless of what the hardware does with plain stores. ## Answering the question well The strong answer has three beats. First, be concrete about the hardware: aligned 8-byte access on x86-64 is atomic, so tearing is not observable there, and say so plainly rather than hedging. Second, be precise about the specification: it still permits tearing because it targets 32-bit-word implementations too, and cite the historical 32-bit HotSpot behaviour including what `volatile long` compiled to. Third, land the practical rule: write to the specification, watch alignment off-heap, and note that visibility usually mandates the same synchronisation anyway. Candidates who only say "long isn't atomic, always use volatile" sound like they are reciting; candidates who only say "it can't tear on 64-bit, don't worry" are reasoning from one machine. The gap between the two is the answer.
- How was a `volatile long` kept atomic on the old 32-bit HotSpot x86 port?By using an access wide enough to be atomic on that hardware rather than two 32-bit moves — an 8-byte SSE2 move, or a locked `cmpxchg8b` where a read-modify-write form was needed. That is a concrete example of `volatile` costing more than a plain access: the annotation forced the compiler to pick a different instruction sequence purely to honour the specification's atomicity guarantee.
- Can a 64-bit value tear on x86-64 despite the hardware guarantee?Yes, if it is not naturally aligned. The processor's atomicity guarantee applies to aligned accesses; an 8-byte access that straddles a boundary can be split. Object fields are laid out aligned by the JVM, but off-heap access through `ByteBuffer` or foreign-memory APIs at arbitrary offsets is not, so concurrent readers of an unaligned 64-bit slot can observe a mixture.
- If tearing cannot be observed on your production hardware, is `volatile` on a shared `long` pointless?No — atomicity is usually the smallest of the guarantees you need. Without `volatile` or a lock, a reader may never observe the update at all, and surrounding accesses can be reordered relative to it. The annotation you would add for tearing is the same one visibility and ordering already require for any genuinely shared mutable field.
saying these in an interview costs you the question
- Concluding that because tearing cannot be reproduced on x86-64, the specification rule is obsolete
- Claiming the JVM guarantees 64-bit atomicity on 64-bit platforms (it does not; the platform happens to provide it)
- Assuming any `ByteBuffer.getLong`/`putLong` is atomic regardless of alignment
- Believing `volatile long` only affects tearing and has no visibility or ordering effect
- Saying a plain `long` write and a `volatile long` write compile to identical code on every platform