What does a JVM actually have to emit at the end of a constructor to deliver the frozen-field guarantee, and why do reading threads need nothing on their side?
answer
- store-store barrier at constructor end
- x86: free (no store-store reordering)
- ARM/PPC: real but light fence
- reader safe via address dependency
- escape analysis can remove the barrier
basics
~20 sThe JVM must keep the final-field stores from moving after the object reference becomes visible. HotSpot emits a store-store barrier at constructor end — free on x86's strongly ordered stores, a real instruction on ARM. Readers need nothing because address dependencies preserve the load order.
solid answer
~60 sThe whole cost of the guarantee sits on the writing side. The JIT must not sink the stores of final fields past the store that publishes the reference, and the hardware must not reorder them either. HotSpot implements this by placing a **store-store barrier** at the end of a constructor that assigned final fields. What that compiles to depends on the architecture. On x86 the store buffer is FIFO and stores are not reordered with other stores, so the barrier degenerates to a compiler-only restriction — no machine instruction, effectively zero cost. On weakly ordered machines such as AArch64 or PowerPC it becomes a real fence instruction, though still a cheap one relative to a full barrier. The reading thread emits nothing because the read of the final field is *address-dependent* on the read of the object reference: the CPU cannot load the field before it knows where the object is. Every mainstream architecture respects that dependency, so no acquire fence is needed. This asymmetry — tiny writer cost, zero reader cost — is why the guarantee could be made unconditional in the specification.
code
text · 10 lines# after inlining `new Holder(42)` and publishing the reference
allocate object
store [obj + hdr] = class
store [obj + off_id] = 42 ; final field store
---- storestore barrier ---- ; freeze; x86: no instruction emitted
store [shared_ref] = obj ; publication
# reader
load r1 = [shared_ref]
load r2 = [r1 + off_id] ; address depends on r1 -> ordered by hardwarego deeper
Know that the cost is paid by the constructing thread, and that reading a final field is as cheap as reading any other field.
Name the store-store barrier at the end of the constructor and note that it is essentially free on x86 and a light fence on weakly ordered CPUs.
Explain both the compiler and hardware reorderings being prevented, why address dependency frees the reader, and contrast the cost with a volatile write.
Use the asymmetry to argue design: immutable objects with final fields are the cheapest safe-sharing mechanism available, since publication cost is one-time and read cost is zero, which is why they scale better than lock- or volatile-based sharing on read-heavy paths.
## The two reorderings the guarantee must forbid For a thread to see a fully built object through a plain reference, two independent reorderings have to be prevented: 1. **Compiler-side.** The JIT sees, within the constructing thread, that the field stores and the later store of the reference are independent. Nothing in single-thread semantics stops it from moving field stores after the publishing store, or from folding initialization into the caller after inlining the constructor. It must be told not to. 2. **Hardware-side.** Even with the stores in program order in the emitted code, a CPU with a non-FIFO store buffer may make the reference store globally visible before the field stores. ## What HotSpot emits HotSpot's answer is a **store-store barrier placed at the end of the constructor** — the point that corresponds to the freeze in the specification — whenever the constructor assigned final fields. Semantically it says: every store before this point becomes visible before any store after it. That is exactly enough, and deliberately no more: - It is not a full fence. Loads are unconstrained and store-load ordering is untouched, so the expensive part of a `volatile` write — draining the store buffer to order a later load — is not paid here. - It applies per constructor, not per field write. On x86 and x86-64 the memory model already guarantees that stores are not reordered with other stores, so the barrier lowers to nothing at the machine level; it survives only as a scheduling constraint inside the compiler. On AArch64 or PowerPC it lowers to a lightweight store-ordering fence. On any of them the amortized cost is negligible compared with allocating and initializing the object in the first place. There is a nuance about escape analysis: if the JIT proves an object never escapes the allocating thread, it may scalarize it and the barrier disappears entirely, because there is no other thread to order against. ## Why the reader is free The reader executes two loads: load the object reference from some field, then load the final field at an offset from that reference. The second load's *address* is computed from the first load's result. This is an address dependency, and real hardware cannot speculate past it in a way that produces an out-of-order result — the CPU literally does not know which address to fetch until the first load returns. Every architecture in current mainstream use preserves dependent load ordering. (The historical exception discussed in memory-model literature is the DEC Alpha, whose split caches could return a stale dependent load; the specification's phrasing about chains of dereferences exists in part because of that class of machine, and a conforming JVM there would have to emit a reader-side barrier.) The consequence is that a final-field read compiles to exactly the same instruction as a plain field read. There is no runtime tax on reading immutable objects, which is the whole reason the guarantee is worth having: `String`, boxed types, and every immutable value class in an application are read constantly and published rarely. ## What the JIT additionally gets from `final` Because a final field is not expected to change after the freeze, the JIT may treat repeated reads as returning the same value and constant-fold reads of `static final` fields with known values. That optimization is precisely what makes an escaped-`this` bug behave inconsistently: the interpreter may re-read the field each time and observe it change, while compiled code caches the first read. The optimization is legal only because the program is required not to publish the object early. ## Reflective writes Writing a final field after construction — through reflection or `Unsafe`-style access — breaks the assumption the barrier and the JIT's caching rely on. The specification is explicit that a thread which read the field before the modification need never observe the new value. Modern JDKs restrict such writes for exactly this reason: the freeze is a promise the runtime has already spent optimizations against. ## Summary of the cost model One cheap store-store barrier per constructor that assigns final fields, sometimes eliminated entirely, sometimes zero instructions on the target architecture; nothing at all on the reader. That asymmetry is what let the specification hand out the guarantee unconditionally rather than making it opt-in.
- How does the barrier at the end of a constructor differ in cost from a volatile write?The constructor barrier orders stores against stores only. A volatile write additionally has to order the write against subsequent loads, which on x86 requires draining the store buffer with a locked instruction or an explicit fence — far more expensive. That is why publishing immutable objects through final fields is cheaper than making every field volatile.
- Can the JIT ever omit the barrier entirely?Yes. If escape analysis proves the constructed object cannot be reached by any other thread, there is nothing to order against and the barrier is dropped, often along with the allocation itself if the object is scalarized. The barrier is also unnecessary if the constructor assigned no final fields.
saying these in an interview costs you the question
- Claiming reading a final field is slower or requires a fence.
- Describing the guarantee as implemented by a full memory fence or a lock.
- Assuming the barrier is emitted per final-field assignment rather than once at the end of the constructor.
- Saying x86 needs an explicit instruction for store-store ordering.
- Thinking reflection can rewrite a final field with the same guarantees intact.