The Shenandoah collector relocates objects while application threads are running. Explain the load reference barrier it uses, the Brooks forwarding pointer it replaced, and what invariant they enforce.
answer
- Brooks pointer = permanent extra word, indirect on every access
- LRB (JDK 13+) = barrier on reference LOAD from heap
- Self-healing: barrier writes the corrected ref back
- To-space invariant: mutators never touch from-space
- Racing copies resolved by CAS on the forwarding word
basics
~30 sA load reference barrier is JIT-inserted code that runs when a reference is loaded from the heap: if the referent lives in a region being evacuated, the barrier resolves or copies it to its new location and writes the corrected reference back. That enforces the to-space invariant — application threads only ever work with up-to-date copies. It replaced the Brooks pointer, an extra header word every access had to indirect through.
solid answer
~1 min**Brooks pointers (the original design).** Every object carried an extra word pointing to its current copy — to itself normally, or to the new copy once evacuated. Every access indirected through that word, so a thread holding a stale reference still reached the live copy. Costs: one extra word per object, and an indirection on every read *and* write. **Load reference barriers (JDK 13 onward).** The barrier moved to the point where a **reference is loaded out of the heap**. The barrier checks the global GC state; in the common case, no cycle is evacuating and it is a cheap check. During evacuation, if the referent sits in a collection-set region, the barrier resolves it — copying the object to a new region if nobody has yet — and then **self-heals**: it writes the updated reference back into the field it was loaded from, so subsequent loads hit the fast path. Forwarding information lives in the object header during the cycle rather than in a permanent extra word. **The invariant** is the *to-space invariant*: after any load, a mutator holds only references to to-space (current) copies, so ordinary field reads and writes need no further checks and can never land on a stale copy. When two threads race to copy the same object, a compare-and-swap on the forwarding word picks the winner and the loser discards its copy.
code
java · 12 lines// Application code:
Node next = node.next; // <-- reference LOAD from the heap: barrier here
int v = next.value; // plain field read, no barrier needed
next.value = v + 1; // plain field write, no barrier needed for movement
// Conceptual barrier on the load (emitted by the JIT, not written by you):
// ref = raw_load(node.next)
// if (gc_state_requires_action && in_collection_set(ref)) {
// ref = resolve_or_evacuate(ref); // CAS installs forwarding on first copy
// cas_store(&node.next, raw, ref); // self-healing write-back
// }
// use refgo deeper
Recall that the collector inserts a check when a reference is loaded so a moved object is still found correctly.
Explain the load barrier's steps and that it replaced a per-object forwarding word, and name the memory saving.
Name the to-space invariant, the CAS race resolution, and self-healing, and quantify where the throughput cost lands (reference-load-heavy code).
Reason about the trade explicitly: barrier design determines the throughput tax and mutator-assist behaviour, which is what you weigh against the latency benefit when adopting the collector.
## The core difficulty Moving an object while the application runs creates a window in which two copies exist and threads may hold references to either. If a thread writes a field on the old copy after the collector has finished copying, that write is lost. If a thread reads the old copy, it may see stale data. Any concurrent compactor must therefore intercept application memory access with **barriers**: short code sequences the JIT emits inline around heap accesses so the collector and mutators agree on which copy is authoritative. Shenandoah has had two answers to this, and interviewers ask about both because the change is instructive. ## Brooks forwarding pointers — the original design Every object was allocated with one extra word immediately before its normal header. In the steady state that word pointed at the object itself. When the collector evacuated the object, it allocated a copy in a new region and then set the *old* copy's forwarding word to point at the new one. The rule was: never touch an object directly; always dereference the forwarding word first. A thread holding a stale reference to the old copy would follow the forwarding word and end up at the new copy, so reads and writes always reached the authoritative object. The scheme is elegant but expensive: - **Memory**: one extra word on *every* object in the heap, permanently — a significant footprint tax on small-object-heavy workloads. - **Time**: an indirection on effectively every access, read and write alike. - **Complexity**: keeping the invariant intact around the JIT's optimisations, and around writes in particular, was intricate. ## Load reference barriers — the current design JDK 13 replaced Brooks pointers with **load reference barriers**. The barrier moved from "every access to an object" to "every load of a *reference* out of the heap." The idea: if you guarantee that a reference is corrected at the moment it is loaded, then everything downstream — field reads, field writes, comparisons, calls — operates on an up-to-date reference and needs no barrier at all. What the barrier does, in order: 1. **Check global GC state.** A thread-local word records whether the collector is in a phase requiring barrier action. Outside evacuation the barrier is a cheap test-and-branch, and the JIT keeps the fast path inline. 2. **Check the region.** If the loaded reference points into a region that is not in the collection set, return it unchanged. 3. **Resolve or evacuate.** If it points into a collection-set region, read the forwarding information in the object's header. If the object has already been copied, use the new address. If not, the *mutator thread itself* copies the object into a new region and installs the forwarding pointer with a **compare-and-swap**. If the CAS fails, another thread won the race; this thread abandons its copy (the region's allocation is simply wasted) and uses the winner's address. 4. **Self-heal.** Write the corrected reference back into the memory location it was loaded from, so future loads from that field take the fast path. Over the course of a cycle, mutator traffic itself repairs much of the heap, and the concurrent update-references phase mops up the rest. Because forwarding data lives in the object header only while the object is being relocated, the permanent per-object word disappeared. That was the practical headline of the change: same latency properties, lower memory footprint, simpler and faster barriers. ## The to-space invariant Both designs exist to uphold an invariant, and naming it is what separates a memorised answer from an understood one. Shenandoah maintains the **to-space invariant**: once a reference has been loaded through the barrier, it points at the object's current (to-space) copy. Mutators never operate on from-space copies. Therefore: - Field writes go to the authoritative copy and cannot be lost when evacuation completes. - Field reads see current data. - No write barrier is needed for correctness of *movement* (a separate snapshot-at-the-beginning write barrier exists to support concurrent marking, which is a different problem). Contrast this with the Brooks design, which allowed mutators to hold from-space references and fixed things up on every dereference. Moving to a to-space invariant is what let the fast path shrink to a state check. ## Practical consequences worth stating - **Throughput cost is on loads.** Reference-load-heavy code — pointer chasing through linked structures, large object graphs — pays more than arithmetic-heavy code. Typical measured overheads land in the single-digit-to-low-double-digit percent range and are workload dependent. - **Mutators may do collector work.** A thread that loads a reference into the collection set may itself copy the object. This is a deliberate design choice: it spreads evacuation cost across threads and keeps the collector from being the sole bottleneck, but it means an application thread can occasionally take a small, unpredictable detour. - **Reference comparison is safe.** Because loads are corrected, comparing two references for identity compares to-space addresses on both sides. - **Wasted copies exist by design.** Racing evacuations discard the losing copy; that wasted space is reclaimed when the region is recycled, and it is a reason the collector needs free headroom to work with. ## How to answer under pressure Say what the barrier intercepts (reference loads from the heap), what it does (resolve/evacuate, then self-heal the source field), what invariant it maintains (to-space), what it replaced (a permanent per-object forwarding word requiring indirection on every access), and what it costs (throughput on load-heavy paths). That is the complete answer.
- Why is a barrier on reference loads sufficient, when a naive design also barriers writes?Because correcting the reference at load time establishes the to-space invariant: every reference a mutator holds already points at the current copy, so any subsequent read or write on it necessarily targets the authoritative object. A write barrier would be redundant for relocation correctness. Shenandoah does keep a separate write barrier, but that one exists to support concurrent marking (snapshot-at-the-beginning), not object movement.
- What happens when two application threads load a reference to the same not-yet-evacuated object at the same time?Both may allocate a copy in a new region and attempt to install a forwarding pointer with a compare-and-swap on the object's header. Exactly one CAS succeeds; that copy becomes authoritative. The losing thread abandons its copy — the space is simply wasted until the region is recycled — and proceeds with the winner's address, so both threads end up on the same to-space object.
- Why did removing the Brooks pointer matter in practice?It removed one machine word from every object in the heap, which on workloads dominated by small objects is a real footprint reduction, and it removed a mandatory indirection from ordinary accesses. The fast path became a state check on reference loads instead of a dereference on every access, improving both memory usage and throughput while keeping the same pause characteristics.
A mail-forwarding rule that not only redirects the letter but also corrects the sender's address book, so the next letter goes straight to the new address.
saying these in an interview costs you the question
- Saying the barrier runs on every field access, including primitives — it runs on reference loads
- Claiming Shenandoah still adds a permanent forwarding word to every object after JDK 13
- Describing the barrier as pure bookkeeping — it can copy the object on the mutator's thread
- Confusing the load reference barrier (relocation) with the snapshot-at-the-beginning write barrier (marking)
- Assuming stale references are left in fields forever; the barrier self-heals the source location