HotSpot's ZGC compiles a load barrier into application code. When does that barrier run, what does its fast path test, what can happen on the slow path, and what does it mean that the barrier is 'self-healing'?
answer
- Fires on reference loads from the heap, not primitives
- Fast path: test color vs good color, branch
- Slow path: mark, or forwarding-table lookup / mutator-copies-with-CAS
- Self-heal = write the fixed pointer back into the slot
- Load (not write) barrier because objects MOVE
basics
~20 sEvery read of a reference field from the heap runs a load barrier. The fast path tests the pointer's color against the current good color; a bad color takes a slow path that marks the object or looks up its new location, then writes the corrected pointer back into the field it was read from — self-healing.
solid answer
~60 sThe barrier fires on **reference-typed loads from the heap** — reading an object field or array element whose type is a reference. Loads of primitives, and re-uses of a reference already sitting in a local/register, are not barriered. **Fast path:** a couple of instructions — test the loaded pointer's color bits against the current good color and branch. In the common case the color is good and execution continues. **Slow path:** the color is stale, so the barrier does whatever the current phase requires. During marking, it marks the object and pushes it on the mark stack. During relocation, it consults the forwarding table for the object's new address; if the object is in the relocation set but not yet copied, the *application thread itself* copies it and installs the forwarding entry with a compare-and-set. **Self-healing:** after fixing the pointer, the barrier stores the corrected, good-colored pointer back into the memory location it was loaded from, so that location never takes the slow path again. Slow-path work is therefore paid at most once per reference field per cycle.
go deeper
Know that ZGC inserts a check on every reference read so the program never uses a pointer to a moved or unmarked object.
Describe the fast path (color test plus branch) versus the slow path (mark or remap), and state that the barrier applies only to reference loads from the heap.
Explain forwarding tables, mutator-assisted relocation with CAS, and why self-healing bounds slow-path cost and removes the need for a remapping pause.
Reason about the throughput cost profile — pointer-chasing workloads pay most — and about how mutator-assisted relocation distributes GC work into request latency rather than into pause histograms.
## Where the barrier sits ZGC keeps GC state inside pointers, and it relocates objects while the application runs. Both facts mean the application can be holding a reference whose color is stale — the object may have been moved, or may not yet be marked in the current cycle. The load barrier is the piece of machinery that guarantees the application never *acts* on such a reference. The interpreter, C1 and C2 all emit it. It applies to **loads of reference-typed values out of the heap**: `obj.field` where the field is a reference, `array[i]` for an object array, and equivalent internal accesses. It does *not* apply to: - loads of primitives (`int`, `long`, `double` fields) — those carry no GC state; - uses of a reference already loaded into a local variable or register, because it was already healed when it was loaded; - stores, in the non-generational design. (Generational ZGC adds a separate *store* barrier for remembered-set maintenance — a different barrier with a different job.) ## The fast path The fast path is deliberately tiny: mask/compare the loaded value's color bits against the currently good color, and branch if they differ. On x86-64 this is on the order of a test-and-branch plus the address materialization, and the branch is overwhelmingly not taken, so it predicts well. This is the throughput tax of ZGC. It is real but modest for most workloads; it hurts most in pointer-chasing inner loops — traversing a linked structure billions of times — and least in code dominated by primitive arithmetic or I/O waits. ## The slow path When the color is bad, the barrier calls out to the runtime and does phase-dependent work: - **During concurrent marking.** The referenced object has not yet been marked with this cycle's color. The barrier marks it and pushes it onto a mark stack so the collector will traverse its fields. This is what keeps concurrent marking correct in the presence of a mutating object graph: any reference the application actually loads becomes known to the collector. - **During concurrent relocation.** The reference points into a page in the relocation set. The barrier looks the object up in the **forwarding table** (an off-heap side structure per relocated page mapping old address → new address). Three cases: the object was already copied, so take the new address; the object has not been copied, so the *mutator itself* copies it and installs the forwarding entry with a compare-and-set — if the CAS loses to a GC worker or another mutator, the winner's address is used and the loser's copy discarded; or the object is not in the relocation set at all and only needs its color refreshed. - **After relocation, before the next cycle finishes remapping.** The reference is simply stale in color and possibly in address; the forwarding tables from the last cycle answer the question until the next marking pass has remapped everything. ## Self-healing The crucial optimization: having computed the correct, good-colored pointer, the barrier **writes it back into the heap slot it was loaded from**. Now every subsequent load of that field takes the fast path. Two consequences follow. First, slow-path cost is bounded: each reference-holding memory location can cost slow-path work at most once per phase transition, not once per read. A hot field read a billion times is repaired once. Second, the heap converges on being fully remapped *as a by-product of the application running*, which is why ZGC does not need a dedicated stop-the-world remapping pass — anything the application never touches is fixed by the next concurrent marking traversal instead. The write-back must tolerate races: several threads can heal the same slot concurrently, and a mutator may have overwritten the slot with a different reference entirely. The healing store is therefore conditional (a compare-and-set of the expected stale value), so it never resurrects an overwritten reference. ## Why a load barrier rather than a write barrier Collectors that mark concurrently but do not *move* objects can get away with a write barrier that records mutations. ZGC moves objects, so the dangerous operation is not writing a reference — it is *following* one. Only a barrier on loads can intercept a mutator before it dereferences a pointer to memory that has been evacuated. That is also what lets ZGC free an evacuated page immediately after copying its live objects, without waiting for every stale reference in the heap to be found and fixed. ## What this looks like in practice You do not write or configure the barrier; it is emitted by the compilers. But it explains observable behaviour: a throughput dip relative to a non-concurrent collector on reference-heavy code, application threads occasionally doing GC work (copying) on their own dime, and — combined with immediate page reclamation — ZGC's ability to keep free memory available without a compaction pause.
- Does the barrier run when reading an int field or when re-using a reference already in a local variable?No to both. Primitive loads carry no GC state, so there is nothing to check. A reference already held in a local variable or register was healed when it was originally loaded from the heap, so subsequent uses need no barrier. The JIT also coalesces redundant barriers when it can prove a reference was already checked.
- If an application thread can end up copying an object itself, what does that mean for latency measurement?It means some relocation cost is charged to application threads rather than to GC threads, so it appears as scattered microsecond-scale work inside request handling rather than as a pause in the GC log. It is bounded — one copy per object, and only for objects in the relocation set that the thread actually touches — but it is a reason GC-pause histograms alone do not fully describe ZGC's latency impact.
A forwarding-address check at the front desk: the first visitor asking for a resident who has moved is told the new room and the desk updates the directory entry, so nobody after them has to ask.
saying these in an interview costs you the question
- Calling it a write barrier or describing it as intercepting stores — ZGC's core barrier is on loads because objects move.
- Thinking the slow path runs on every read of a moved object rather than once, thanks to self-healing.
- Claiming application threads never do GC work; under ZGC a mutator can relocate an object itself.
- Saying the barrier applies to all field loads including primitives.
- Describing self-healing as an unconditional store, ignoring that it must CAS to avoid clobbering a concurrent write.