Compare snapshot-at-the-beginning and incremental-update as correctness strategies for concurrent garbage-collection marking: what does each one's write barrier record, and how does the choice affect floating garbage and end-of-cycle work?
answer
- SATB = pre-barrier, saves the OLD value
- Incremental update = post-barrier, records the NEW value
- SATB: floating garbage, predictable termination
- Incremental update: less garbage retained, mutation-sensitive remark
- G1/Shenandoah = SATB; CMS = incremental update
basics
~20 sA snapshot-at-the-beginning barrier saves the overwritten (old) reference before a store, so anything live when marking began stays marked - producing floating garbage but predictable termination. An incremental-update barrier records the newly stored reference or re-greys the target, marking a more current picture at the cost of rescanning.
solid answer
~50 s**Snapshot-at-the-beginning (SATB, Yuasa)** puts a **pre-write** barrier on reference stores: before `obj.f = newVal`, it enqueues the *old* value of `obj.f`. Those saved references are drained as extra roots. The collector therefore marks over a logical snapshot of the heap as it was when marking started: **anything reachable at cycle start is marked live**, even if it dies during the cycle. That is **floating garbage**, reclaimed in the next cycle. The benefit is monotone progress and predictable termination. G1 and Shenandoah use SATB. **Incremental update (Dijkstra/Steele)** puts a **post-write** barrier that records the *new* reference or re-greys the object that received it. The collector chases the current graph, so it retains less floating garbage - but re-greying can create new work late in the cycle, so the final-mark pause has more to rescan and termination is less predictable. Classic CMS worked this way, rescanning dirtied cards in a final pause. Tradeoff in one line: **SATB trades heap headroom for pause predictability; incremental update trades pause predictability for precision.**
code
text · 11 lines// SATB pre-write barrier (Yuasa): preserve what is being destroyed
if (marking_active) {
oop old = *field;
if (old != NULL) satb_buffer.enqueue(old); // drained as an extra root
}
*field = new_value;
// Incremental-update post-write barrier (Dijkstra/Steele): notice what was created
*field = new_value;
if (marking_active && is_white(new_value))
shade_gray(new_value); // or dirty the holder's card for rescango deeper
Know that concurrent collectors need a write barrier on reference stores, and that one style saves the old value while the other notices the new one.
State which value each barrier records and name the direct consequence: floating garbage for the snapshot approach, rescanning for the update approach.
Discuss the operational profile - heap headroom requirements versus mutation-sensitive remark pauses - and place the JVM collectors on each side.
Frame it as choosing where variance lives: bounded marking work with extra retained memory, or tighter reclamation with a pause that scales with mutation rate, and connect that to heap sizing and latency objectives.
## The problem both solve Concurrent marking is unsafe if, at termination, a scanned (black) object references an unmarked (white) object with no unscanned (gray) path leading to it. Both of those conditions are necessary, so a collector only needs to break one. SATB and incremental update are the two families, one per condition. ## Snapshot-at-the-beginning SATB reasons about a **logical snapshot** of the object graph taken when marking started. Its guarantee: *everything reachable at the start of the cycle will be marked live*. To maintain that, the barrier runs **before** every reference store and saves what is about to be destroyed: ``` if (marking_in_progress) { oop old = *field; if (old != NULL) satb_queue.enqueue(old); } *field = new_value; ``` Each thread has a thread-local satb buffer; full buffers are handed to marking threads, which treat the saved references as additional roots and trace from them. Deleting the last live path to an object therefore cannot hide it - the deletion itself recorded it. Newly allocated objects also need handling, since they did not exist in the snapshot. Collectors treat them as implicitly live for this cycle, for example by tracking a top-at-mark-start pointer per region and considering everything above it live. **Consequences.** - *Floating garbage*: an object that was live at cycle start but dies during the cycle is still marked. It survives to the next cycle. The heap therefore needs headroom; a heap sized too tightly against live data will run cycles back to back or fail over to a full collection. - *Predictable termination*: work only ever comes from the snapshot plus a bounded set of saved references, so marking converges. No re-scanning of already-black objects is needed. - *Cost profile*: a pre-barrier includes a load of the old value on every reference store, which is why it is conditioned on marking being active. G1 and Shenandoah both use SATB. ## Incremental update Incremental update reasons about the **current** graph. Its barrier fires on the store and records the new reference, or marks the target black object dirty so it will be scanned again: ``` *field = new_value; if (marking_in_progress && is_white(new_value)) { shade_gray(new_value); // Dijkstra style: shade the new referent // or: mark the holder's card dirty for rescanning (Steele style) } ``` Because an already-scanned object that acquires a new referent is effectively demoted back to the frontier, the traversal is not monotone in the same way. The collector re-scans dirty regions, and mutators can dirty them again while it does. **Consequences.** - *Less floating garbage*: objects that die during the cycle are more likely to be recognized as dead, since the collector tracks current reachability rather than a start-of-cycle snapshot. - *Harder termination*: a highly mutating workload keeps producing new work. In practice the collector gives up on doing it all concurrently and finishes in a **stop-the-world remark** that rescans dirtied cards and stacks. That final pause scales with mutation rate, which makes it the least predictable part of the cycle - historically the reason CMS remark pauses could be surprisingly long. - *Cost profile*: a post-barrier avoids loading the old value, but the rescanning it implies moves work into the pause. CMS is the classic JVM example and is why its remark phase was the pause to worry about. ## Choosing between them The question is where you would rather pay. - If the goal is **bounded, predictable pauses**, SATB is attractive: marking work is bounded by the snapshot, and the final pause mostly drains queues and re-scans stacks rather than chasing mutation. You pay in retained floating garbage, so you provision more heap. - If the goal is **maximum reclamation per cycle** and pauses may vary, incremental update collects more per cycle but risks a mutation-sensitive final pause. Modern low-pause JVM collectors have gone the SATB route for marking, and address the relocation half of the problem separately with load barriers and forwarding rather than by changing the marking strategy. ## Things that trip people up - Floating garbage is **not** a correctness bug. Nothing live is lost; some dead objects are kept one cycle longer. - SATB does not mean marking sees a physical copy of the heap. Nothing is copied; the "snapshot" is maintained incrementally by the barrier. - Neither strategy removes the final safepoint. Thread stacks are not covered by heap write barriers, so a short pause remains. - Both barriers are conditioned on marking being in progress, so their cost is not paid continuously - but the branch and the code size are always there.
- Is floating garbage a correctness problem?No. It is a space and timing cost, not a safety violation - the collector keeps objects that are already dead until the next cycle, but never frees anything live. It matters operationally because it raises the effective live-set size, so a heap sized with no headroom will trigger cycles more often and can degrade into a full collection under allocation pressure.
- Why did CMS remark pauses have a reputation for being unpredictable?Because incremental-update marking pushes work created by mutation into a final stop-the-world phase: dirtied cards and thread stacks must be rescanned to reach a consistent marking state. The amount of that work scales with how much the application mutated during the concurrent phase, so a write-heavy burst directly lengthens the pause.
- Does SATB slow down every reference store in the application?The barrier is conditioned on marking being in progress, so outside a marking cycle it costs a predictable branch on a flag rather than the full path. During marking it adds a load of the old value plus a possible enqueue. The aggregate throughput tax is real but modest for most workloads, and largest for code that rewrites references in tight loops.
SATB is taking attendance at the door when the meeting starts and counting everyone on that list as present for the whole meeting. Incremental update is walking the room repeatedly and re-checking wherever someone moved - more accurate, but you may never stop walking while people keep moving.
saying these in an interview costs you the question
- Saying SATB records the new value or that incremental update records the old value
- Treating floating garbage as a bug that loses data
- Believing SATB physically copies or snapshots the heap
- Claiming either barrier eliminates the stop-the-world final-mark pause
- Assuming barriers run at full cost even when no marking cycle is active