In three-colour marking (white unreached, grey pending, black scanned), when may a black object safely point at a white one?
answer
- the rule comes in two strengths
- the edge itself is not the problem
- the frontier must still be able to arrive
- white chain hanging off a grey object
- strong forbids, weak protects
basics
~20 sA black-to-white edge is safe when the white target is still reachable from some grey object through a chain of white objects — the weak tri-colour invariant, which requires only that the pending frontier can still walk to the target.
solid answer
~50 sThe rule has two strengths. The **strong** form forbids the edge outright: no black object may reference a white one. The **weak** form permits it, provided every white object that a black object references is still reachable from some grey object along a path that runs only through white objects. The insight is that the real requirement was never "no such edge" but "the frontier can still get there": while a grey-rooted white path to the target survives, the marker will reach and shade it as it drains its work list. The strong form is simply the cheapest way to guarantee that, because it can be checked at a single store with no reasoning about paths at all. A design that leans on the weak form must guarantee the protecting path some other way, so relaxing the invariant moves the work rather than removing it.
go deeper
Know that marking has a safety rule about references from scanned objects to unreached ones, and that it is what makes a concurrent trace trustworthy.
Be able to state the strong form and say why a barrier can enforce it from a single store without looking at the rest of the graph.
Recognise the weak form when you see it, and use it to tell a legal black-to-white edge from an actual lost-object hazard rather than treating every such edge as a bug.
Judge where enforcement should live. Relaxing to the weak form does not remove work; it relocates the guarantee, and the decision is about which hot path absorbs the cost.
## One requirement, two statements Marking is safe when no reachable object is left white at the end of a cycle. Two invariants are used to guarantee that, and they differ in how much they forbid. - **Strong tri-colour invariant:** there is no reference from a black object to a white object, at any instant. - **Weak tri-colour invariant:** a black object may reference a white object, provided that white object is also reachable from some grey object by a path made up entirely of white objects. The weak form is strictly more permissive; anything satisfying the strong form satisfies the weak one as well. What both guarantee is the same end state: when the grey set drains, every object still white is genuinely unreachable. ## What "protected by grey" means The marker only ever moves forward from grey objects. Scanning one means walking its reference fields, shading each white target grey, and then blackening it. So a white object is certain to be visited as long as some chain of *not-yet-scanned* objects leads to it from the pending set. That is why the protecting path must consist of white objects. Suppose the only route from a grey object `G` to the white object `W` passes through a black object `B`. `B` is off the work list; nothing will follow its fields again. The chain is a dead end, `W` is not protected, and the weak invariant is violated even though a grey object nominally "reaches" it. Worked the other way: `B` is black and holds a reference to white `W`, and grey `G` still holds `G.next = X` where `X` is white and `X.ref = W`. This is legal under the weak form. The marker will eventually scan `G`, shade `X` grey, later scan `X`, and shade `W` grey. The black-to-white edge from `B` never mattered, because the frontier arrived anyway. ## Comparing the two forms | | Strong form | Weak form | |---|---|---| | Forbids | every black-to-white reference | only unprotected ones | | What must be established | a local fact at one store | a property of a path through the heap | | Typical enforcement | shade on the store that creates the edge | guarantee the protecting chain survives, by some other means | | Ease of reasoning | high — no graph argument needed | lower — requires an argument about reachability | ## Why the strong form is the usual choice A barrier sees one store. It knows the object being written, the reference going in, and the reference coming out; it does not know whether some other chain of white objects still leads to the target, and working that out would mean a traversal at store time. The strong form is therefore the one a barrier can actually maintain: shade something grey and the edge becomes harmless immediately, with no appeal to the rest of the graph. The weak form earns its place as a *reasoning tool* and as the basis for designs that enforce the requirement somewhere other than the store. Some collectors intervene when a reference is loaded rather than when it is written, and the property they preserve is the weak statement — the target is kept protected, not the edge prevented. In either case the guarantee delivered to the trace is identical. ## How the weak form gets broken Under the weak invariant a single bad action is rarely enough. The loss needs a combination: 1. a reference to an unreached object ends up in an already-scanned object, and 2. the last white chain from the pending set to that object is destroyed. Until step 2 happens, the situation in step 1 is legal and self-correcting. Once both hold, the object has no route the marker will travel, and it ends the cycle white — reclaimed while the program can still reach it through the scanned object. The two barrier families correspond to preventing one or the other of these, which is also why exactly two families exist rather than a dozen. ## Why any of this is worth knowing Three practical payoffs: - **It stops a false alarm.** Reasoning only with the strong form leads people to treat every black-to-white reference in a heap dump or a trace log as proof of a collector bug. It is not; under the weak form it can be perfectly legal. - **It explains where enforcement is allowed to live.** If the protecting path can be guaranteed another way, the store-side hook is not the only possible place to pay. - **It sharpens what a barrier actually preserves.** The barrier is not preventing an edge for its own sake; it is keeping a target reachable by the frontier. Every argument about which stores need a hook is an argument about that, whether or not it is phrased that way.
- Why must the protecting path run through white objects only?Because the marker advances out of grey objects and then blackens them. A path that leaves a grey object and passes through a black one is a dead end: nothing will follow that black object's fields again, so anything beyond it is not guaranteed a visit. Only an unscanned, white chain hanging off the pending set is certain to be walked.
- If the strong form is easier to enforce, why is the weak form worth knowing?It says what the barrier is actually preserving. Designs that enforce the rule where a reference is read, rather than where it is written, rely on the weaker statement, and so does any argument about which stores genuinely need a hook. Reasoning only with the strong form leads people to call every black-to-white edge a defect, which it is not.
saying these in an interview costs you the question
- Says any black-to-white reference is automatically a lost object.
- Thinks the weak form removes the need for a barrier entirely.
- Believes the protecting path may run through a black object.
- Treats the two forms as different algorithms rather than two strengths of one rule.
- Assumes the weak form exists to retain less garbage.