Why does Go's write barrier shade both the pointer being overwritten and the pointer being stored?
answer
- one write cannot lose an object; two can
- the marker never revisits what it finished
- one half saves what the write destroys
- goroutine stacks get no barriers at all
- so a scanned stack never needs a second look
basics
~20 sBecause either half alone leaves a hole. Shading the newly stored pointer stops an object being smuggled into memory the marker has already finished with; shading the overwritten pointer keeps anything that was reachable at the moment of the write, which is what lets Go scan each goroutine stack once and never re-scan it.
solid answer
~50 sDuring concurrent marking a goroutine can hide a live object with two writes: copy the only reference into an object the marker has already finished with, then delete the original reference from somewhere the marker has not reached yet. A Dijkstra-style barrier that shades the pointer being *stored* closes the first half, but Go deliberately does not put barriers on writes to a goroutine's own stack — they are far too hot — so with insertion barriers alone every stack had to be re-scanned in a stop-the-world pause at the end of marking, and that pause grew with the number and size of goroutine stacks. Go's hybrid barrier adds the Yuasa-style half: shade the pointer being *overwritten*, so anything reachable when the write happened survives the cycle. With both halves a stack can be scanned once, blackened, and never looked at again. The price is floating garbage — objects retained by the barrier that were already dead.
go deeper
You are not expected to derive this, but be able to say that Go inserts extra code around pointer writes during a collection so the collector cannot lose an object the program is still using.
Explain the two conditions that must both hold for an object to be hidden, and say which half of the barrier attacks each. Being able to state that one half preserves what a write destroys is the core of a good answer.
Tie the design to the observable outcome: because both halves are present, goroutine stacks are scanned once and never re-scanned, which is why pauses no longer scale with goroutine count. Name floating garbage as the price.
Frame it as a deliberate exchange of a slightly larger retained heap and per-write throughput for pause times that do not grow with the program's shape — and be ready to say when a workload would rather have the memory back.
## The hazard the barrier exists to prevent A concurrent marker walks the object graph while the application keeps rewriting it. It marks objects it has reached, and once it has finished scanning an object's fields it does not come back to it. That last part is what makes concurrent mutation dangerous. An object can be lost from the marker's view only if **two things both happen** during the cycle: 1. A reference to the object is installed into something the marker has **already finished with** (so it will never be re-examined), and 2. Every reference to that object from anywhere the marker **has not yet reached** is destroyed. After both, the object is reachable from the program but invisible to the marker, and it gets freed while still in use — the worst class of GC bug, showing up later as corrupted data or a mysterious crash. ## The two classical barriers A **write barrier** is a small piece of code the compiler injects around pointer stores so the runtime can observe the mutation. Two designs fix the hazard from opposite ends: - A **Dijkstra-style insertion barrier** shades the pointer being *stored*: whatever you install, the marker will look at. This attacks condition 1. - A **Yuasa-style deletion barrier** shades the pointer being *overwritten*: whatever you destroy is preserved as if a snapshot had been taken when the write happened. This attacks condition 2. (*Shade* here just means: hand the object to the marker as work to do, so it will definitely be scanned before the cycle ends.) ## Why Go uses both The reason is goroutine stacks. Stack slots are written constantly — every local variable assignment, every call — and putting a barrier on stack writes would be ruinously expensive, so Go does not. That leaves unbarriered writers loose in the system, and it has a consequence: with only an insertion barrier, a stack that had already been scanned could still have a fresh pointer written into it without the collector noticing. The only fix was to re-scan every goroutine stack at the end of marking, with the world stopped. A program with hundreds of thousands of goroutines paid a pause proportional to all of them. Go's **hybrid barrier** combines the two halves so that the deletion half covers what the missing stack barriers cannot. Conceptually, on a pointer write it shades the value being overwritten, and it shades the value being stored while the writing goroutine's own stack has not yet been scanned. The runtime's implementation records both pointers from the store into a small per-P write-barrier buffer, and flushes that buffer into marking work in batches rather than doing the shading inline at every store. The payoff is precise: **a goroutine's stack is scanned once during the cycle and is then done — permanently black, never re-scanned.** That is what removed a stop-the-world phase whose length scaled with goroutine count, and it is the single biggest reason Go's pauses are measured in microseconds rather than milliseconds. ## What it costs **Floating garbage.** The deletion half deliberately retains objects that were reachable at the moment a pointer was overwritten, even if they became garbage a nanosecond later. They are not freed this cycle; they are freed by the next one. Concurrent collectors all pay some version of this — you are collecting a graph that has moved on since you started — and Go accepts a slightly larger retained heap in exchange for not re-scanning stacks. **Throughput.** Every barriered pointer store does more work than a plain store, and that cost lands on the application's own goroutines, in proportion to how many pointers they rewrite while marking is in progress. ## Answering it well Say the hazard needs two conditions, not one. Say which half of the barrier attacks which condition. Then give the Go-specific reason the language needs both: unbarriered stack writes, and the stop-the-world stack re-scan that combining the halves eliminated. Finish with the cost — floating garbage — so it is clear you know the trade rather than just the mechanism. ## Frequently muddled points - *One shade is enough because the marker will come back.* It will not. Finished objects are not revisited; that is the whole premise. - *Shading the overwritten pointer frees it.* The opposite — shading keeps it alive for this cycle. - *The barrier runs on stack writes too.* It does not, and that omission is exactly why the deletion half is needed. - *The barrier stops the writing goroutine.* It does not block; it records work for the marker.
- What does Go get in exchange for the deletion half of the barrier?A goroutine stack can be scanned once and then treated as finished for the rest of the cycle. Without the deletion half, unbarriered stack writes meant every stack had to be re-scanned with the world stopped at the end of marking, a pause that grew with the number and size of goroutine stacks. Removing that re-scan is what made Go's pauses sub-millisecond.
- What does the hybrid barrier cost in retained memory?Floating garbage. An object that was reachable when a pointer to it was overwritten is kept alive for the rest of the cycle even if it died immediately after. The next cycle reclaims it. In exchange for a slightly larger retained heap you get correctness without re-scanning stacks.
- Do writes to local variables on a goroutine's own stack go through the barrier?No. Stack writes are unbarriered because they are extremely frequent and a barrier on each would wreck performance. That omission is the reason the deletion half exists: it covers what the missing stack barriers cannot, so a scanned stack stays valid without being looked at again.
- Does the barrier do the marking work inline at the store?Not usually. The compiled barrier records the old and new pointers into a small per-P write-barrier buffer and returns; the runtime drains that buffer into marking work in batches. That keeps the per-store cost closer to a couple of stores and a bounds check than to a full call into the marker.
It is like an auditor working through a warehouse: you must both stop staff hiding a crate in an aisle already counted, and keep a copy of the manifest line they tear up in an aisle not yet reached.
saying these in an interview costs you the question
- Says one shade suffices because the marker revisits objects
- Thinks shading the overwritten pointer frees it early
- Claims Go puts write barriers on goroutine stack writes
- Believes Go re-scans every goroutine stack at mark termination
- Says the barrier blocks the writing goroutine until marking catches up
- Cannot name any cost, only the benefit