After porting a Go service from amd64 to arm64, workers sometimes read a stale package-level config pointer. Why now?
answer
- the port revealed it, it did not cause it
- strong ordering was hiding a missing edge
- the compiler can break it on x86 too
- publish before the readers exist
- cross-compiling does not reproduce it
basics
~20 sThe publish was never synchronized, so no happens-before edge orders the write against the workers' reads. amd64's strong hardware ordering hid the bug; arm64 is weakly ordered and lets the stale value be observed. The race was always there.
solid answer
~50 sNothing in the program establishes a happens-before edge between the goroutine that assigns the package-level pointer and the goroutines that read it, so this is a data race and always was one. amd64 is strongly ordered, so a plain store usually becomes visible promptly and in program order, which makes the window vanishingly small; arm64 is weakly ordered, so both the compiler and the CPU may delay or reorder the store and readers genuinely observe the old value — or a non-`nil` pointer to a struct whose fields still read as zero. The fix is an ordering edge, not a barrier you hand-roll: assign the config before the `go` statements that start the workers (the `go` statement orders everything sequenced before it), publish it by closing a channel the workers receive from, or guard it with a `sync.Mutex` or `sync.RWMutex`, since an `Unlock` is synchronized before the next `Lock` returns.
code
go · 8 linesvar cfg *Config // read by every worker
func startup() {
go func() { cfg = loadConfig() }() // no edge to the readers
for i := 0; i < 8; i++ {
go worker() // worker reads cfg
}
}go deeper
Take away the headline: code that works on one machine can still be wrong. Shared state written by one goroutine and read by another needs a channel, a lock, or to be set before the readers start.
Explain the two layers — the missing happens-before edge in the program, and the weaker hardware ordering that exposed it — and name a concrete fix that uses Go's own synchronization rather than a barrier.
Show a diagnosis method that produces evidence: a reduced harness run many times on real arm64 hardware, an understanding that cross-compilation does not reproduce it, and a fix chosen for the value's actual write pattern.
Own the fleet-level consequence: once architectures are mixed, an amd64-only test pipeline is structurally blind to this class of bug, and the policy question is what concurrency testing you require on each architecture before a service is allowed to run there.
## What broke, and what did not Nothing broke on arm64. The program contained a data race on both machines: a package-level pointer written by one goroutine and read by others with no happens-before edge between the write and the reads. Go's memory model makes no promise at all about such a read, so the amd64 build was not correct — it was lucky. The reason the luck ran out is the hardware memory model. amd64 is strongly ordered (a store buffer with total store order): plain stores become visible to other cores in program order and reasonably promptly, so an unsynchronized publish usually looks like it works. arm64 is weakly ordered: stores may become visible to other cores late and out of order relative to one another unless the code emits the appropriate barriers — which the compiler only emits where the memory model requires it, that is, at real synchronization operations. Two distinct symptoms follow: - A worker reads the pointer and still sees `nil`, long after the assignment ran. - A worker sees a non-`nil` pointer but a struct whose fields still read as their zero values, because the store publishing the pointer became visible before the stores that initialised the struct it points at. The second one is the more alarming and the more instructive: without an edge, "the pointer is set, therefore the object is built" simply does not hold. ## The compiler is the other half of the problem It is tempting to file this as "an arm CPU thing" and reach for a barrier. That is the wrong model. Because Go's memory model gives racy programs no guarantees, the *compiler* is also free to act: it can hoist a repeated load of a package-level variable out of a loop and keep it in a register forever. A worker polling an unsynchronized flag can therefore spin for eternity on amd64 too. Any fix that addresses only the hardware leaves that failure mode in place. ## Confirming it before you change anything This class of bug is timing- and hardware-dependent, so a single green run proves nothing in either direction. What actually generates evidence: - Run on **real arm64 hardware**, not cross-compiled. `GOARCH=arm64 go build` produces an arm64 binary but executing it on the amd64 host — under emulation or not at all — tells you nothing about that host's memory ordering. - Reduce to a small harness that starts the publisher and the readers and asserts the value each reader observed, then run it many times: `go test -run TestPublish -count=1000`. A visibility bug shows up as a handful of failures across a thousand runs rather than a clean reproduction. - Vary the load. Adding readers, or running the harness on a busy machine, widens the window; a quiet machine can hide it entirely. - Treat "it has passed ten thousand times on amd64" as no evidence. Absence of an observed race is not the absence of a race. ## The fixes, in the order to prefer them **1. Publish before you start the readers.** If the config is read-mostly and written once at startup, assign it in package initialization or in `main` *before* the `go` statements that launch the workers. Package-level initialization completes before `main.main` starts, and the `go` statement is synchronized before the new goroutine's first line — so the ordering is free and there is no lock on the hot read path. This is usually the right answer and it also removes the concurrency from the design rather than managing it. **2. Hand it over a channel.** If the value genuinely cannot be ready before the workers start, have the loader close (or send on) a channel the workers receive from before their first read. Closing is synchronized before every receive that observes the closure, so it publishes to any number of readers at once. **3. Guard it with a mutex.** If the value is replaced during the process's life, a `sync.RWMutex` around the pointer is correct and boring: an `Unlock` is synchronized before the next `Lock` returns, and `RUnlock`/`RLock` give the same ordering for readers. The read cost is real but so is the correctness. What is *not* a fix: adding a `time.Sleep` before the first read, making the variable larger or smaller, re-reading it twice, or declaring it `var` at package scope in the belief that package variables are somehow special. None of those creates a happens-before edge. ## The judgment to carry away When a port to a new architecture surfaces intermittent staleness, the first hypothesis is not a toolchain bug — it is that a pre-existing race has become observable. Fix it as a missing ordering edge, in Go's own vocabulary, and the fix is correct on every architecture including the one where the bug never showed.
- Why is a long history of green amd64 runs not evidence that the code is correct?Because the guarantee comes from Go's memory model, not from an architecture. amd64's total store order makes an unsynchronized publish usually visible in program order, so the race almost never manifests — but the program has no happens-before edge, so the compiler and any weaker hardware are free to break it. Absence of an observed failure is not absence of a race.
- Can the compiler alone break this on amd64, with no reordering by the CPU?Yes. In a racy program the compiler may hoist a repeated load of a package-level variable out of a loop and keep it in a register, so a goroutine polling an unsynchronized flag can spin forever even on strongly ordered hardware. That is why a memory barrier is not the fix — an ordering edge in Go's own terms is.
- The config must be replaced at runtime, not just set once. What now?Then it needs a real lock: a `sync.RWMutex` with readers taking `RLock` and the reloader taking `Lock`. An `Unlock` is synchronized before the next `Lock` returns, and the same holds for the read side, so each reader observes a fully constructed value. Keep the critical section to reading or swapping the pointer, not to using it.
- How would you reproduce it in CI once the fleet is mixed?Run the concurrency-sensitive tests on a real arm64 runner, not a cross-compiled binary, and repeat them: `go test -count=1000` on the reduced harness surfaces a handful of failures where a single run shows none. Keep the arm64 job mandatory, since an amd64-only pipeline is structurally blind to this class of bug.
saying these in an interview costs you the question
- Calls it an arm64 code-generation bug in the compiler
- Says only the arm64 build needs fixing
- Adds a sleep before the first read of the pointer
- Assumes a pointer-sized store is atomic and therefore ordered
- Treats ten thousand green amd64 runs as proof of correctness
- Reaches for a hand-rolled memory barrier instead of an ordering edge