skip to content

A go test -race failure names a write and a previous read on a shared []byte — how do you get from that report to the fix?

level: seniorimportance: should knowfreq 48%

answer

  1. read the creation frames first
  2. same address beats same variable name
  3. two slices, one backing array
  4. sharing mistake, not locking mistake
  5. hand off, copy, or lock every access

basics

~20 s

Read both stacks and the goroutine-creation frames to find who shares the backing array and where it was handed over. Then fix ownership: guard every access with the same lock, copy the bytes out, or hand the buffer off — locking only the writer leaves the race intact.

solid answer

~50 s

The report gives four facts: the offending write with its stack, the previous read with its stack, the address they share, and the `created at` frames for each goroutine. I start with the creation frames, because that is where the sharing was introduced — usually a `go` statement handing a slice to a worker while the caller keeps using it. The shared address matters more than the variable names: two differently named `[]byte` values that came from the same `append` or reslice are the same backing array, so the report can name two variables you thought were separate. Then I fix ownership rather than sprinkling a mutex: either the buffer is handed off and the sender stops touching it, or the reader gets a copy made under the lock, or every access — reads included — takes the same mutex. Locking only the writer silences nothing and fixes nothing.

code

go · 14 lines
go
type buf struct {
	mu   sync.Mutex
	data []byte
}

func (b *buf) add(p []byte) {
	b.mu.Lock()
	b.data = append(b.data, p...)
	b.mu.Unlock()
}

func (b *buf) snapshot() []byte {
	return b.data // unlocked read, and it escapes the lock entirely
}

go deeper

for a junior

Know that the report shows two accesses and that at least one is a write, and that the fix must cover both sides. Being able to point at the file and line in each stack is enough at this level.

for a middle

Explain why the shared address matters more than the variable names, and why a slice handed to another goroutine still aliases the caller's backing array after append or reslicing.

for a senior

Demonstrate ownership thinking: reconstruct who owned the bytes and when, choose between handoff, copy, and full locking on that basis, and be explicit that a green rerun is evidence only if the test still executes both sides.

for a principal

Push the discussion to how such sharing gets designed out — API shapes that do not hand out internal buffers, ownership conventions written into review standards — rather than accumulating mutexes on state that never needed to be shared.

## Take the report apart first A `go test -race` failure is unusually informative and most people read only half of it. There are four pieces of evidence. **The offending access.** `Write at 0x... by goroutine 21`, with a stack. This is what tripped the check. **The previous conflicting access.** `Previous read at 0x... by goroutine 7`, with its own stack. At least one of the two is a write, or there would be no report. **The address.** Both lines carry the same address, and that is a stronger statement than "the same variable". A `[]byte` is a header pointing at a backing array; two slices produced by reslicing, by `append` that did not reallocate, or by passing a slice into a function are different headers over the *same* memory. So the report legitimately pairs accesses through two variables with different names, and the fix has to be reasoned about in terms of the array, not the names. **The `created at` frames.** `Goroutine 21 (running) created at:` gives the stack of the `go` statement that started it. This is usually the most valuable part, because a data race is nearly always a *sharing* mistake rather than a locking mistake, and the `go` statement is where the sharing happened. ## Reconstruct the ownership story With those four facts, write down in one sentence who owns the bytes. Typical stories on a shared `[]byte`: - The caller filled a buffer, started a goroutine to process it, and then reused or appended to the same buffer for the next item. Both now write the same array. - A method returned the internal slice to a caller (`return b.data`), and the caller reads it after the lock inside the method was released. The lock is real and useless: the array escaped it. - Two goroutines append to the same slice field. `append` reads the length, may reallocate, and writes back the header — three racing operations, not one. ## Choose a fix by ownership, not by reflex There are only three honest fixes, and choosing the wrong one produces code that still races or code that is slower for no reason. **Hand it off.** One goroutine owns the buffer at a time. Send it on a channel and never touch it again on the sending side. This is the cheapest fix and the one that survives future edits, because there is no discipline to remember. **Copy.** If a reader needs a stable view, copy the bytes into a fresh slice under the same lock the writer uses, and return the copy. Returning the internal slice from under a lock is not a fix — the array outlives the critical section. **Lock everything.** If the buffer really is shared mutable state, every access — every read included — must take the same mutex. A frequent half-fix is to lock the mutating method and leave the accessor unlocked because "reads are safe". They are not: a read concurrent with an unsynchronized write is precisely what the detector reported. ## Confirm the fix the same way you found it Re-run the specific test under `-race`, narrowed with `-run` while you iterate. Two cautions. First, a green run after the change is evidence, not proof — the same coverage limits apply as before, so make sure the test still executes both sides of the sharing after your edit. It is easy to "fix" a race by accidentally removing the concurrency from the test. Second, resist the fix that only moves the access: copying a value into a local variable at the racing read does not help, because the racing load still happens. ## Two things not to conclude A race-free program is not automatically a correct one. Taking the same mutex around two operations separately still leaves a check-then-act bug; the detector has no opinion about that. And the goroutine that appears in the report is not necessarily the guilty one. The write may be perfectly legitimate; the bug is that something else was allowed to read the same array concurrently. Fixing the access that happens to be reported second is a common way to make the report move rather than disappear.

  • The accessor takes the mutex and returns b.data. Why does the race persist?
    Because the slice header it returns points at the shared backing array, and the caller reads that array after the mutex is released. The critical section protected the header copy, not the bytes. Either copy under the lock and return the copy, or hand ownership over so nobody else writes it.
  • The report names two slice variables you believed were unrelated. What explains that?
    They share a backing array. Reslicing, passing a slice to a function, or an `append` that fit in existing capacity all produce a second header over the same memory. The detector works on addresses, so it correctly pairs accesses that the variable names make look independent.
  • How do you confirm the fix without fooling yourself?
    Re-run the same test under `-race`, and check that the test still executes both sides of the sharing after your edit. A green run because you accidentally serialised the test proves nothing. It also helps to keep the failing test as a regression test rather than deleting it once it goes green.

saying these in an interview costs you the question

  • Locks only the writer because reads are safe
  • Returns the internal slice from under the mutex
  • Assumes two named slices cannot be the same memory
  • Ignores the goroutine creation frames in the report
  • Copies at the racing read and calls it fixed
  • Treats race-free as equivalent to correct