Concurrent lookups on one *os.File return each other's bytes, yet go test -race is clean. Why, and how do you fix it?
answer
- well-formed records, wrong ones
- two calls, one cursor
- the shared state is not in Go memory
- the detector only sees what it can instrument
- make the position an argument
basics
~20 sSeek and Read are two operations over one file cursor shared by every goroutine holding that handle, so calls interleave and each reader lands on someone else's position. The race detector watches Go memory, not that cursor, so it reports nothing. Read positionally with os.File.ReadAt.
solid answer
~50 sThe symptom is well-formed but wrong data: each lookup returns a complete fixed-width record, just not the one it asked for, and it only happens under concurrency. The cause is that `Seek` and `Read` are two separate calls against a single cursor belonging to the open file, so `SeekA, SeekB, ReadA, ReadB` gives both goroutines the wrong bytes. `go test -race` stays silent because there is no unsynchronised access to a Go variable — the shared state is the kernel's file offset, which the detector cannot see. That is the general lesson: the race detector finds unsynchronised memory access on paths a test actually executes, and proves nothing about logical races outside Go memory. The fix is to stop sharing a cursor: call `f.ReadAt(buf, off)`, which takes the position as an argument and is defined for parallel use, or hand each caller its own `io.NewSectionReader(f, off, n)` when it needs a reader. A mutex around the `Seek`/`Read` pair is also correct but serialises every lookup.
code
go · 9 lines// BROKEN when called from more than one goroutine.
func readRecord(f *os.File, i int64) ([]byte, error) {
if _, err := f.Seek(i*recordSize, io.SeekStart); err != nil {
return nil, err
}
buf := make([]byte, recordSize)
_, err := f.Read(buf) // another goroutine may have moved the cursor
return buf, err
}go deeper
Take away that a file cursor belongs to the open file and is shared by every goroutine using that handle, so a seek followed by a read is not one indivisible step.
Explain the interleaving concretely and name the positional alternative. Be able to say why the returned data is well formed but wrong, which is what makes the bug survive validation.
Diagnose it without a tool telling you: reason about what the race detector instruments, design the content-asserting test, and rank the fixes by throughput and descriptor cost rather than stopping at the first one that works.
Set the expectation that clean tooling output is evidence, not proof, and that shared mutable position in an API is a concurrency ceiling you are shipping to every caller. That review rule prevents the class, not just this instance.
## The failure An index-file reader serves lookups from computed offsets: fixed-width records inside one large file, one open `*os.File` shared by all the goroutines handling requests. Under load, callers start getting records that belong to other callers. Nothing is malformed — every response is a valid record — so schema validation and checksums over the record body pass. Only a caller that compares the record's own key against the key it asked for notices. The broken shape is the obvious one: ```go f.Seek(i*recordSize, io.SeekStart) f.Read(buf) ``` Two calls; one cursor. The cursor is state on the open file, not on the goroutine. The interleaving `SeekA, SeekB, ReadA, ReadB` is ordinary and gives goroutine A the bytes at B's offset, and B the bytes just after them. The bug is invisible at one request per second and constant at a hundred. ## Why the race detector says nothing This is the part that separates candidates. Go's race detector instruments **memory accesses in Go code** and reports a pair of accesses to the same address from different goroutines with no happens-before relationship between them. Here there is no such pair: `buf` is local to each call, the `*os.File` value itself is only read, and the mutable state everyone is fighting over — the file offset — lives in the kernel, behind a system call. Internally the runtime's poller even serialises the individual read operations on a descriptor, so the two system calls do not overlap; they simply happen in the wrong order relative to the seeks. Nothing in Go memory is accessed unsafely, so a clean `-race` run is accurate and useless. Generalise it: `-race` is a dynamic detector. It reports races that actually occurred on the paths the test exercised, so a clean run over an untested path proves nothing, and a logical race outside Go memory — a file cursor, a database row, a file on disk, an external counter — is invisible to it by construction. ## Finding it deliberately Since the detector will not help, the test has to assert content rather than await a report. Write records whose payload encodes their own index, then run many goroutines looking up known indices and assert that each returned record identifies itself as the one requested. Run it with enough concurrency and iterations and the mismatch shows up immediately. Two supporting signals: the failure disappears entirely when you wrap the pair in a mutex, and it disappears when you drop to a single goroutine — both point at unsynchronised use of shared position rather than at corruption on disk. ## The fixes, in order of preference **Positional reads.** `f.ReadAt(buf, off)` carries the offset in the call, so there is no shared position to interleave. The `io.ReaderAt` contract explicitly allows parallel calls on the same source, and `*os.File` implements it with the kernel's positional read. This is one line, needs no lock, and needs no second file descriptor. **A section reader per caller.** When the consumer wants an `io.Reader` — a decoder, a parser — give it `io.NewSectionReader(f, off, n)`. The returned `*io.SectionReader` has its own cursor, reports `Size()`, and ends with `io.EOF` at the end of its section rather than the end of the file. Constructing one per lookup is cheap and still uses the single shared descriptor. Do not share one section reader between goroutines: it has a cursor, and that is the problem you just left. **A mutex around the pair.** Correct, and sometimes the only option if you are stuck behind an interface that only offers `io.ReadSeeker`. The cost is that every lookup now queues on a resource the operating system would happily serve in parallel, so it caps throughput for a reason that has nothing to do with the disk. **One open file per goroutine.** It works, and it trades a correctness problem for a resource problem: descriptors are a per-process limit, and a handler that opens a file per request will find it. Reach for it only when you need independent cursors for a genuinely sequential workload. ## The review rule to take away Any two-call sequence where the first call sets state that the second consumes is an atomicity question the moment more than one goroutine can reach it. `Seek` then `Read` is the file-shaped instance. When you see one, ask whether the API offers a single-call form that takes the state as an argument — here it does, and using it makes the concurrency question disappear rather than answering it with a lock. ## What an interviewer is checking Whether you can explain a clean `-race` run instead of trusting it, whether you know positional reads exist and why their contract permits parallel use, and whether you can rank fixes by what they cost rather than naming the first one that works.
- How would you write a test that catches this reliably?Build the file so each record encodes its own index, then run many goroutines requesting known indices and assert the returned record identifies itself. The failure is a value mismatch, not a detector report. Verify the diagnosis by confirming it vanishes with a mutex and with a single goroutine.
- What does a clean go test -race run actually prove?Only that no unsynchronised access to Go memory occurred on the code paths that test executed. It is a dynamic detector, not a static proof: untested paths are unexamined, and logical races over state outside Go memory — a file cursor, a row in a database — are invisible to it.
- Why not just open the file once per goroutine?It gives each one an independent cursor and works, but file descriptors are a per-process limit, so a handler doing this per request eventually exhausts them and adds an open and close to every lookup. Positional reads solve the same problem with one descriptor and no extra syscalls.
- The consumer insists on an io.Reader. Does that force you back to the lock?No. `io.NewSectionReader(f, off, n)` returns a reader with its own cursor over a byte range of the same handle, so each caller gets private position state and the shared file's cursor is never touched. Create one per lookup — just never share a single one between goroutines.
saying these in an interview costs you the question
- Says a clean -race run proves the code is concurrency-safe
- Puts the mutex around Read only, leaving Seek outside it
- Assumes each goroutine has its own file cursor
- Blames disk corruption or the file system for the wrong bytes
- Opens a new file descriptor per request as the first fix