skip to content

os.WriteFile rewrites a file in place, so how can a concurrent os.ReadFile of it return half a document and a nil error?

level: seniorimportance: should knowfreq 33%

answer

  1. one of the three open flags does the damage
  2. the file is briefly telling the truth
  3. a short read is a successful read
  4. the error surfaces in the wrong package
  5. checking the size first only moves the window

basics

~20 s

os.WriteFile truncates the file to zero before writing, so for a moment it really is empty or holds only a prefix. A reader that opens then reaches a genuine end of file, gets fewer bytes, and sees no error.

solid answer

~50 s

`os.WriteFile` opens with `O_TRUNC` and then writes. Between the truncation and the last byte, the file on disk is empty or a prefix — that is not a fiction, it is the file's real state, and any reader that opens it then sees exactly that. `os.ReadFile` reads until end of file and treats end of file as success, so it returns a short slice with a **nil error**. Nothing in the I/O layer flags the problem; the failure surfaces one layer up as a parser complaining about unexpected end of input, which sends people debugging the parser. Adding an `os.Stat` size check before the read does not help: the size you observe is already stale by the time you open, and the file can shrink or grow in between. The conclusion is that in-place rewriting is the wrong primitive for a file other readers touch at arbitrary times.

code

go · 9 lines
go
// Writer: truncates to zero at open, then writes.
if err := os.WriteFile("gen/model.go", out, 0o644); err != nil {
	return err
}

// Reader, running at the same moment in another process:
data, err := os.ReadFile("gen/model.go")
// err can be nil while len(data) is 0, or a prefix of either version,
// because reaching the current end of file is a successful read.

go deeper

for a junior

Know the shape of the trap: writing a file replaces it in place, so for a moment the file is shorter than either version, and a reader can see that.

for a middle

Explain the mechanism from the flags: O_TRUNC zeroes the length at open, the data follows in one or more writes, and a read that stops at the current end of file returns a nil error with fewer bytes.

for a senior

Diagnose it from the symptom — intermittent decoder failures that correlate with the writer, never reproducible afterwards — and reject the stat-then-read check on the grounds that it only relocates the window rather than closing it.

for a principal

Decide the rule for the system: which paths are allowed to be rewritten in place at all, which are contracts other processes read at arbitrary times, and what the team's default is when a partially written file would be visible or unrecoverable.

## The scenario A code generator emits `.go` files with `os.WriteFile`. Something else — a watcher, a build step, another service — reads those files whole with `os.ReadFile`. Occasionally the reader fails with a parse error partway through a file, and rerunning it makes the problem vanish. The file on disk, inspected afterwards, is perfectly fine. ## Why it happens `os.WriteFile` is not a swap. It is: ``` OpenFile(name, O_WRONLY|O_CREATE|O_TRUNC, perm) Write(data) Close() ``` `O_TRUNC` sets the file's length to zero at open time. Then the data goes out, possibly in more than one underlying write for a large slice. During that interval, the file's length is zero, then some prefix, then finally the full new contents. Those intermediate states are not a race in the abstract sense — they are what the file genuinely is at that instant, visible to anything that looks. There is no locking. Unix filesystems do not serialise a reader against a writer for you, and Go's `os` package adds none. The reader is not doing anything wrong; it is reading a file that is currently short. ## Why there is no error This is the part that makes the bug expensive. `os.ReadFile` reads until end of file, and end of file is its success condition — it returns `err == nil`, never `io.EOF`. If it reaches end of file after 300 bytes because that is all there is right now, that is a completely successful read of 300 bytes. So the I/O layer reports nothing. The error appears one level up, in whatever consumes the bytes: a JSON decoder saying unexpected end of input, a Go parser complaining about an unexpected end of file, a template failing to compile, a checksum mismatch. Every one of those points at the parser, not the filesystem, which is why teams lose hours to it. ## The signature of the bug - Intermittent, and correlated in time with the writer running. - Never reproducible under a debugger or a rerun, because by then the write has finished. - The file on disk is valid when you go and look. - The failure rate rises with file size — a larger payload means a longer window between truncation and the final write — and with how often the writer runs. - Two readers can disagree, one succeeding and one failing on the same path at the same second. To reproduce it deliberately, run the reader in a tight loop while the writer rewrites the file repeatedly, and assert on the length of what comes back rather than on the parse. You will see zero-length and prefix-length reads directly. ## Why an os.Stat size check does not fix it The instinct is to stat the file first and only read if the size looks right. That does not work, for two independent reasons. First, it is a check followed by a use, with a gap in between. The size `os.Stat` reports describes the file at the moment of the stat; by the time you open and read, the writer may have truncated it. You have moved the window, not closed it. Second, `os.ReadFile` does its own internal stat purely as an **allocation hint** — it sizes the initial slice from it and then reads to end of file regardless, growing if there is more and returning less if there is less. The size is never a contract on how many bytes you get. Files that report a size of zero, such as many synthetic files, still read correctly for exactly this reason. So even the number that `os.ReadFile` itself sees does not constrain the result. A length check *after* the read is more honest — "I expected at least N bytes" — but it is a heuristic, not a fix, and it cannot distinguish a truncated read from a legitimately shortened file. ## The same hazard inside one process Nothing about this requires two processes. Two goroutines, one calling `os.WriteFile` and one calling `os.ReadFile` on the same path, hit it identically. The Go race detector will not find it: there is no shared memory being accessed unsynchronised, only a shared file, and `-race` instruments memory, not filesystem ordering. Reasoning about the ordering yourself is the only tool. ## The real conclusion `os.WriteFile` is an excellent tool for a file you own exclusively: output nobody reads until you say so, a fixture written before a test starts, a file rewritten while the system is quiescent. It is the wrong tool the moment another reader can arrive at an arbitrary time, because it has no story for making the new contents appear as a unit — it deliberately destroys the old contents first. The second, quieter consequence of the same design: if the write fails midway, the old contents are already gone. So the exposure is not only to concurrent readers but to any crash during the write. Whenever the contents of a path must always be a complete, valid document, replacing the file as a whole rather than rewriting it in place is the property you need, and it is a different technique with its own mechanics. ## What an interviewer is checking That you can trace a parser-level symptom back to an I/O-level cause; that you know a short read is a successful read; that you do not reach for a stat-then-read check and call it solved; and that you can articulate what guarantee `os.WriteFile` never claimed to give.

  • Would calling os.Stat and comparing Size before os.ReadFile catch the truncated read?
    No. The size describes the file at the moment of the stat, and the writer can truncate between that call and your open — you have moved the window, not removed it. os.ReadFile also treats size only as an allocation hint and reads to whatever end of file it finds, so no observed size constrains the byte count you get back.
  • Why does the failure show up as a parse error rather than an I/O error?
    Because reaching end of file is os.ReadFile's success condition; a 300-byte read of a file that is currently 300 bytes long is entirely successful and returns a nil error. Only the consumer of those bytes notices that the document is incomplete, so the first visible symptom is in the decoder or parser.
  • Does this hazard exist between two goroutines in one process, and would the race detector find it?
    Yes to the first, no to the second. Two goroutines on the same path race exactly the same way. The race detector instruments memory accesses, not filesystem operations, so a `-race` build reports nothing here — there is no unsynchronised shared variable, only a shared file.
  • Beyond concurrent readers, what else does truncate-then-write expose you to?
    Loss. The truncation happens at open, before any new byte is written, so a crash, a full disk or a signal midway leaves a truncated file and no copy of the previous contents. The returned error tells you the write failed; it does not mean the file is unchanged.

saying these in an interview costs you the question

  • Assumes os.WriteFile replaces the file atomically
  • Blames the parser because the visible error is a parse error
  • Thinks a short read must return an error
  • Adds an os.Stat size check and declares it fixed
  • Expects the race detector to find a file-level race
  • Believes the OS locks a file while it is being written