skip to content

A raw syscall.Read in a file-tree indexer intermittently returns EINTR — why, and how must the caller handle it?

level: seniorimportance: should knowfreq 34%

answer

  1. nothing actually failed
  2. a signal arrived mid-call
  3. the runtime itself sends signals
  4. restartability differs by platform and call
  5. retry with the bytes not yet transferred

basics

~20 s

EINTR means a signal reached the thread before the call finished, and Go's runtime preempts goroutines with signals. The caller must retry with the bytes not yet transferred, handling short reads in the same loop, rather than reporting failure.

solid answer

~50 s

EINTR is not a real I/O failure: a signal arrived while the thread was blocked in a slow system call, so the kernel returned early. Go's own runtime is a signal source — asynchronous preemption and CPU profiling both deliver signals to running threads — so a goroutine calling `syscall.Read` directly can see EINTR even though nothing external signalled the process. Whether the kernel restarts the call automatically depends on the platform and the specific call, which is why the bug shows up on the Linux build box and never on the developer's Mac. The fix is a retry loop: `for { n, err := syscall.Read(fd, buf); if err == syscall.EINTR { continue }; ... }`, retrying with the bytes not yet transferred. Use `errors.Is` if the error may have been wrapped. The standard library's `os` and `net` wrappers already retry internally, which is the strongest argument for using them instead of a bare descriptor.

code

go · 17 lines
go
func readFull(fd int, buf []byte) (int, error) {
	total := 0
	for total < len(buf) {
		n, err := syscall.Read(fd, buf[total:])
		if err == syscall.EINTR {
			continue
		}
		if err != nil {
			return total, err
		}
		if n == 0 {
			return total, io.EOF
		}
		total += n
	}
	return total, nil
}

go deeper

for a junior

Know that EINTR means a signal interrupted the call and nothing actually failed, and that the response is to retry rather than to report an error upward.

for a middle

Explain where the signals come from in a Go process — runtime preemption and CPU profiling, not just external signals — and write the retry loop correctly, advancing past bytes already transferred.

for a senior

Diagnose it as a platform-dependent truncation bug: reproduce under load or with profiling on, confirm with a system-call trace, then fix it in one helper and lock the behaviour down with a test that injects the errno.

for a principal

Decide the policy: which parts of the codebase are allowed to touch raw descriptors at all, so that interruption handling lives in one reviewed place instead of being re-derived, badly, by each team that needs a syscall.

## What EINTR actually means When a thread blocks in a *slow* system call — reading from a pipe, socket, terminal or a device that may not have data ready — and a signal is delivered to that thread, the kernel can abandon the call and return `EINTR`. Nothing failed. No data was lost. The call simply came back early so the signal handler could run. Whether it returns early at all depends on how the handler was installed and on the specific call. A handler installed with `SA_RESTART` lets the kernel transparently restart many calls. But restart is not universal: calls with a timeout, I/O on descriptors with a receive timeout set, and several polling and sleeping calls are not restartable even with `SA_RESTART`, and the exact list differs between Linux, macOS and the BSDs. That per-platform difference is why this is a portability bug, not just a robustness bug. ## Why Go programs see it without any external signal The usual mental model — "we do not use signals, so EINTR cannot happen" — is wrong for Go, because the runtime itself sends signals: - **Asynchronous preemption.** Since Go 1.14 the runtime interrupts a goroutine that has been running too long by sending a signal (SIGURG) to its thread, so the garbage collector and scheduler are not held hostage by a tight loop. The Go 1.14 release notes call out the consequence explicitly: programs using `syscall` or `golang.org/x/sys/unix` will see more slow system calls fail with EINTR. - **CPU profiling.** Turning on a CPU profile delivers SIGPROF at high frequency to running threads. A service that is fine normally starts returning EINTR the moment someone profiles it in production. So the frequency is a function of load, of GC pressure, of how many goroutines are running, and of whether profiling is on. All four differ between a laptop and a busy production host, which is exactly the shape of "works on my machine". ## The correct handling Retry. That is the whole answer, but the details matter: ```go func readFull(fd int, buf []byte) (int, error) { total := 0 for total < len(buf) { n, err := syscall.Read(fd, buf[total:]) if err == syscall.EINTR { continue // interrupted before transferring; retry } if err != nil { return total, err } if n == 0 { return total, io.EOF } total += n } return total, nil } ``` Four points hide in there: 1. **Retry immediately.** EINTR is not a transient network failure; backoff and jitter are wrong here and only add latency. 2. **Retry with what is left.** A call that was interrupted after transferring bytes reports those bytes; restarting from the beginning of the buffer would duplicate data. Advance the slice. 3. **A short read is a separate concern.** `read` is allowed to return fewer bytes than asked for with no error at all. Code that only handles EINTR and assumes a full buffer otherwise still corrupts data — and this is the failure that looks identical in production, so handle both in the same loop. 4. **Unwrap if you wrapped.** Comparing `err == syscall.EINTR` is fine directly off the wrapper, but as soon as the error travels through a `fmt.Errorf("...: %w", err)` you need `errors.Is(err, syscall.EINTR)`. Writes need the identical treatment: an interrupted or partial `write` must resume from the offset already sent. ## Diagnosing it The symptom in a file-tree indexer is a file whose contents are silently truncated at an arbitrary offset, on one platform, under load. Three things pin it down: - `strace -f -e trace=read,write` on Linux (or `dtruss` on macOS) shows the interrupted call and the signal that interrupted it, interleaved. - A CPU profile makes the bug *more* frequent rather than less, which is a strong tell: the diagnostic tool is itself a signal source. - A unit test that fabricates the condition: have the wrapper return `syscall.EINTR` from a seam and assert with `errors.Is(err, syscall.EINTR)` that the loop retried rather than surfaced it. This is the test that keeps the fix from being deleted later as "redundant". ## The best fix is often not to be there The standard library already solved this: `os.File` and `net.Conn` retry EINTR internally and, for pollable descriptors, hand the wait to the runtime's network poller so the goroutine parks instead of blocking a thread. If you took a raw descriptor only to call `read` on it, adopt it with `os.NewFile` and use the `os.File` methods. Keep raw calls for what genuinely has no standard-library wrapper — and put the retry loop in exactly one helper that everything else in the package goes through.

  • Why does this bug appear on the Linux CI box but never on the developer's machine?
    Two independent reasons. Which calls the kernel restarts automatically after a signal differs between Linux, macOS and the BSDs, so the same code is interruptible on one and not the other. And the signal rate depends on load, goroutine count, GC activity and whether CPU profiling is on — all higher on a busy build or production host than on an idle laptop.
  • Should the retry use exponential backoff, as you would for a failed network call?
    No. EINTR reports that a signal preempted the call, not that a resource is unavailable, so there is nothing to wait out. Backoff only adds latency and, under a steady signal rate such as active CPU profiling, can starve throughput. Retry immediately, and make sure the loop still has an exit — a cancelled context or a closed descriptor should break out rather than spin forever.
  • Beyond EINTR, what other partial-transfer case must the same loop cover?
    A short read or short write: the call succeeds with no error but transfers fewer bytes than the buffer holds, which is normal for pipes, sockets and terminals. Code that retries only on EINTR and otherwise assumes a full buffer silently truncates data. Track the running total and slice the remaining buffer on every iteration, which handles both cases with one piece of logic.
  • How would you avoid writing this loop at all?
    Adopt the descriptor with `os.NewFile` and use `os.File`'s methods, or use `net.Conn` for sockets. The standard library retries EINTR internally and, for pollable descriptors, parks the goroutine on the runtime's network poller instead of blocking an OS thread. Reserve raw calls for facilities with no standard-library wrapper, and funnel those through a single helper.

It is a phone call cut off because the doorbell rang: nobody hung up on you, and the right response is to dial back and resume, not to declare the conversation failed.

saying these in an interview costs you the question

  • Treats EINTR as a fatal I/O error and aborts
  • Adds exponential backoff before retrying an interrupted call
  • Retries from the start of the buffer, duplicating transferred bytes
  • Assumes only external signals like SIGINT can cause it
  • Handles EINTR but still assumes read fills the whole buffer
  • Says SA_RESTART makes every system call restart automatically