Why doesn't a blocked os.File.Read return when the context.Context governing the job is cancelled, and what do you do instead?
answer
- The signature has no context in it
- Already parked in the kernel
- Shrink the unit, check between units
- Deadlines need a pollable descriptor
- Close refcounts, so it does not interrupt
basics
~20 sos.File.Read has no context parameter and the read is already in the kernel, so it returns only on data, EOF or error. Cancel between reads instead: read in bounded chunks and check ctx.Err() on each iteration.
solid answer
~50 s`Read` takes a buffer, not a context, and once the goroutine is parked in the syscall nothing in Go can interrupt it. So you make the unit of work small enough to cancel between units: loop over fixed-size chunks and check `ctx.Err()` (or select on `ctx.Done()`) before each `Read`. For descriptors the runtime poller owns - pipes, terminals, sockets - `os.File.SetReadDeadline` works, and a watcher goroutine that sets a deadline in the past on cancellation unblocks the read with `os.ErrDeadlineExceeded`. On an ordinary disk file that same call returns `os.ErrNoDeadline`, because regular files are not pollable. Closing the file from another goroutine is not an escape hatch either: `os.File` reference-counts the descriptor so it is not really closed until the in-flight read finishes, which is exactly the case you were trying to break out of.
code
go · 20 linesfunc copyCtx(ctx context.Context, dst io.Writer, src io.Reader) error {
buf := make([]byte, 32*1024)
for {
if err := ctx.Err(); err != nil {
return err // cancelled between chunks, not inside one
}
n, err := src.Read(buf)
if n > 0 {
if _, werr := dst.Write(buf[:n]); werr != nil {
return werr
}
}
if err == io.EOF {
return nil
}
if err != nil {
return err
}
}
}go deeper
Know that Read's signature takes only a byte slice, so no context can reach it, and that the standard workaround is reading in chunks with a check between them.
Explain why a parked syscall cannot observe a closed Done channel, and show the loop with the check placed before each Read. Mention that deadlines are the alternative and only work on pollable descriptors.
Diagnose it from the outside: the job reports cancelled, descriptors stay open, goroutine stacks show everyone parked in Read. Then choose between chunking, a deadline bridge, and accepting an uninterruptible operation you bound at a higher layer.
Own the honesty rule: if some work cannot be interrupted, the system must not report it as stopped. Decide what the operator is promised on cancel, and make sure abandoned goroutines cannot write into state the caller has already torn down.
## Why the read is deaf `func (f *os.File) Read(b []byte) (n int, err error)` has no context in its signature, and that is not an oversight that a wrapper could fix. When the goroutine calls it, the runtime hands the thread to the kernel's `read`. At that point the goroutine is not running Go code, it is not in a `select`, and there is no scheduler point at which a closed `Done` channel could be noticed. Cancellation is a closed channel; a parked syscall cannot see a channel. The cost shows up in exactly one shape of program: a batch job walking a tree of large input files. Each file read is fast on local disk and arbitrarily slow on a network mount. When an operator cancels the job, the database writes and outbound calls stop at once, and the file loop does not. The job looks hung; the descriptor count in `lsof` stays flat instead of falling, because the reads never return and the deferred `Close` calls never run. ## The technique that actually works: cancel between units Make the blocking unit small, then check the context between units. Reading a 4 GB file in 32 KiB chunks means the longest you can be uncancellable is one chunk read, which on any healthy device is sub-millisecond. The check goes at the top of the loop body, before the read - not after, and not once before the loop. `ctx.Err()` is the cheap form (it returns non-nil exactly when Done is closed); a `select` with a `default` case does the same thing more verbosely. Return the context error so the caller can distinguish a cancelled job from a corrupt input, and make sure the `defer f.Close()` you wrote at open time is what actually releases the descriptor - cancellation will never do that for you. This chunking idea generalises: whenever a library gives you a blocking call with no context, look for a variant that does bounded work (`Read` on a small buffer, one row at a time, one directory entry at a time) and put the check in the loop that drives it. ## Deadlines, and why they do not help on a regular file `os.File` has `SetDeadline`, `SetReadDeadline` and `SetWriteDeadline`, taking an absolute `time.Time`. When the descriptor is one the runtime's network poller can own - a pipe, a terminal, a socket - a deadline works, and a deadline set in the *past* unblocks an already-parked read immediately with `os.ErrDeadlineExceeded`. That gives you a real bridge from cancellation to I/O: a small goroutine that waits on `ctx.Done()` and then sets a past deadline. On an ordinary file on disk the same call returns `os.ErrNoDeadline` ("file type does not support deadline"). Regular-file I/O does not go through the poller, so there is nothing to time out. This is the fact that catches people: they write the deadline bridge, test it against a pipe, and it silently does nothing in production where the input is a file. Always check the error from `SetReadDeadline`. ## Why closing the file is not the escape hatch The next idea people try is to have the cancel path call `f.Close()` and let the read fail. `os.File` reference-counts its descriptor precisely so a descriptor cannot be closed out from under an in-flight operation and then reused by an unrelated `open` - a real and nasty class of bug. So the `Close` marks the file closed for *future* operations and the actual `close(2)` is deferred until the outstanding read finishes. The blocked read is not unblocked; you have added a second stuck goroutine. ## What to say about the shape of the fix There is no general way to interrupt a syscall from Go, so the honest answer at senior level is that you choose one of three: make the units small and check between them (almost always right); use deadlines when the descriptor is pollable; or accept that the operation is uninterruptible and bound it at a higher layer - let the goroutine finish and leak briefly while the job returns, rather than pretending you stopped it. What you must not do is claim the work stopped when only the visible layers did. If you leave a goroutine behind, say so, and make sure it cannot write into state the caller has already torn down.
- What does os.File.SetReadDeadline return for an ordinary file on disk?`os.ErrNoDeadline`. Regular files are not handled by the runtime poller, so there is nothing to time out. Deadlines only apply to pollable descriptors such as pipes, terminals and sockets, where setting one in the past unblocks a parked read with `os.ErrDeadlineExceeded`.
- Can another goroutine call f.Close() to break the stuck read out?No. `os.File` reference-counts the descriptor so it cannot be closed while an operation is in flight and reused by an unrelated open. The `Close` takes effect for later calls, but the real close waits for the read, so the blocked goroutine stays blocked.
- How small should the chunks be?Small enough that one read's worst-case latency is an acceptable cancellation lag, large enough to keep syscall overhead down. Tens of kilobytes is the usual compromise, and it is what `io.Copy` uses internally. Tune it by the device: a network mount argues for smaller chunks than local NVMe.
saying these in an interview costs you the question
- Claims wrapping the read in a select makes it cancellable
- Expects SetReadDeadline to work on a regular file
- Thinks closing the file from another goroutine unblocks the read
- Checks ctx.Err() once before the loop instead of each iteration
- Says the runtime will interrupt the syscall on cancellation