One destination in an io.MultiWriter starts failing mid-transfer. What happens to the others?
answer
- the fan-out is strict, not best effort
- argument order becomes a priority order
- no rollback for the sinks already written
- a short write turns into io.ErrShortWrite
- optional sinks need a wrapper that always says len(p), nil
basics
~10 sThe fan-out stops at the failing destination and returns its error, so destinations after it never get those bytes and the whole transfer aborts. A short write with no error becomes io.ErrShortWrite.
solid answer
~50 s`io.MultiWriter` writes to its destinations in argument order and checks each result. On the first non-nil error it returns immediately; if a destination accepts fewer bytes than it was given while reporting success, it returns `io.ErrShortWrite`. Either way the destinations after that point are skipped for that call, and the copy loop above sees a write error and abandons the transfer entirely — so an optional audit sink or mirrored upload can take down the primary path it was only meant to shadow. There is no rollback either: destinations earlier in the list already have those bytes. The fix is to decide per destination whether it is essential. Essential ones go into the fan-out directly and early. Optional ones get wrapped in a best-effort writer that records the failure, stops using the broken sink, and always reports a full write, so the failure is surfaced after the transfer instead of during it.
code
go · 13 linestype bestEffort struct {
w io.Writer
err error // first failure, checked after the transfer
}
func (b *bestEffort) Write(p []byte) (int, error) {
if b.err == nil {
if _, err := b.w.Write(p); err != nil {
b.err = err // stop using the broken sink
}
}
return len(p), nil
}go deeper
Know the headline: a fan-out write is all or nothing per call, and one bad destination makes the whole write fail. That is enough to avoid casually adding a sink to a transfer that matters.
Explain the loop's exact rules — stop on the first error, turn a short write into io.ErrShortWrite, skip everything after — and why destinations earlier in the list keep the bytes they already took.
Show the production reasoning: classify each destination as essential or optional, wrap the optional ones so they cannot abort the transfer, surface their recorded failures afterwards, and prove all of it with a stub that fails on demand.
Own the policy question of what an added observer is allowed to cost. Decide, for the codebase, whether observation may ever fail a primary path, and make that a reviewable rule rather than a choice each author makes at the call site.
## One sink shouldn't take down the transfer The scenario is common: an existing transfer works, and someone adds a second destination to it — a mirror, an audit copy, a checksum sent to another service, a debug capture. `io.MultiWriter` makes that a one-line change, and that is exactly why the failure semantics catch teams by surprise. ### What the fan-out actually does on failure A `Write` on the value returned by `io.MultiWriter` loops over its destinations in the order they were passed. For each one it calls `Write` and inspects the result: - If the destination returns a non-nil error, the loop returns that error immediately. - If the destination returns `n < len(p)` with a nil error, the loop returns `io.ErrShortWrite`. - Otherwise it continues to the next destination, and returns `len(p), nil` when all of them have accepted the full slice. Three consequences follow. First, **there is no isolation**: any single destination can fail the whole call. Second, **there is no rollback**: destinations before the failing one already consumed those bytes, so after an abort your sinks disagree about how much of the stream they hold. Third, **order is visible**: destinations after the failure never see that slice at all, which means the argument order encodes a priority you may not have intended to express. Above the fan-out, whatever is driving the transfer sees a write error and stops. The observable symptom is a transfer that fails with an error naming a subsystem nobody thought was on the critical path — "broken pipe" from a log shipper, a permission error from an audit file — while the primary destination was perfectly healthy. ### Diagnosing it This class of bug is nearly impossible to reproduce by accident, because in development every destination is a local file or an in-memory buffer and neither ever fails. Reach for a stub destination instead: a type whose `Write` returns an error after the first N calls, or accepts only part of each slice. Assemble the real fan-out with the stub in one slot, run a real transfer through it, and assert what you expect — that the primary sink is complete, or that the transfer failed, whichever the design says. That test is the whole design conversation made executable, and it belongs in the repository next to the code that composes the writers. The same stub proves the short-write branch, which is the subtler half. A destination that accepts part of a slice and reports no error is not obviously broken; it simply cannot take more right now. The fan-out converts it into `io.ErrShortWrite`, and the transfer dies with an error message that points at nothing in particular. ### Designing for it Start by classifying each destination. **Essential** means the transfer is worthless if this sink misses bytes — the actual file being written, the response going back to a caller, a ledger you must not gap. **Optional** means you would like the bytes there, but you would rather complete the transfer than fail it. Essential destinations go straight into `io.MultiWriter`, listed before the optional ones, so that if something does fail they have already been written for that slice. Optional destinations get a wrapper that absorbs their failures: on an error it records the cause, marks itself broken so it stops calling the dead sink, and — critically — returns `len(p), nil` so the fan-out and the copy above it carry on. The recorded error is checked after the transfer completes and reported through whatever channel makes sense: a log line, a metric, a returned partial-success value. What you must not do is discard it silently, because "the audit copy has been failing for three weeks" is a bad thing to discover later. The other option for an optional sink is to decouple it entirely: put a hand-off in the fan-out that pushes chunks onto a bounded channel, and let a separate goroutine drain that channel into the real destination. Now the sink's failures and its latency are both its own problem, at the cost of a buffer and a policy for what to do when it fills — drop, block, or count. That is worth it for a slow or remote sink and overkill for a local file. ### The judgment to show The wrong instinct is to treat the strict behaviour as a bug in the standard library. It is the safe default: silently dropping bytes from a destination you explicitly asked for would be far worse, and a library cannot know which of your sinks is optional. Encoding that knowledge — with a wrapper, an order, or a decoupling — is the caller's job, and being explicit about it is the difference between an observability addition and a new failure mode in the primary path.
- After an abort mid-transfer, what state are the destinations in?Divergent. Destinations before the failing one accepted that slice, the failing one accepted part or none of it, and the ones after it saw nothing — and nothing is undone. If the sinks are meant to be identical copies you now need an explicit reconciliation step: truncate and restart the mirror, mark it incomplete, or record the byte offset at which it diverged so a later job can repair it.
- Does listing the optional sink last in io.MultiWriter solve the problem?Only partly. It guarantees the essential destinations receive each slice before the optional one is attempted, so no essential sink is skipped. But the failing sink still returns an error out of the fan-out, the copy loop above still stops, and the transfer still fails. Ordering limits the damage per call; it does not keep the transfer alive. That takes a wrapper that reports success regardless.
- How do you reproduce this in a test when local files never fail?Write a stub destination whose `Write` fails after a set number of calls, or accepts only part of the slice with a nil error. Put it in one slot of the real `io.MultiWriter` and run a transfer through. Assert the outcome you designed for: with a best-effort wrapper the transfer completes and the recorded error is non-nil; without one, it aborts and the later destinations are short. Cover the short-write branch too — it is the one people miss.
- When would you decouple a sink with a goroutine instead of wrapping it?When it is slow or remote, not just unreliable. `io.MultiWriter` writes sequentially on the calling goroutine, so a sink with network latency adds that latency to every chunk of the primary transfer. Feeding it through a bounded channel drained by its own goroutine removes it from the critical path, at the cost of memory and an explicit policy for a full buffer — drop and count, or block. For a local file it is not worth the machinery.
saying these in an interview costs you the question
- Expects the fan-out to skip a failing destination and continue
- Assumes the destinations stay in sync after an aborted write
- Thinks a short write with a nil error is treated as success
- Puts an unreliable remote sink directly into the fan-out
- Swallows the optional sink's error and never reports it anywhere