skip to content

Your racing download returns the winner's bytes while two mirror fetches are still running. What damage can those losers do?

level: seniorimportance: nice to knowfreq 33%

answer

  1. return does not end them
  2. their writes land after you left
  3. two writers, one destination file
  4. each attempt needs private state
  5. cancel on return, cleanup in the attempt

basics

~20 s

They run on after the caller has returned, so anything they touch is touched too late: a shared destination file gets a second writer, caller-owned memory gets written after the function returned, and their sockets and buffers stay held. Give each attempt private state and cancel the shared context on return.

solid answer

~50 s

Returning does not stop them, so the losing fetches finish minutes later and every side effect they still perform lands after the caller has moved on. Two kinds of damage: correctness — writing into a shared destination file or a caller-owned slice or struct field is a write with no synchronisation against the code that already consumed the winner's result, so you get a data race and, with a file, a corrupted archive built from two mirrors; and cost — each loser holds a connection, a response body and a whole archive in memory until it finishes. Contain both: give each attempt its own buffer or its own temporary file and promote only the winner's by rename, hand every attempt the same cancellable context and cancel it when you return so in-flight transfers abort, keep the result channel buffered so they can still exit, and log each attempt's outcome as it lands so you can see how long the losers really run.

code

go · 15 lines
go
func attempt(ctx context.Context, mirror, name string) result {
	f, err := os.CreateTemp("", "pkg-*")
	if err != nil {
		return result{mirror: mirror, err: err}
	}
	path := f.Name()

	sum, err := download(ctx, mirror, name, f) // honours ctx
	f.Close()
	if err != nil {
		os.Remove(path) // the loser cleans up its own file
		return result{mirror: mirror, err: err}
	}
	return result{mirror: mirror, path: path, sum: sum}
}

go deeper

for a junior

The key fact to hold onto: goroutines you started keep running after your function returns. Anything they write, they write later, so do not point them all at the same file or the same variable.

for a middle

Explain why a loser's write to shared state is a genuine data race rather than merely wasted work, and how private per-attempt state plus a result channel removes the sharing entirely.

for a senior

Show the containment as a package: one shared cancellable context cancelled on return, buffered results so cancelled losers can exit, private buffers or temp files, self-cleanup in the attempt, and per-attempt outcome logging to prove the losers really stop.

for a principal

Own the rule about when racing is legitimate at all: attempts must be redundant and idempotent, and the duplicate resource cost during the loser tail is a real budget line that has to be justified against the tail-latency it buys.

## The premise people get wrong A `return` in the calling function ends *that* function. It does not end the goroutines it started. In a mirror race the caller is gone as soon as one archive lands, and two downloads are still in flight — reading sockets, appending to buffers, and eventually running whatever code follows their fetch. Everything that code does happens **after** the operation it belonged to has completed. That is the whole hazard, and it splits into correctness and cost. ## Correctness: writes that land too late **A shared destination.** The obvious design is to have every attempt stream into the file the caller asked for. Now two or three writers append to the same file, interleaved, and the caller has already checksum-verified and returned. The artifact on disk is a blend of two mirrors' bytes and fails verification later, in a completely different part of the program, with no clue where it came from. The fix is structural: **each attempt writes to its own temporary file**, and the winner's file is renamed into place by the caller after verification. Losers' temp files are removed — and because a loser may finish after the caller returned, its own cleanup must be its own responsibility, not the caller's. **Caller-owned memory.** The same mistake in miniature: attempts appending into a shared `[]byte`, assigning to a field of a struct the caller owns, or writing into a `map` the caller reads. The caller reads the winner's data and returns; a loser writes the same memory a second later. There is no happens-before edge between those two, so it is a genuine data race — the kind that reports cleanly under a `-race` build precisely because it does happen at runtime rather than being a theoretical interleaving. Attempts must communicate only through the result channel; the value they send is theirs until the receiver takes it, and that send is the synchronisation. **Anything non-idempotent.** If an attempt increments a counter, writes a cache entry, marks a package as installed, or emits an event, the race performs that action N times, and N-1 of those are for results nobody used. Cache entries are the one people accept — a loser populating a mirror cache is often harmless or even useful — but it must be a decision, not an accident. ## Cost: what a loser holds while it finishes Until it returns, each losing goroutine holds a TCP connection out of a bounded pool, an in-flight HTTP response, and the bytes it has downloaded so far — which for an archive race is a full copy of the artifact per loser. Race three mirrors on every request in a service and steady-state memory and connection use are roughly tripled for as long as the losers take to finish. Cancelling them is not a nicety; it is what keeps the pattern's cost close to one download plus a tail. ## Containment, concretely 1. **Derive one cancellable context** from the caller's, give it to every attempt, and cancel it on the way out of the function. A loser blocked in an HTTP request returns promptly with a cancellation error instead of downloading the rest of an archive nobody wants. 2. **Keep the result channel buffered to the number of attempts** so a cancelled loser can still deliver its (now useless) result and exit rather than parking on a send. 3. **Private state per attempt** — its own buffer, its own temp file, its own error. Nothing shared and mutable. 4. **Self-cleanup in the attempt.** Since a loser outlives the caller, it must close its own body and remove its own temp file in its own defers. 5. **Do not race side-effecting work.** The pattern assumes attempts are redundant and idempotent. Racing three writes or three payments is a different program. ## The evidence you want Emit one structured log line per attempt as its result arrives — mirror, outcome, duration — plus which one won. That single artefact answers the questions that decide whether the race stays: how long the losers keep running after the winner returned (are they actually being cancelled?), whether the same candidate wins every time (in which case you are paying for duplicates to no benefit), and whether a candidate always reports back instantly with an error. Without those lines the losers are invisible: they consume real resources during a window that ends after the request they belonged to has been logged as complete, so nothing in the normal request trace shows them at all.

  • Why must a losing attempt clean up its own temporary file rather than the caller doing it?
    Because the caller returns before the loser finishes, so at return time the loser's file may not exist yet — and moments later it does, with nobody left to remove it. Ownership has to sit with the goroutine that outlives the call: create, close and remove in the attempt's own defers.
  • Is a loser writing into a slice the caller already read really a data race, or just wasted work?
    A real race. The caller's read and the loser's write are on different goroutines with no synchronisation between them, so the program has undefined behaviour and a `-race` build will report it. It is also the kind that actually fires in production, because the loser genuinely does run after the caller's read.
  • How would you tell whether the losers are being cancelled promptly?
    Log each attempt's outcome and duration as its result arrives, alongside the winner's. If losers keep reporting long after the winner's duration, your cancellation is not reaching them — typically a fetch that ignores the context, or a read with no deadline — and you are paying for full downloads you decided not to use.

saying these in an interview costs you the question

  • Assumes returning stops the other goroutines
  • Has every attempt write to the same destination file
  • Lets attempts assign into a caller-owned struct field
  • Cleans up losers' temp files in the caller
  • Races attempts that have non-idempotent side effects