skip to content

Racing three mirrors for one archive, how do you return the first successful download rather than the first attempt to finish?

level: middleimportance: should knowfreq 42%

answer

  1. failing is faster than succeeding
  2. the broken candidate answers first
  3. keep receiving instead of returning
  4. count the attempts, do not range
  5. one joined error only when all fail

basics

~20 s

Receive in a loop instead of once. Return the first result whose error is nil; collect the failures and keep receiving. Only after all attempts have reported do you give up, returning the accumulated errors together.

solid answer

~40 s

A single receive returns whichever attempt finishes first, and a mirror that is down refuses the connection in milliseconds — so the broken candidate reliably wins the race and hands you its error while the healthy ones are still downloading. The fix is to keep receiving: loop exactly N times over the buffered result channel, return immediately on the first result with a nil error, and append the others to a slice of failures. If the loop completes without a success, return `errors.Join(errs...)` so the caller sees why every mirror failed rather than just the fastest excuse. The loop count matters — iterate over the number of attempts started, not until the channel closes, because nobody closes a channel with N independent senders on it.

code

go · 9 lines
go
var errs []error
for range mirrors { // one receive per attempt started
	r := <-ch
	if r.err == nil {
		return r.body, nil
	}
	errs = append(errs, fmt.Errorf("%s: %w", r.mirror, r.err))
}
return nil, errors.Join(errs...)

go deeper

for a junior

Know the trap: the quickest answer is often a failure, because refusing a connection is far faster than transferring a file. Returning whatever arrives first is not the same as returning what worked.

for a middle

Explain the corrected receiving loop precisely — a fixed number of receives, an early return on the first nil error, accumulated failures joined at the end — and why ranging over the channel would hang.

for a senior

Show the operational consequence: total failure now takes as long as the slowest attempt, so the call needs an outer deadline, and per-attempt outcome logging is what tells you a mirror is failing fast rather than being slow.

for a principal

Own the semantic choice explicitly. First success wins suits redundant attempts; first error wins suits jointly required ones, and picking the wrong one silently makes an availability feature into an availability risk.

## First to finish is not first to succeed The naive race takes one value: ```go r := <-ch return r.body, r.err ``` and it has a perverse property: **failure is fast**. A mirror whose host refuses the connection answers in a millisecond or two; a mirror that resolves to nothing answers as soon as DNS says so; a mirror returning 404 answers after one round trip. Meanwhile a mirror that is *working* has to transfer the whole archive. So the single receive systematically selects the broken candidate. You built a race for availability and got a race that surfaces the fastest error — the pattern is now strictly worse than picking one mirror, because it fails whenever *any* mirror is down instead of when *your* mirror is down. ## The corrected shape Keep the fan-out identical; change only the receiving side. ```go var errs []error for range mirrors { // exactly one value per attempt started r := <-ch if r.err == nil { return r.body, nil // first success wins } errs = append(errs, fmt.Errorf("%s: %w", r.mirror, r.err)) } return nil, errors.Join(errs...) ``` Four details carry the weight: **Loop over the attempts, not the channel.** `for range mirrors` runs exactly as many receives as there are senders. Writing `for r := range ch` instead would block forever after the last value, because nobody closes this channel — there are N independent senders and no single owner that could safely close it. (Closing from a sender would risk a send on a closed channel, which panics.) Counting the receives is the idiom for a channel with many senders and a known count. **Return on the first nil error.** That early return is what makes it a race at all: you leave as soon as one mirror delivers, without waiting for the slow or dead ones. Because the channel is buffered to N, the attempts you abandon can still send and exit. **Aggregate the failures, do not overwrite them.** `errors.Join` builds one error wrapping all of them, so the message names every mirror and its reason and a caller can still `errors.Is`/`errors.As` through the whole set. Keeping only the last error is the common shortcut and it throws away the very information that makes a total failure diagnosable — "all three mirrors failed" is a different incident depending on whether all three said 404 or all three timed out. **Attach the identity when you wrap.** `fmt.Errorf("%s: %w", r.mirror, r.err)` is what turns the joined error into something actionable; without it you get three indistinguishable connection-refused messages. ## What this costs On the success path, nothing: the loop still returns on the first good value. On the total-failure path you now wait for the slowest attempt rather than the fastest, because you cannot know that everything failed until everything has reported. That is usually right — an operation that can succeed should not give up while an attempt is still in flight — but it is why the whole call still needs an outer deadline from the caller's context: without one, an attempt that hangs turns "all mirrors failed" into "the call never returns". ## First success versus first error, as a choice Both semantics are legitimate and they answer different questions: - **First success wins** — the attempts are *redundant*. Any one of them is a complete answer, so a failure is only interesting once they have all failed. This is the download race, and it is the semantics this pattern is named for. - **First error wins** — the attempts are *jointly required*. You need all of them, so one failure dooms the operation and the sooner you learn about it the better; the right move is to stop the siblings immediately. The bug is picking the second by accident, which is exactly what the single receive does. Say out loud which one your call site needs before you write the receiving side. ## The observability that proves it works Log one line per attempt as its value arrives: the mirror, whether it succeeded, and how long it took. Two things fall out. First, you see whether the race earns its duplicate downloads — if one mirror wins 99% of the time, you are paying triple egress for nothing and should reorder or stagger the attempts. Second, you see the *losers'* latencies, which is what tells you a mirror is quietly failing fast rather than being slow; a candidate that always reports back in three milliseconds with an error is a candidate that would have won every naive race.

  • Why loop `for range mirrors` instead of `for r := range ch`?
    Ranging over a channel ends only when the channel is closed, and nobody can close this one: there are N independent senders, and closing from a sender risks another sender panicking on a send to a closed channel. With a known number of senders that each send once, counting the receives is the idiom.
  • What does the corrected version cost on the failure path?
    Latency. You cannot conclude that everything failed until every attempt has reported, so total failure now takes as long as the slowest mirror instead of the fastest. That is the right trade, but it means the whole call must sit under a deadline from the caller's context so a hung attempt cannot turn failure into a hang.
  • Why join the errors instead of returning the last one?
    Because the shape of a total failure is the diagnosis. Three 404s mean the artifact is missing; three timeouts mean the network is. `errors.Join` keeps all of them in one error, and `errors.Is`/`errors.As` still match through it, so callers lose nothing while operators gain the whole picture.

saying these in an interview costs you the question

  • Returns the first result regardless of its error
  • Assumes a fast answer means a good answer
  • Ranges over a channel that nobody closes
  • Keeps only the last error and discards the rest
  • Waits for all attempts even after one succeeded