A price lookup is sent to three redundant sources and the first answer wins; what must happen to the two losers?
answer
- all candidates start, one is kept
- the commitment is permanent, not a fallback
- latency is the minimum, cost is the sum
- cancellation is cooperative, not immediate
- an instant failure is the fastest response
basics
~20 sAll three are subscribed at once; the moment one responds, the selection cancels the other two and delivers only the winner's values. Cancellation is a request the losing sources must honour, so work already in flight, and any side effect it has, may still complete.
solid answer
~50 sFirst-to-respond selection subscribes to every candidate immediately, waits for the first signal, then commits to that source and cancels the rest. From that point the losers' values are never delivered — they are not held as a fallback, and a winner that later fails does not hand the result back to them. Three consequences follow. **Cost multiplies**: every candidate performs its work, so latency is the minimum but load is the sum. **Cancellation is cooperative**: a source that does not check for it runs to completion, so its side effects still land — which makes this rule safe over repeatable reads and unsafe over writes. **The fastest response may be a failure**: a candidate that errors instantly is the quickest to respond, and rules differ on whether the winner is decided by the first signal of any kind or only by the first value.
go deeper
Recall the shape: every candidate is started, the first to respond is kept, and the rest are cancelled. Only the winner's values reach the output.
Explain the trade: latency is the fastest candidate's, load is all of them combined, and the choice is permanent — the cancelled candidates are not a fallback if the winner later fails.
Show the operational view. Cancellation is cooperative so a loser's effects can still land; a candidate that fails instantly may win; and two thirds of your dispatched work produces no visible outcome unless you record it per candidate.
Decide when paying several times the load for the best response time is justified, and set the precondition — interchangeable answers, repeatable reads, no per-call cost — so teams do not race calls that write or charge.
## What the rule actually does First-to-respond selection takes several candidate sources that are expected to produce the *same* answer and turns them into one output. It subscribes to all of them at the same moment, watches for the first signal, and then **commits**: the winner becomes the output, and every other subscription is cancelled. It is a selection rule, not a combining rule — the other candidates' values are not merged, buffered or reserved. The commitment is the part candidates misjudge in interviews. Once the winner is chosen, the losers are gone. If the winner produces one value and then hangs, or fails on its second value, the selection cannot fall back: there is nothing left to fall back to. Redundancy bought you a faster first answer, not a durable one. ## The three costs 1. **Work multiplies by the number of candidates.** Latency becomes the minimum of the candidates' latencies; load becomes the sum. With three sources you have traded three times the downstream load for the best of three response times, and every one of those calls occupies a connection and a worker for as long as it runs. 2. **Cancellation is a request, not a guarantee.** A cancelled subscription tells the source to stop; whether it stops depends on the source honouring it. Work already dispatched may run to completion, and anything it does that is visible outside the pipeline still happens. Selection is therefore safe over repeatable, side-effect-free reads and dangerous over anything that writes, charges, sends or reserves. 3. **A failure is a response.** The fastest candidate is often the one that fails immediately — a closed connection, a rejected request, a missing configuration. Formulations differ on whether the winner is decided by the **first signal of any kind**, including a failure, or only by the **first value**; the difference decides whether one instantly-broken candidate can poison a selection that had two healthy alternatives. Know which rule your pipeline uses before you rely on it. ## What "winning" can mean | the winner is decided by | consequence when one candidate fails instantly | |---|---| | the first signal of any kind | the failing candidate wins and the output fails, despite two healthy candidates | | the first value only | the failure is ignored, and the first candidate to actually produce data wins | | the first value, failures collected | the output fails only when every candidate has failed | ## Operating it - **The losers' outcomes disappear.** Two thirds of the work you dispatched produces no observable result at all, so error rates measured at the output understate what your downstream sources are doing. Record each candidate's outcome separately if you care whether one of them is quietly broken. - **A candidate that wins by failing fast makes the system look fast and broken at once.** Response-time metrics improve while the error rate climbs, which is a confusing pair of signals during an incident. - **Dispatching the second candidate after a short delay** keeps the usual case at one call while still capping the tail: nothing is sent to the alternatives unless the first has not answered within a chosen interval. The cost multiplier then applies only to slow requests rather than to all of them. - **Idempotence is the precondition.** Every candidate must be safe to invoke even when its result will be thrown away, because by design most of them will be. ## Where it fits and where it does not Use it where several sources are genuinely interchangeable and reading is cheap and repeatable: redundant read replicas of the same data, a local cache raced against the authoritative source, several mirrors of the same static content. Avoid it where the candidates are not equivalent (their answers differ, so which one wins changes the result), where a call costs real money per invocation, and where a candidate performs any effect that cannot be safely repeated or abandoned halfway. And separate it from the rules next to it. It is not lockstep pairing, which waits for every source; it is not latest-value combination, which keeps every source alive and re-emits on each arrival; it is not sequential appending, which runs candidates one after another and would give you a fallback but not a faster answer. Selection is the only one of the four that deliberately throws work away.
- The winning source fails on its second value. Can the selection fall back to a loser?No. The losers were cancelled at the moment of commitment, so there is nothing to return to. Recovery has to be built separately — for example by treating the whole selection as one unit and retrying it, which starts a fresh race among all the candidates rather than resuming a cancelled one.
- How do you keep the cost multiplier off the common case?Dispatch only the first candidate, and start the alternatives only if it has not answered within a chosen interval. Fast requests then cost one call each, and the multiplier applies only to the slow tail — which is where the redundancy was worth paying for in the first place.
- Why is idempotence a precondition for this rule?Because by construction most candidates' work is discarded, and cancellation is cooperative, so a loser's effect may still land even though its value is never delivered. Only work that is safe to perform and throw away — repeatable reads with no external effect — can be raced safely.
saying these in an interview costs you the question
- Says the losing sources are kept as a fallback if the winner fails
- Assumes cancelling a source stops its work immediately
- Thinks only one candidate actually does any work
- Believes a candidate that fails instantly cannot win the selection
- Treats selection as safe over calls that write or charge
- Claims the losers' values are merged into the output