A unit's function writes a row into an external system - why does that make a duplicate attempt on it unsafe?
answer
- the runtime drops records, not effects
- two copies, two external writes
- discarded may not mean stopped
- emit records, let one writer apply them
- cancellation lands mid-effect
basics
~20 sBecause the runtime can discard the losing copy's records but not what that copy already did outside the job. Two copies means two rows in the external system, and dropping one copy's output does not remove them.
solid answer
~50 sA **duplicate attempt** - a second copy of a unit that is running far behind its peers, of which the runtime keeps whichever finishes first - is safe only because the loser's output is thrown away before anything downstream sees it. That guarantee covers records flowing through the runtime and nothing else. A write the function performs directly into another system bypasses that path entirely: the runtime has no handle on it and cannot take it back when it discards the copy. So the row is inserted twice, the message is sent twice, the charge is applied twice. Worse, engines differ on whether the losing copy is actively cancelled or merely ignored, and a copy that is only ignored keeps running and keeps writing. The remedy at this level is to keep outside-world effects out of the unit body, or to switch duplication off for that step where the runtime lets you choose.
go deeper
Remember the boundary: the runtime can throw away the records a losing copy produced, but it cannot undo anything that copy did to a system outside the job.
Explain why two copies of the same unit cause two external writes, and name the safe shape - emit records and let one write path apply them once.
Show that you would audit which steps touch outside systems before relying on duplicated work, and that you know a losing copy may keep running rather than being killed.
Set the policy: which classes of step may be duplicated at all, and whether the platform enforces that side-effecting work sits on a single controlled write path rather than inside arbitrary unit code.
## The guarantee the runtime actually makes When a runtime starts a second copy of a lagging unit and keeps whichever copy finishes first, the safety of that trick rests on one narrow promise: **exactly one copy's output records become part of the step's result, and the other's are dropped before anything downstream observes them.** Because both copies computed the same piece from the same input, dropping one is harmless - the kept copy's records are indistinguishable from the discarded ones. That promise covers exactly one channel: records that travel through the runtime's own output path. It says nothing whatever about anything else the unit's code did while it ran. ## What escapes the promise If the function inside the unit reaches out and changes another system, that change is outside the runtime's reach. It cannot be rolled back when the copy loses the race, because the runtime does not know it happened. Concretely: - **A row inserted into another database** is inserted twice, once by each copy, and discarding one copy's records removes neither row. - **A message published to another system** is published twice, and any consumer of it acts twice. - **A call to a paid or rate-limited service** is made twice, doubling the cost and the quota consumed. - **A file written directly to a shared path** is written by two processes at once - the classic corruption case, where two writers interleave into one destination or one truncates what the other has written. - **A counter incremented somewhere else** is incremented twice, and nothing in the job's own output reveals it. There is a subtler version that catches experienced people. The unit's function may read from an outside system and cache the result, or take out a lock, or claim an item from a work queue. Two copies then contend with each other: one copy holds what the other is waiting for, and a unit that was merely slow becomes a unit that is stuck. ## The losing copy may not actually stop Designs differ on what "discarded" means, and this is the part most people get wrong: - Some runtimes **cancel** the losing copy as soon as the winner completes, interrupting it where it stands. - Others simply **ignore its result** and let it run to its natural end. Under the second behaviour, a losing copy keeps executing its function and keeps issuing external writes for as long as it takes to finish, which on a genuinely sick machine can be a long time after the step has moved on. And even a cancelled copy is cancelled *somewhere* - in the middle of its work, having already applied some of its effects and not others. There is no clean point at which an interrupted copy has done nothing. ## What you do about it In rough order of how much you should prefer them: 1. **Take the effect out of the unit.** Have the unit emit records describing what should happen and let one designated write path apply them once, on the runtime's own output channel. Then duplication is safe again because the only channel in play is the one the runtime controls. 2. **Turn duplication off for that step**, where the runtime lets you make the choice per step rather than per job. You lose the remedy for machine-caused slowness on that one step and keep it elsewhere. 3. **Make the write land on the same result when repeated**, so that a second application changes nothing. This is a large subject in its own right - the general treatment of repeated delivery and of keys that make a write repeatable belongs elsewhere - but the local point stands: if repeating the write is genuinely harmless, the hazard disappears. What you should *not* do is accept the duplicate effect on the grounds that it is rare. The mechanism fires precisely when a machine is unhealthy, which is when you are least able to reason about what the other copy managed to do. ## Where this sits Keep two related questions separate from this one. What the outside world ends up seeing after a job crashes and its work is replayed is a recovery question, not a straggler question. And the general statement of how many times a system may deliver or apply something is distributed-systems material that applies far beyond processing engines. The point here is narrower and sharper: **the one remedy that costs the author nothing stops being free the moment a unit does something the runtime cannot take back.** An interviewer asking this wants to hear that you know which channel the runtime's guarantee covers, and that you can spot a unit whose body quietly leaves it.
- Is a unit that only emits records into the job's own output affected in the same way?No. Output that flows through the runtime's own path belongs to whichever copy wins, and the loser's records are dropped before anything downstream sees them. That is the case the mechanism was designed for. The hazard is specifically the effect that bypasses that path and lands somewhere the runtime has no control over.
- The losing copy is discarded - does that mean it has been stopped?Not necessarily. Some designs cancel it outright, others simply ignore its result and let it run to completion. Under the second behaviour it keeps issuing whatever external calls its function makes until it ends. And a copy that is cancelled is cancelled part-way through, having applied some effects and not others.
- Two copies of a unit both claim an item from an external work queue. What goes wrong?They contend rather than duplicate. One copy takes the item and the other blocks or fails waiting for it, so the remedy for slowness has produced a stall. Worse, if the winner is the copy that gets discarded, the item has been consumed by work whose output was thrown away.
saying these in an interview costs you the question
- Believes the runtime can undo the losing copy's external writes
- Assumes a discarded copy is always stopped immediately
- Thinks duplication is only started for units that write nothing
- Says the risk is acceptable because the effect is rare
- Treats duplicate copies as coordinating with each other over a lock