Finance reports double-counted revenue from a pipeline the team believed could not repeat an effect — how do you decide which destinations to make repeat-proof and what the next notch costs?
answer
- an effect, not a delivery
- earned per destination, per hop
- internal coherence is a different claim
- rank by retractability, not by ease
- freshness follows the commit interval
basics
~20 sThe strong promise names an effect at one named destination, not a delivery across a pipeline, so it has to be decided hop by hop. Inventory each effect by whether it can be retracted and who already consumed it, then price the next notch in freshness, storage and coupling.
solid answer
~50 sStart by correcting the claim. What is buyable is an effect that looks the same at a named destination whether the work ran once or twice; nothing stops the second delivery, it is made not to matter. A pipeline is therefore only as repeat-proof as its weakest hop, and the team's belief almost certainly came from a property of the engine's own internal accounting, which says nothing about a destination already written to. Then inventory the effects: a table that can be rebuilt, an accumulating counter that cannot, a payment, a message to a third party, a downstream job that has already read. Rank by retractability and blast radius, not by how easy each is to fix. Price the next notch honestly — output visible only at commit boundaries, identifier columns and deduplication indexes, coupling to what the destination can do, more moving parts on restart. Finally, decide who is told when a destination stays duplicate-tolerant, because the repair has to reach readers, not just rows.
go deeper
Recall that the strong promise is about what the destination ends up holding, not about stopping a second delivery, and that it has to be arranged for each place the pipeline writes.
Explain why an engine's internally consistent recovery and a duplicate at the destination are compatible, and name the three writing-side mechanisms that close the gap.
Trace a real pipeline hop by hop, say which write would repeat after a restart, and quantify what the affected consumers would double-count over that span.
Own the allocation: which destinations earn the guarantee, what freshness, storage and coupling each one costs, who is told when a destination stays duplicate-tolerant, and how the exception list stays true as the pipeline grows.
## First, correct the claim the team is making The phrase in circulation is *exactly-once*, and the useful reading of it is narrow: **the destination ends up in the state it would have been in had the work run once, even though the work in fact ran twice** — nothing prevents the second delivery, the second delivery is simply made not to matter. Two consequences follow immediately, and both are usually what has gone wrong when a report like this arrives: - it names an **effect at a named destination**, not a property of a pipeline. Each hop earns it separately, and the chain is only as strong as the hop that earns it least; - it is not the same claim as **the engine's own accumulated state reflecting each record once after recovery**. That second claim is often true, is what most runtime documentation is describing, and is compatible with the destination holding a row twice — the rows left the process and were never part of the state that recovery restored. So the investigation is not "is the engine broken". It is "which hop writes in a way that a replay changes". ## Inventory the effects, not the services Walk the pipeline and write down every point where something becomes visible outside the job, and classify each: | Class | Example | What a replay does | What earning the promise costs | |---|---|---|---| | Rebuildable | A derived table replaced whole each run | Nothing, if the write replaces by key | Usually already free | | Accumulating | A running counter maintained by increment | Permanently wrong until someone recomputes | Rewrite as a recomputed value, or key plus replace | | Published once | An append-only landing area | A second copy every reader must handle | Reader-side deduplication, paid forever | | Unretractable | A payment, a message to a third party | Real-world duplicate | Caller-supplied identifier on the other side, or accept a risk | | Already consumed | A downstream job that read the duplicate | The damage propagated before you noticed | Notification and recomputation, not a write-side fix | The last row is the one that turns a technical fix into an incident. Correcting the table does not un-send the alert, un-fire the invoice, or un-snapshot the copy a downstream team took this morning. Any plan that stops at the row is incomplete. ## Price the next notch honestly Moving one destination from duplicate-tolerant to repeat-proof is not free, and the costs are predictable: 1. **Freshness.** If the mechanism is committing output and the recorded read position as one unit, output becomes visible only at commit boundaries. The destination's freshness is then set by the commit interval, not by the record rate — and shortening the interval to recover freshness raises per-commit overhead and, for a stateful job, the cost of taking recovery points more often. 2. **Storage and read cost.** Identifier columns, uniqueness constraints, deduplication views and the indexes underneath them. Reader-side collapsing is paid on every query by every consumer, for as long as the data exists. 3. **Coupling.** The design now depends on what the destination can do. Replace the destination and the guarantee has to be re-earned, which is a constraint on future architecture, not just on this job. 4. **Operational surface.** Staged output that must be published, intent records that must be reconciled, and a restart path with more states in it. Each is a thing that can be half-done at three in the morning. Runtimes also differ in how much they contribute: some offer destination integrations that arrange the coupling, some leave the whole construction to the author, and none can supply a key the record does not contain. Do not price this from the runtime you know best. ## Decide, write it down, and publish the exceptions The deliverable is not a mechanism, it is a short table: for each destination, which class it is in, what guarantee it carries, and who is told when it is duplicate-tolerant. Two organisational habits make it hold: - **Spend where the effect cannot be retracted or recomputed.** Money and messages first, accumulating counters second, rebuildable tables last — they usually cost nothing to fix and are the ones teams fix first because they are easy. - **State the tolerated case out loud.** A destination that may double-count during a replay is a perfectly reasonable design; a destination that may double-count and whose consumers were never told is the thing that produced this report. The answer to "why did finance see it twice" is almost never a broken recovery mechanism. It is that the promise was believed to be a property of the pipeline when it was only ever a property of a particular write, and nobody had written down which writes had it.
- The team points at a setting that turns the strong guarantee on. Why is that not the end of the conversation?Because such settings govern what the runtime can arrange for the hops it participates in — typically its own state and destinations it has an integration for. A hop written by hand, a third-party call, or a destination that cannot replace by key is outside it. Ask which specific writes the setting covers, and treat every write it does not name as duplicate-tolerant until shown otherwise.
- How do you handle the consumers who already read the duplicate?Identify them before repairing the data, because the repair changes numbers they may have copied. For each, decide whether they can recompute from the corrected source, need an explicit correction, or took a copy nothing can reach. Publish the affected span of input and the window of time; a consumer that knows which day is suspect can usually fix itself.
- Is a duplicate-tolerant destination ever the right long-term answer?Frequently. Approximate telemetry, rebuildable indexes and staging areas that a later step collapses all tolerate repeats, and paying for a stronger guarantee there buys nothing. What makes it right is that it is chosen, recorded and visible to consumers, rather than assumed because nobody looked.
saying these in an interview costs you the question
- Treats the strong guarantee as a property of the pipeline rather than each write
- Quotes the engine's internal recovery claim as an end-to-end promise
- Fixes the table and never contacts the consumers who already read it
- Prices the next notch without mentioning the freshness it costs
- Hardens the easy destinations first and leaves the unretractable effect
- Leaves a duplicate-tolerant destination undocumented for consumers