skip to content

How do you make an automated case deliver two identical payment messages at the very same moment, and what does that catch?

level: seniorimportance: nice to knowfreq 24%

answer

  1. Sequential replay is the easy half
  2. Two copies in flight together
  3. Release them from one shared point
  4. Repeat it; a race is probabilistic
  5. The gap between deciding and recording

basics

~10 s

Hold identical copies behind one release point so consumer instances process them in parallel. Concurrency exposes the gap between deciding a copy is new and recording its effect, which a sequential replay never opens.

solid answer

~50 s

First check the window is reachable: if the deployed shape hands every copy of one identity to a single place that processes them one at a time, they cannot overlap and the case would be theatre. Where they can overlap, stand up the parallel consumption production actually runs — several consumer instances, or one draining its backlog with several handler threads. Prepare several publishing threads, block them all on one shared release point, then release them together so the copies enter the transport within microseconds. Use more than two copies; two collide rarely, eight collide often. Wait until every copy is consumed, then assert exactly one effect as a delta from a captured baseline — never which copy won or in what order. The defect it finds is the gap between deciding a copy is new and recording that it was handled.

code

pseudocode · 14 lines
pseudocode
copies  = 8
barrier = newBarrier(parties = copies)
msg     = { messageId: "m-3a17", type: "PaymentCaptured", paymentId: "p-88", amount: 250 }

before = ledgerEntryCount(paymentId = "p-88")

repeat copies times in parallel:
    barrier.await()                          # all publishers block here
    publish(queue = "payments", message = msg)   # identical body, identical identity

awaitCondition(() => consumedCount(messageId = "m-3a17") == copies)

assert ledgerEntryCount(paymentId = "p-88") == before + 1
# repeat the whole block 20 times; report the iteration that failed

go deeper

for a junior

Be ready to recall that duplicates can arrive together, not only one after the other, and that a case sending two copies back to back is usually still testing them one at a time.

for a middle

Explain the mechanics: parallel consumption, a shared release point rather than a publish loop, more than two copies, and an assertion on the effect count rather than on the interleaving.

for a senior

Show the judgment: establish whether the window is reachable in the deployed shape before writing anything, treat an intermittent red as a real defect rather than an unstable case, and report the failing iteration.

for a principal

Own the call about whether this class of case belongs in the regular suite at all, on a schedule, or nowhere — and be able to argue that from the shape of the deployment and the cost of a duplicated effect.

## The half that a sequential replay proves A sequential replay — publish, wait for the effect, publish the identical copy, wait again, assert one effect — proves the consumer *can* recognise a message it has already handled. More precisely it proves that once the first delivery has been completely applied and the fact of that application is durably recorded, a later arrival of the same identity is a no-op. That is the ordinary case and it deserves a case of its own. What it structurally cannot reach is the interval **while the first copy is still being handled**. Nearly every duplicate guard has the same shape: decide whether this identity has been seen, do the work, record that it has been seen. Between the decision and the record there is a window, and a second copy that arrives inside that window is told "not seen yet" and does the work again. A sequential replay steps over the window by construction, because it waits for the first delivery to finish before it sends the second. | | Sequential replay | Concurrent replay | | --- | --- | --- | | What it asserts | One effect after two deliveries | One effect after many simultaneous deliveries | | Defect it finds | A guard that is missing altogether | A guard that has a window in it | | Result character | Deterministic | Probabilistic; needs repetition | | Cost to write | Low | Moderate — needs parallel consumption and a release point | ## Building the simultaneous case 1. **Check the window is reachable at all.** If the deployed shape routes every copy of one identity to a single place that handles them one at a time, two copies can never be in flight together and the case is theatre. The duplicates that *can* overlap are the ones produced by a retried publish, by a replay of a stored range, or by several instances draining the same backlog. 2. **Stand up real parallel consumption.** More than one consumer instance, or one instance that drains its backlog with several handler threads — whatever production runs is what the case should run. Concurrency invented inside the case process is not the concurrency you ship. 3. **Release the copies together.** A publish loop is not simultaneity; the first copy is often fully handled before the last is sent. Prepare several publishing threads, block them all on one shared release point, then release, so the copies enter the transport within microseconds of each other. 4. **Raise the copy count above two.** Two copies collide rarely. Eight or sixteen collide far more often, cost almost nothing more, and the invariant you assert does not change. 5. **Settle, then assert the invariant.** Wait until every copy has been consumed, then assert exactly one effect as a delta from the captured baseline. Never assert which copy won, in what order they were handled, or how long handling took — those vary legitimately. ## Reading a probabilistic result - **A single red is a red.** A defect that reproduces one run in twenty is a defect that reproduces one run in twenty. "Flaky" properly describes a case whose own verdict varies while the behaviour under test is correct; this case's verdict varies because the behaviour is wrong. - **Put the iteration in the failure output.** "Failed on iteration 7 of 20, two ledger rows for one identity" tells the next reader how hard the window is to hit and what was actually observed. - **Turn the dial instead of quarantining.** If it reproduces too rarely to investigate, raise the copy count, add background load, or run more iterations — do not park the case. - **A pass is weak evidence, not proof.** Green means the window was not hit at this concurrency on this run. Keep the case on a schedule rather than treating one green as a closed question. - **Keep it separate from the sequential case** so a red result names which of the two properties broke. ## When not to write it - Every effect of the message is naturally absorbing, so there is nothing a second application could double. - The deployed shape genuinely serialises handling of one identity, making the window unreachable. Record that reasoning where the case would have lived, and revisit it when the deployment shape changes — more instances, keyless publishing, or an internal handler pool all reopen it. - The effect is guarded by a constraint the storage layer enforces on the write itself. The concurrent case is still cheap insurance, but it is no longer the highest-value case you could write this week.

  • This case fails on one run in twenty while the product is unchanged. Is that a flaky case?
    No. A race that reproduces one run in twenty is a defect that reproduces one run in twenty, and the case is doing its job. Flaky properly means the verdict varies while the behaviour is correct. Report the failing iteration and what was observed, raise the copy count or add load to reproduce it faster, and resist parking it in quarantine.
  • Your consumers handle every copy of one identity serially, so the copies can never overlap. What do you do with the case?
    Drop it and record why, in the place the case would have lived. Forcing concurrency by calling the handler on two threads inside the case process tests a path delivery never produces, so a red there means nothing and a green means less. Revisit the decision when the deployment shape changes: more instances, keyless publishing or an internal handler pool all reopen the window.
  • Why assert only the effect count rather than which of the copies was the one applied?
    Which copy wins is legitimately non-deterministic and varies run to run, so asserting on it produces failures that describe nothing wrong. The invariant that actually matters is that one effect exists after many deliveries. Assert the delta from the baseline and leave the interleaving unconstrained.

Publishing copies in a loop is like sending runners off one at a time and calling it a race; a shared release point is the starting gun that puts them on the track together.

saying these in an interview costs you the question

  • Believes a sequential replay already covers simultaneous arrival
  • Calls the handler twice in-process and calls that concurrent delivery
  • Runs the race case once and treats the pass as proof
  • Quarantines the case as flaky when it exposes a real race
  • Asserts which copy won rather than the single surviving effect
  • Publishes copies in a tight loop and assumes they overlap