After a fix, a parked queue message is resubmitted for processing. What must the case check beyond the effect appearing once?
answer
- The failed attempts may have done real work
- The world moved on while it sat parked
- Count totals across the whole episode
- Retire the parked copy after replay
- Prove the case fails on unfixed code
basics
~20 sCheck what the earlier failed attempts left behind, and whether the resubmitted message is now stale: partial work must not double the totals, newer state must not be overwritten, and the parked copy must not stay replayable.
solid answer
~50 sA replay after a fix is a recovery scenario, not merely a second delivery. Each failed attempt may have completed steps before it failed — a row written, an outbound call made, a notification sent — so count totals across the whole episode rather than only after the replay. Ask whether the message is still valid: while it sat parked, later messages for the same entity may already have been applied, and the old one must either be rejected as stale or leave the newest state intact, never resurrect a superseded value. Assert the parked copy is consumed or marked, so a scheduled drain cannot apply it a second time. And prove the case earns its place by checking it fails against the unfixed build; a replay case written after the fix usually passes for the wrong reason.
code
pseudocode · 15 lines# 1. drive the real failure, so partial work genuinely exists
pid = publish_unprocessable(entity: "acct-42")
await_until(deadline: 60s) { parked_record(pid) != null }
calls_before = downstream_calls(entity: "acct-42")
# 2. the world moves on while the message sits parked
apply_valid_update(entity: "acct-42", version: 7)
# 3. fix deployed; replay by the path an operator would use
replay(parked_record(pid))
await_until(deadline: 30s) { parked_record(pid) == null } # copy retired
assert ledger_rows(entity: "acct-42", cause: pid) == 1 # counted for the episode
assert downstream_calls(entity: "acct-42") == calls_before # no extra outbound call
assert current_version(entity: "acct-42") == 7 # newer state not clobberedgo deeper
Know that a parked message is not gone: someone can put it back through processing after the bug is fixed, and the outcome has to be right despite the attempts that already failed.
Explain the two ways a replay differs from a first delivery — earlier attempts may have completed part of the work, and newer messages for the same entity may already have been applied — and what each means for the assertions.
Show the honest setup: drive the real failure, capture totals first, move the world on in between, replay by the operator's path, retire the parked copy, and prove the case fails against the unfixed build.
Own the recovery policy itself. Decide when a parked message is replayed, when it is corrected at the source, and when state is rebuilt instead — and make sure whichever is chosen is the one the automated case exercises.
## A replay is a recovery, not a redelivery When a message is parked, the system stops. When someone replays it after a fix, the system restarts from a point it has already partly visited, into a world that has moved on. Those two differences — history behind the message, and time in front of it — are what make the replay its own case rather than another delivery of the same input. | | A fresh delivery | A replay after a park | | --- | --- | --- | | Prior work on this message | none | one or more attempts that failed partway | | State of the entity | as the message expects | possibly updated by later messages | | Code that processes it | unchanged | changed, by the fix under test | | The copy in the parking destination | does not exist | must not stay replayable afterwards | A case that asserts only "after the replay, the record exists once" checks none of the four rows. It can pass while the earlier attempts left a duplicate outbound call behind, while the replay quietly reverted a newer correction, and while the parked copy sits waiting for the next operator to replay it again. ## Four things the case must check 1. **Partial work from the failed attempts.** A handler that writes a row, calls a downstream service and then fails on the final step has already done two of three things — possibly several times over. Count outcomes across the entire episode: one ledger row *in total* for this message, and the same number of outbound calls before and after the replay if the earlier attempts already made them. Asserting "one row after the replay" is the classic miss, because it is measured from the wrong starting point. 2. **Staleness.** Suppose the parked message carries a correction for an account, and while it sat parked two later messages for that account were processed. Replaying the old one must not roll the account back. The design has to say which behaviour is correct — reject as superseded, apply and let the newer value win, or merge — and the case asserts the chosen one by driving newer traffic in between and checking the final state is still the newest. 3. **The parked copy is retired.** After a successful replay the record must be removed or clearly marked as replayed, with the time and who did it. Without that, the same recovery gets performed twice: a second operator, or a scheduled drain of the parking destination, finds a record that looks untouched. 4. **The case fails on the unfixed build.** A replay case written after the fix landed will pass for the wrong reason — the message is now processable, so the whole recovery path is skipped. Run it once against the code with the fix reverted, or shape the assertion so it cannot pass without the fix, and record that you did. ## Getting the setup honest The temptation is to hand-place a record in the parking destination and replay that. It is quick, and it destroys the case. A hand-placed record has no failed attempts behind it, no partial side effects, no real attempt count and no genuine metadata, so the two hardest checks — leftover work and retirement of the real record — become untestable. - **Drive the actual failure.** Publish the unprocessable message, wait until it parks, and take the parked record the system produced. - **Capture the totals before replaying**: rows for the entity, outbound calls, notifications sent. They are the baseline the post-replay assertions compare against. - **Move the world on in between.** Apply at least one newer, valid message for the same entity so staleness is actually exercised rather than assumed away. - **Replay by the path an operator would use**, not by republishing a freshly built message. Republishing tests a normal delivery and quietly skips the recovery mechanism you meant to cover. - **Assert the correlation identifier survives** the round trip, so the recovery is traceable end to end in the record that a post-mortem will read. ## When the right answer is "do not replay it" Sometimes recovery is not replay. If the message was wrong at the source, the correct action may be to publish a corrected message and discard the parked one; if downstream state can be re-derived, the correct action may be a rebuild rather than a per-message retry. Both are legitimate, and both change what the case asserts — that the parked record is discarded with a reason, or that the rebuild produced the same totals. What is never acceptable is an undocumented habit of replaying everything in the parking destination after every deployment, because that is exactly the operation nobody has tested and the one that turns a contained failure into duplicated work.
- Why not simply place a record in the parking destination yourself and replay that?Because a hand-placed record has no history: no failed attempts, no partial side effects, no real attempt count and no genuine metadata. The two checks that matter most — leftover work from the earlier attempts, and retirement of the record the system itself created — cannot be made at all. Drive the real failure and replay what the system parked.
- How do you decide what correct behaviour is when a replayed message is older than the entity's current state?It is a design decision the case only records. Ask whether the message carries a version or timestamp the handler compares against current state, and whether the domain wants last-writer-wins, reject-if-superseded, or a merge. Write the chosen rule down, then drive newer traffic during the parked window and assert the final state matches that rule rather than whatever the code happens to do.
- What proves the replay case is testing the fix rather than passing for free?Running it once against the code with the fix reverted and watching it fail. Otherwise the message is simply processable now, the recovery path is never exercised, and the case degrades into a plain delivery test that would keep passing if the parking behaviour were removed entirely.
saying these in an interview costs you the question
- Counts effects only after the replay, not across the episode
- Hand-places a record in the parking destination and replays it
- Ignores newer updates applied while the message sat parked
- Leaves the parked copy in place after a successful replay
- Republishes a fresh message instead of using the recovery path