skip to content

How does a test prove an unprocessable queue message was parked rather than redelivered forever?

level: middleimportance: should knowfreq 44%

answer

  1. Both outcomes look identical at first
  2. Wait for evidence, not for time
  3. Read the attempt count twice
  4. A later good message must land
  5. Filter by this run's identifier

basics

~20 s

Parking leaves evidence that endless redelivery cannot fake: a record in the parking destination carrying this run's correlation identifier, a processing-attempt count that stops growing between two readings, and a later good message that gets processed.

solid answer

~50 s

The two outcomes look identical for the first few seconds, so the case has to wait for evidence rather than for time. Three observations together settle it. First, a record for the injected message appears in the parking destination, found by the correlation identifier the case generated, carrying a failure reason and an attempt count. Second, the processing-attempt count stops moving: read it, wait a settle window, read it again and assert the two readings are equal — a single reading cannot distinguish stopped from slow. Third, publish a known-good message behind the bad one and assert its effect lands, which a consumer still grinding on the poison message can never do. Bound the whole wait with a deadline derived from the declared attempt ceiling and its declared spacing, and report each observation as its own assertion.

code

pseudocode · 16 lines
pseudocode
baseline = parked_count()
bad_id   = "poison-" + random_id()
good_id  = "good-"   + random_id()

publish(channel: "orders.incoming", id: bad_id,  body: undecodable_bytes())
publish(channel: "orders.incoming", id: good_id, body: valid_order())

await_until(deadline: 45s) { parked_record(bad_id) != null }   # a decision was recorded
await_until(deadline: 45s) { order_exists(good_id) }           # the consumer resumed

n1 = attempts_for(bad_id)
settle(5s)
n2 = attempts_for(bad_id)

assert n1 == n2                          # redelivery has stopped, not merely slowed
assert parked_count() == baseline + 1

go deeper

for a junior

Be ready to say that a message a consumer cannot handle ends up in one of three places — delivered again, moved aside deliberately, or quietly lost — and that a passing test on its own does not tell you which.

for a middle

Explain the evidence rather than the timing: a parking record carrying the run's own identifier, an attempt count that stops moving between two readings, and a later good message that gets processed.

for a senior

Show how the case fails informatively — separate assertions for parked and for resumed, a deadline derived from the declared attempt ceiling, and a count baseline so an old parked record cannot pass the case on your behalf.

for a principal

Own the standard that every consumer ships with a case proving its unprocessable path, and that the team agrees in advance what counts as evidence, so a green suite stops being mistaken for an argument that failures are contained.

## Why elapsed time is not evidence Moments after a bad message is published, a consumer that will park it and a consumer that will grind on it forever look exactly alike from outside. Nothing has been produced, nothing final has been logged, and no error has reached your case. A case that sleeps a few seconds and then asserts "the test saw no exception" passes in both worlds — and passes just as happily in a third, where the message was silently discarded and no one will ever know. The job is to wait for **evidence**: a state only one of those outcomes can produce, reached inside a bounded deadline, with a specific failure message when it is not. ## Three signals, and what each one rules out | Signal the case checks | What it rules out | What it leaves open | | --- | --- | --- | | A parking record carrying this run's correlation identifier | endless redelivery, and silent discard | whether the consumer resumed other work | | The attempt count read twice, a settle window apart, and equal | an unbounded redelivery loop | whether the reader died rather than moved on | | A known-good message published behind it, and its effect observed | a wedged consumer making no progress | little — this is the strongest of the three | | The message no longer visible on the input channel | almost nothing | in-flight work, silent discard, an operator draining it | The last row is the trap, because it is the observation people reach for first. A message vanishing from the input side only means *something* took it: it is equally consistent with a consumer still chewing on it, a reader that dropped it without a word, and a human clearing the channel while your case ran. On its own it proves nothing. The first three are complementary rather than redundant. The parking record proves a decision was made and recorded. The stable attempt count proves the decision ended the redelivery loop instead of merely slowing it. The good message that follows proves the consumer is alive and moving, which is the difference between "parked" and "parked, then wedged" — a real and surprisingly common outcome when the failure is handled at a level that stops the reader instead of advancing past the message. ## Ordering the assertions so the failure reads well 1. Record the parked-count baseline, then publish the unprocessable message with a fresh correlation identifier. 2. Publish a known-good message behind it, tagged with its own identifier. 3. Wait until a parking record for the bad identifier exists, or the deadline expires. 4. Wait until the good message's effect exists, or the deadline expires. 5. Read the attempt count, wait a settle window, read it again, assert the readings match. Keep these as separate assertions with separate messages. "Parked, but the consumer never resumed" and "consumer resumed, but nothing was parked" are two different defects with two different owners, and a case that collapses them into one timeout tells whoever is on call nothing at all. ## Choosing the deadline from the design, not from superstition Derive the deadline from the consumer's own declared ceiling. Suppose the design says at most five attempts and the longest declared gap between attempts is eight seconds: parking cannot complete before roughly forty seconds, so a twenty-second deadline guarantees an intermittent failure and a five-minute one hides regressions in the spacing. Those numbers are an illustration, not a property of any system — read the ceiling and the spacing your own design declares and compute from them. ## False positives that make a green case worthless - **A leftover parked record** from an earlier run satisfies an assertion that searches for "a record" rather than "the record with this identifier". Always filter by the identifier the case generated, and baseline the count. - **A silently discarded message** looks like a park to any case that only checks the message stopped coming back. The parking record is what separates a designed park from data loss. - **A single attempt-count reading** cannot tell stopped from slow; two readings a settle window apart can. - **A consumer that parked and then died** passes the first two checks and fails the third — which is precisely why the third exists. - **A case that asserts on the newest parked record** starts failing the moment two runs overlap, and the failure will be blamed on the consumer rather than on the case. When all three signals are in place, the case says something worth saying: this input cannot be processed, the system recognised that, it recorded the decision where a human can find it, and it went back to work.

  • The parking record appears but the good message published behind it is never processed. What does that tell you?
    Parking and progress are separate facts, and only one of them held. The record was written, but the reader did not move past the bad message — typically the failure was handled at a level that stopped the reader, so everything queued behind it is stalled. The case should report that as its own defect rather than as a bare timeout.
  • Why is the message disappearing from the input channel weak proof that it was parked?
    Disappearance only means something took it. It is equally consistent with a consumer still working on it, with a reader that discarded it silently, and with someone draining the channel by hand while the case ran. Only the parking record plus resumed progress distinguish a designed park from quiet data loss.

saying these in an interview costs you the question

  • Sleeps a few seconds and asserts nothing crashed
  • Treats disappearance from the input channel as proof of parking
  • Reads the processing-attempt count once and calls it stopped
  • Searches for any parked record rather than this run's
  • Cannot tell a designed park from a silently discarded message