A queue consumer parks messages it cannot process without alerting anyone. What must an automated case assert so the park is not read as success?
answer
- Parking is contained failure, not success
- Exact ceiling, not eventually stopped
- The record must support a later replay
- Assert the signal an alert reads
- Baseline the parked count first
basics
~20 sAssert three facts: redelivery stopped at the declared attempt ceiling, the parked record carries a reason, an attempt count and the correlation identifier, and the park moved an operator-visible signal. An unannounced park is contained failure, not success.
solid answer
~50 sA park is work the system gave up on, so a case that checks only that the message reached the parking destination certifies silence. Assert the ceiling first: compare the number of processing attempts against the bound the design declares — if it says five, assert five, because an unbounded loop also "eventually stops" when your case times out. Assert the record's content next: the failure reason, the failing stage, the attempt count, the original payload and the correlation identifier, since that is exactly what a replay after the fix will need. Then assert the announcement: the counter or event an alerting rule reads must move, and where the case can reach the rule itself, that it fires. Finally, baseline the parked count before injecting, so an old record cannot fake or mask the result.
code
pseudocode · 14 linesdeclared_ceiling = 5 # from the consumer's own design note
signal = "consumer.parked.total" # the counter the alerting rule reads
before = counter_value(signal)
pid = publish_unprocessable(channel: "orders.incoming")
await_until(deadline: 60s) { parked_record(pid) != null }
rec = parked_record(pid)
assert rec.attempts == declared_ceiling # bounded, and bounded where we said
assert rec.reason is not empty
assert rec.failing_stage is not empty
assert rec.original_body == injected_body # enough to replay after a fix
assert counter_value(signal) == before + 1 # somebody will be toldgo deeper
Know that a parked message is work the system gave up on, not work it completed, and that a test which sees the park and stops there has proved nothing about whether anyone finds out.
Explain the three separate assertions — attempts equal to the declared ceiling, a parked record carrying reason and attempt count, and a signal moving — and why each catches a different regression.
Demonstrate the judgement: derive the ceiling from the design rather than from observation, assert on the exact signal the alerting rule reads, baseline the count, and clean up so the alert protecting real work stays trustworthy.
Own the tradeoff between a strict alert that fires on any abandoned work and a threshold that tolerates a trickle. Decide who is accountable for a parked message, how quickly it must be triaged, and how automated environments avoid eroding that rule.
## A park is contained failure, not success When a consumer parks a message it cannot process, a piece of work has been abandoned. That is the right engineering outcome — far better than a reader wedged forever on one bad message — but it is not a success, and the danger of a well-built parking path is precisely that it makes failure comfortable. Nothing crashes, throughput looks normal, dashboards are green, and the only trace is a record in a destination nobody opens. A case that asserts "the message was parked" and stops there is not testing the failure path; it is ratifying the silence. The case therefore has to assert three separate things, and each one catches a different regression. ## The three assertions 1. **Redelivery stopped at the declared ceiling.** Read the number of processing attempts for the injected message and compare it to the bound the design states, not to "more than one" and not to "it eventually stopped". Eventually-stopped is what an unbounded loop looks like when your deadline expires first. Asserting the exact number is also what catches the two most common drifts: a ceiling raised to a large value during an incident and never lowered, and a ceiling that quietly became one, so a transient failure is now parked on its first attempt and real work is being thrown away. 2. **The parked record carries the diagnosis and everything a replay needs.** A record consisting of a payload and nothing else means the person who finds it three days later has to re-derive what went wrong. The record should carry the original body and metadata, the correlation identifier, the failure reason, the stage that rejected it, the attempt count and the time of the last attempt. Assert on those fields by name; they are the contract between the failure path and whoever has to recover from it. 3. **The park announced itself.** Something an operator can see must change: a counter incremented, a structured log record emitted at a level that is actually collected, an event raised on the channel humans watch. Assert the change, not the intention. | What the case asserts | The regression it catches | | --- | --- | | Attempts equal the declared ceiling | a ceiling raised during an incident and left there | | The record carries reason, stage and attempt count | a park that cannot be diagnosed or replayed later | | The signal an alerting rule reads moves | a park nobody is ever told about | | The parked count returns to baseline after clean-up | a destination filling with deliberate failures until the alert is ignored | ## Making the announcement assertable Teams stumble here because the alerting rule often lives outside anything an automated case can reach. Work down this ladder and take the highest rung available: - **Best:** drive the rule itself — inject enough parked messages to cross its threshold in an environment where notifications are routed to a sink the case controls, and assert the notification arrived. - **Good:** assert on the exact signal the rule reads. If the rule is written against a counter, assert that counter moved by one; the rule and the case then agree on a single name, and renaming the counter breaks both together instead of silently disarming the alert. - **Adequate:** assert the structured log record exists with the parked identifier and reason, and cover the rule's existence separately as a review or configuration check. - **Not enough:** assert that the code contains a call to something alert-shaped. That passes when the signal is emitted at a level nobody collects. A useful sharpening question during design: *if this park happened at three in the morning, what would wake somebody, and how long would it take them to find the payload?* Whatever answers that question is what the case should assert on. ## Baselines, thresholds and clean-up Inject against a known baseline. Read the parked count before publishing, expect it to rise by exactly one, and remove or mark your record afterwards. Two habits follow from that. First, a case that leaves its deliberate failures behind trains the team to ignore the parked-count alert, which is the same alert this whole question exists to protect. Second, a threshold alert that fires on "any parked message" will fire on your suite; that is a signal to route the automated environment's alerts separately, not a reason to weaken the rule in production. Done properly, the case makes a promise on behalf of the failure path: bad input stops being retried after a known number of attempts, the abandonment is recorded with enough detail to recover from, and a human finds out that it happened.
- The alerting rule lives outside anything your case can reach. What do you assert instead?Assert on the exact signal the rule reads — usually a named counter or a structured log field — so the case and the rule share one name and a rename breaks both together rather than silently disarming the alert. Then cover the rule's own existence and threshold as a separate configuration check, and say plainly in the case name that it verifies the signal, not the notification.
- Why assert the exact attempt ceiling rather than just that redelivery stopped?Because an unbounded loop also stops when your deadline expires, so "it stopped" passes in the failure case you are trying to catch. The exact number also catches drift in both directions: a ceiling raised during an incident and never lowered, and a ceiling reduced to one, which parks transient failures that would have succeeded on a second attempt.
- What is wrong with a suite that leaves its deliberately parked messages behind?It fills the parking destination with failures nobody needs to act on, so the count-based alert protecting real work becomes noise and the team learns to dismiss it. Each case should remove or mark the record it created, and the automated environment's alerts should be routed separately so the rule can stay strict where it matters.
saying these in an interview costs you the question
- Asserts the message was parked and checks nothing else
- Accepts eventually stopped instead of the declared ceiling
- Parks a bare payload with no reason or attempt count
- Assumes somebody watches the parking destination
- Leaves deliberately parked records behind after every run