A case fails when an activation email misses its arrival deadline. How do you tell a slow delivery from a lost one?
answer
- Absence looks the same either way
- Look for evidence outside the assertion
- The handoff record answers a different question
- Keep watching after the case has failed
- One run cannot separate late from lost
basics
~20 sAt the deadline both look identical: nothing has arrived either way. Separate them with evidence outside the assertion — the record that a send was accepted, a correlation value carried in the email, and continued observation after the case fails.
solid answer
~50 sThe assertion cannot distinguish them, because absence at one instant is the same observation in both cases. Gather evidence around it instead. First, check the product's own record that a send was accepted for this recipient — if there is none, the fault is upstream of the channel and timing is irrelevant. Second, tag each case with a correlation value so any message that turns up later can still be attributed to the run that caused it. Third, keep observing after the case has gone red, with a sweep an order of magnitude longer than the deadline, and record whether the message eventually landed and when. That yields three distinct diagnoses — never sent, sent and late, sent and absent — and only the third supports the word *lost*. Then aggregate: one run is a sample, not a verdict.
code
pseudocode · 12 linescorrelation = "run17-case04"
trigger_signup(recipient, tag = correlation)
if not await_message(recipient, correlation, deadline):
accepted = product_handoff_record(recipient, correlation)
if not accepted:
fail("no send accepted - upstream of the channel")
else:
fail("accepted at " + accepted.at + "; not observed within " + deadline)
# the case is already red - keep the question open off the critical path
schedule_sweep(recipient, correlation, window = 10 * deadline)
# sweep finds it -> late sweep finds nothing -> absentgo deeper
Be ready to say that a missed deadline does not prove the message was lost — only that it had not arrived at the moment the case stopped looking. Absence and lateness read the same at that instant.
Explain the two evidence sources and what each is good for: the product's own record that a send was accepted, and a continued watch on the receiving side after the case has failed. Say what neither one can prove.
Show you have diagnosed this on a real estate: correlation values that survive the failure, an extended sweep with a much longer window, failure text that names the window it searched, and the discipline of aggregating before declaring loss.
Own where each finding is routed: who receives a late-arrival signal versus a never-arrived one, what evidence the pipeline captures by default so nobody has to reproduce a failure to investigate it, and how much post-failure observation is worth funding.
## Absence is one observation with two causes At the instant a case gives up, "slow" and "lost" produce exactly the same reading: the mailbox does not contain the message. No amount of care inside the assertion separates them, because the assertion is a single sample of a system that is still in motion. If a case reports "the activation email was never sent", it is stating a conclusion its evidence does not support, and that mislabelling is expensive — it routes an investigation at the product when the fault is a queue, or shrugs off a genuine loss as "the channel being slow again". Separating the two is therefore not an assertion problem. It is an **evidence** problem, and the evidence has to come from outside the moment of failure. ## Three evidence sources, and what each can prove | Evidence | What it proves | What it cannot tell you | |---|---|---| | The product's own record that it accepted a send for this recipient | The trigger produced a send; the product's part completed | Nothing about whether delivery finished | | A correlation value carried in the message and matched on the receiving side | Any message you later find is *this* run's, not a neighbour's | Nothing while no message exists | | Continued observation after the case has failed, over a much longer window | Whether the message eventually landed, and how late | Nothing about why it was late | Read together they give three distinguishable diagnoses instead of one, and the three go to different places: 1. **No accepted-send record.** The failure is upstream of the channel entirely. Deadlines, batching and queue depth are all irrelevant; the defect is somewhere between the user action and the handoff, and it belongs to the product team. 2. **Accepted, and the message turns up in the extended observation.** The delivery finished, just later than the case was willing to wait. This is a latency finding: either the allowance was drawn from too little data or the channel has genuinely slowed. 3. **Accepted, and nothing turns up in the extended observation.** Now, and only now, is "lost" a supportable claim, and the investigation moves to why the delivery path dropped it — a separate question with separate owners. ## Designing the case so the evidence survives - **Give every case a recipient or correlation value of its own.** A late message that cannot be attributed to a run is useless evidence, and a shared identifier means a neighbour's message can be mistaken for yours. - **Do not stop observing when the case fails.** Schedule an extended sweep — an order of magnitude longer than the deadline — that records arrival or its absence against the same correlation value. The case is already red; the sweep costs the pipeline nothing on the critical path. - **Carry both timestamps into the failure report.** "Not observed within 45 seconds; observed at 4 minutes 12 in the post-run sweep" is a different sentence from "not observed within 45 seconds; not observed within a 20-minute sweep", and the next person needs to read which one it was. - **Resist re-triggering the action to see if it works this time.** A second send makes attribution ambiguous and can collide with whatever the product does about repeat requests, destroying the evidence you were about to collect. - **Do not treat the accepted-send record as a pass.** It answers a different question. Asserting on it alone silently deletes the whole delivery path from coverage. ## One run is not a verdict A single miss is a sample from a tail you may not have characterised. The distinction between late and lost is a property of the channel, and channel properties only appear in aggregate. - A scatter of arrivals just past the deadline, run after run, says the deadline is too tight rather than that anything is being lost. - A stable proportion of messages never observed at all, across many runs, is a loss signal worth escalating even when most runs are green. - A step change — messages that used to land in seconds now landing in minutes, all of them — points at the delivery path having changed, and is visible in aggregated arrival records before it is visible as a red build. That is why the extended sweep is worth funding: it converts every failure into a data point instead of a shrug. Over a month it tells you which of the three diagnoses your estate actually suffers from, which is a question no individual failing run can answer. ## The shape of a good failure message The failure a team can act on names the observation and the observation window, distinguishes what was checked from what was concluded, and attaches the correlation value so a human can go and look. The failure a team learns to ignore says only that something did not happen. On a channel that is asynchronous by construction, precision about *what was not observed, and for how long* is the whole difference between a diagnosis and a guess.
- The product has no record of accepting a send for that recipient. What does that change?It moves the investigation upstream of the channel entirely. Deadlines, batching and queue depth stop being relevant, because nothing was ever handed to the delivery path. The defect sits between the user action and the handoff — a condition that was not met, an error swallowed, the wrong recipient resolved — and the case's timing configuration is a red herring.
- How many runs do you need before calling it loss rather than lateness?More than one. A single miss is a sample from a tail you may not have characterised. Aggregate the extended sweeps: a stable proportion of messages never observed at all, across many runs, is a loss signal, whereas a scatter of arrivals landing just past the deadline says only that the deadline is too tight for the channel as it currently behaves.
A parcel that has not arrived by lunchtime and a parcel that fell off the van look identical from the doorstep. The difference is only visible in the sender's receipt and in whether it turns up by Thursday.
saying these in an interview costs you the question
- Calls the message lost the moment the deadline passes
- Stops observing as soon as the case has gone red
- Re-triggers the action immediately, destroying any chance of attribution
- Treats the accepted-for-delivery record as proof of arrival
- Draws a loss conclusion from a single failed run
- Reports only that something did not happen, with no window or identifier