skip to content

questions

4

A sign-up email never reaches the mailbox the run polls. How do you tell a product defect from a delivery-path failure?

level: seniorimportance: must knowfreq 52%

answer

  1. One symptom, several possible failing links
  2. Check the links in order, stop at the first gap
  3. The handoff identifier proves the product sent
  4. Delivery records separate deferred from refused
  5. Accepted does not mean the poller can see it

basics

~20 s

Walk the handoff chain and stop at the first missing link: did the product decide to send, did it hand off and get an identifier, what did the sending service report - and only last, the mailbox.

solid answer

~50 s

Turn one symptom into a chain of checkpoints and find the first one that is missing. The product's own record says whether a send was attempted at all; the identifier returned at handoff says the message left the product; the sending service's delivery records for that identifier say whether the receiving side accepted, deferred or refused it; the mailbox the run polls is only the last link. A product defect fails at one of the first two checkpoints - no attempt, or no handoff - while a delivery-path failure fails at the third or fourth with the earlier ones intact. The engineering work is making that chain readable from inside the case: capture the identifier at handoff, look up the delivery outcome for it, and attribute a delivery-path failure to the delivery path, so nobody spends a morning reading feature code because a sending allowance ran out.

code

pseudocode · 15 lines
pseudocode
trigger_signup(recipient)

handoffId = product.outbound_record(recipient).handoffId
if handoffId is null:
    fail("product never handed the message off")        # product defect

message = capture.poll_for(recipient, deadline)
if message is null:
    outcome = delivery_records.lookup(handoffId)
    if outcome in (DEFERRED, REFUSED, BLOCKED):
        report_delivery_path_failure(outcome)           # not a feature regression
    else:
        fail("accepted but absent: compare recipient at handoff with poller")

assert message.activation_link is present

go deeper

for a junior

Know that a message which does not arrive has several possible causes - the product may never have sent it, or the delivery path may have refused or delayed it - and that the mailbox alone cannot tell you which.

for a middle

Explain the chain of checkpoints and what each one proves: the product's send decision, the identifier returned at handoff, the delivery record for that identifier, and finally the mailbox the run polls.

for a senior

Show that you build the chain into the case - capture the handoff identifier, look up the delivery outcome for it, and attribute the failure to the delivery path so a feature team is not sent to read code because an allowance ran out.

for a principal

Own the attribution boundary: who is accountable when messages stop arriving, what evidence the organisation keeps in order to settle that quickly, and how much instrumentation is worth paying for to end a recurring argument.

## One symptom, four possible failing links "No message arrived" is the least informative failure a notification case can produce, because every stage of a long chain reports it identically. The repair is to stop treating it as one observation and start treating it as the **last** observation in an ordered sequence, each link of which can be checked independently: 1. **Did the product decide to send?** The application's own record - the audit row, the outbox entry, the log line written at the decision point - answers this without involving anything external. 2. **Did the product hand the message off?** A successful handoff to a sending service returns an identifier for that message. If the product has no identifier, the message never left it. 3. **What did the delivery path do with it?** The sending service records, per identifier, whether the receiving side accepted it, deferred it, refused it permanently, or produced a complaint afterwards. 4. **Did it reach the mailbox the run polls?** Only now is the capture mailbox a meaningful observation. Read them in order and stop at the first one that is missing. That link, not the symptom, is the finding. ## What the first missing link tells you | First missing link | Verdict | Who owns it | | --- | --- | --- | | No send decision recorded | Product defect - the flow never got that far | The feature team | | Decision recorded, no handoff identifier | Product defect - the send failed or was never attempted | The feature team | | Identifier exists, delivery record says refused or blocked | Delivery-path failure - reputation, recipient, or a receiving-side rule | Whoever owns sending | | Identifier exists, delivery record says deferred | Delivery-path failure - the message is queued, not lost | Whoever owns sending | | Delivery record says accepted, mailbox empty | Between the receiver and the mailbox - routing, folder, or a mismatched recipient | Usually the run's own setup | That last row is the one that catches people. Acceptance is a statement by the receiving system that it took custody, not that a person will find the message where the poller is looking. The usual causes are mundane: the message was filed somewhere the poller does not read, an alias or routing rule ahead of the mailbox diverted it, or - most often of all - the address the case is watching is not byte-identical to the one the product actually used. ## Making the chain readable from inside the case None of this triage is affordable if it means a human logging in to three systems. What makes it cheap is one piece of instrumentation: - **Expose the handoff identifier.** The product stores the identifier it received when it handed the message off, against the account under test, somewhere the run can read it. Everything downstream is then a lookup for exactly this message rather than a guess from timestamps in a shared log. - **Look the outcome up before failing.** When the poll deadline expires, the case queries the delivery record for that identifier instead of failing blind. The failure message then says which link broke. - **Record the recipient at both ends.** Log the address the product used alongside the address the poller watched, so the mismatch case is settled by reading the failure rather than by re-running. - **Keep a coarse category on the product's own send error.** Even a single word - refused, timed out, rejected - carried into the sign-up error saves the whole investigation when the cause was upstream. ## Attributing the failure A missing message is not automatically a red build against the feature. Once the chain identifies the broken link, the failure should be attributed to whoever can fix it: a message the product never handed off is the feature team's, and a message accepted and then deferred belongs to whoever owns sending. Reporting both the same way trains everyone to disbelieve the suite - the second time a team reads feature code for a morning because an allowance ran out, they stop reading it at all, including the time the defect is real. ## The two cases that look identical and are not Two situations produce the same empty mailbox and deserve opposite responses: - **The product silently swallowed a failed send.** The decision is recorded, no identifier exists, and nothing was raised to the user. This is a genuine and often serious defect: real customers are being told their account exists while no activation message was ever sent. - **The delivery path deferred the message.** The identifier exists, the record says deferred, and the message is still in flight. Nothing in the product is wrong; the case's expectation about when it would arrive was. Both look like "no email". Only the chain separates them, which is the whole argument for building it before the first time it is needed rather than during the outage.

  • The delivery record says accepted but the mailbox the run polls is still empty. Where do you look next?
    Between the receiving system and the poller. Compare the recipient recorded at handoff with the address the poller is watching - a case-built address that differs by a character is the commonest cause. Then check for a routing or alias rule ahead of the mailbox, and for filing into a folder the poller does not read. Only after those is the product worth reopening.
  • How should a delivery-path failure appear in the run's results?
    Attributed to the delivery path rather than to the feature: the run should name the checkpoint that failed and route it to whoever owns sending. Reporting an exhausted sending allowance the same way as a broken sign-up flow teaches a team to disbelieve the suite, which costs far more than the failed case did.
  • Which single piece of instrumentation makes this chain cheapest to walk?
    The identifier returned when the message is handed off, recorded by the product against the account under test and exposed where the run can read it. With it, every later record - accepted, deferred, refused - is a direct lookup for exactly this message rather than an inference from timestamps in a log shared with every other run.

Like a parcel that never turned up: the shop's order record, the courier's collection scan and the delivery attempt notice each rule out a different stretch of the journey, and only reading them in that order tells you who to call.

saying these in an interview costs you the question

  • Reports every missing message as a product regression
  • Checks only the capture mailbox before raising a defect
  • Assumes acceptance at handoff means the recipient received it
  • Cannot say whether the product even attempted the send
  • Re-runs the case and moves on when it happens to pass
open as a page

Why can an automated sign-up suite that emails invented addresses damage the sending domain's reputation?

level: juniorimportance: should knowfreq 32%

basics

~20 s

Invented recipient addresses bounce. Receiving mail systems count permanent bounces and complaints against the domain that sent them, so a suite running nightly builds a poor record - and real product mail starts landing in junk folders.

open as a page

How does a mail-sending rate cap show up in a suite whose parallel workers each trigger an email?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Not as a clean error at the point of send. Later cases fail while earlier ones passed, failures cluster in time rather than by feature, and refused or deferred sends surface as ordinary product errors or missing messages.

open as a page

A sign-up email lands in some recipients' junk folder. What can an automated check honestly assert about that?

level: middleimportance: nice to knowfreq 20%

basics

~20 s

Assert only what the team owns - link targets, a text alternative, a working unsubscribe path, the visible sender address. The classification verdict belongs to the receiving side and shifts without notice, so track it as a signal, not a gate.

open as a page