skip to content

After a push connection drops and reopens mid-case, which assertions are still valid?

level: seniorimportance: nice to knowfreq 22%

answer

  1. Part of the evidence is simply missing
  2. Some claims need a complete record
  3. Absence claims are the first casualty
  4. Mark the interruption inside the record

basics

~20 s

Only assertions that survive missing evidence. A gap leaves the case holding an incomplete record, so existence checks on items received after the reopen can stand, while counts, ordering and any claim that nothing was sent are void.

solid answer

~50 s

First the case has to know a gap happened at all - record disconnect and reconnect as marked entries in the same buffer as the deliveries, so the record reads as a timeline with a labelled hole. After that the rule is mechanical: any assertion that depends on having seen **every** item is void. Counts, ordering, first-item and nothing-was-sent all fall. What survives is a positive existence check matching an item received after the reopen, and any verification done through a different route once the flow has settled. The honest response to a gap is to end the case inconclusive and loud rather than green. Either report it with a distinct environmental result, or - when the triggering action is safely repeatable - re-arm and re-trigger with a fresh identifier. Never assume the transport quietly replayed what you missed.

code

pseudocode · 11 lines
pseudocode
# lifecycle markers live in the same record as the deliveries
channel.on_open(     -> observed.append({ at: now(), kind: "connected" }))
channel.on_close(r   -> observed.append({ at: now(), kind: "gap_start", reason: r }))
channel.on_reopen(   -> observed.append({ at: now(), kind: "gap_end" }))
channel.on_message(m -> observed.append({ at: now(), kind: "delivery", body: m }))

had_gap = any(e in observed where e.kind == "gap_start")

if had_gap and assertion_needs_complete_record:
    end_case(result = "inconclusive",
             note  = "evidence gap on the push connection")

go deeper

for a junior

Know that a push connection can drop and come back inside a single case, and that anything sent while it was down is simply absent from what the case observed.

for a middle

Explain which checks depend on having seen everything - counts, ordering, first-item and nothing-was-sent - and why those are precisely the ones a gap destroys.

for a senior

Show how the case detects the gap in the first place, and what it does next: an inconclusive result with the gap named, or a re-trigger with a fresh identifier when the action is safely repeatable.

for a principal

Own the policy for the whole suite: which checks are written gap-tolerant by default, how an evidence gap is reported so it never reads as a product failure, and what gap rate means the environment rather than the cases needs fixing.

## First, notice the gap A case cannot reason about missing evidence it does not know is missing. So the first requirement is not an assertion policy at all - it is instrumentation. Record connection lifecycle transitions as entries in the **same** buffer as the deliveries, in arrival order: connected, disconnected with whatever reason was given, reconnected. The buffer then reads as a timeline with a labelled hole in it, and every later decision becomes checkable rather than assumed. A case that reconnects silently and carries on asserting is asserting over a set it cannot describe. It will still go green most of the time, which is the problem: the gap only bites when the thing it swallowed was the thing under test. ## What a gap destroys The rule is mechanical. **Any assertion that depends on having seen every item is void.** | Assertion | Survives a gap? | Why | |---|---|---| | An item carrying my identifier arrived after the reopen | Yes | It was genuinely observed | | That item's amount field was as expected | Yes | The gap removed evidence; it did not corrupt what arrived | | Exactly four items arrived during the case | No | The count is of what was seen, not of what was sent | | The confirmation preceded the settlement | No | Either could have fallen in the hole | | Nothing of the cancelled kind was pushed | No | The gap removed exactly the evidence that would refute it | | The settled values read back through another route are correct | Yes | Does not depend on the connection at all | The absence row deserves its own emphasis, because it fails in the worst possible direction. A count or an ordering claim at least tends to fail loudly when items go missing. An absence claim gets *stronger* as evidence disappears: the gap removes precisely the item that would have refuted it, so a dropped connection converts a real defect into a confident green. ## Three honest responses 1. **End the case as inconclusive and loud.** The product may be fine; the case simply could not observe. Where the result store carries a status distinct from failed, use it, and put the gap in the message. Where it carries only pass and fail, fail - but name the gap in the first line of the failure so triage takes seconds. 2. **Re-arm and re-trigger**, when the triggering action is safely repeatable. Reopen, confirm the subscription is registered again, and perform a fresh action with a fresh correlation identifier. Never reuse the old identifier and never re-assert over a buffer that spans the hole. 3. **Design the assertion gap-tolerant from the start.** A positive existence check on an identifier the case creates *after* the reconnect is unaffected by anything the case missed. Where the flow allows it, that is the cheapest option of the three. ## Do not lean on replay Some transports can resume delivery from the last item a receiver saw, and it is tempting to treat a reconnect as therefore lossless. That is a property of the transport and of how the deployment configured it, not a property a case may assume. If a case does rely on it, it has to prove it rather than trust it: check that items either side of the gap are contiguous by whatever sequence value the payload itself carries, and treat any hole in that sequence as a gap after all. A resume you did not verify is indistinguishable from a resume that did not happen. ## Gaps are a signal about the environment An occasional drop inside a long case is normal; a suite where a noticeable share of cases hit a gap is telling you something about the deployment, the network path or the resource limits on the connection count, and no amount of assertion cleverness fixes it. That makes the gap worth counting as its own metric, separately from pass and fail. Two questions follow from the number: - Is the rate rising with parallel width? Then connections are competing for a limited resource, and the fix is capacity or fewer concurrent connections, not the cases. - Is the rate concentrated in long cases? Then something is enforcing a lifetime on the connection, and cases should be shortened or the flow verified through a route that does not need a connection held open. ## What this leaves the case author The practical discipline is short. Instrument the lifecycle so a gap is visible. Classify every assertion in the case as complete-set or existence. If any complete-set assertion is present and a gap occurred, the case cannot report a product verdict - so make it say so. And prefer, wherever the flow allows, an existence check on data the case itself created, because that is the only assertion shape here that a dropped connection cannot quietly weaken.

  • Why is "no cancellation update was pushed" the first assertion to fall after a gap?
    It is a claim about the whole set, and it fails in the worst direction. A count or an ordering claim tends to fail loudly when items go missing; an absence claim quietly gets stronger, because the gap removes exactly the evidence that would refute it. A dropped connection therefore turns a real defect into a confident green.
  • The transport can resume from the last item the case saw. Does that rescue the assertions?
    Only if the case proves it. Whether a resume happened, and how far back it reached, are properties of the transport and its deployment, not things a case may assume. If you rely on it, assert it: check the items either side of the gap are contiguous by whatever sequence value the payload carries, and treat any hole as a gap after all.
  • Should a gap fail the case, or mark it inconclusive?
    Distinguish them wherever the result store can carry more than pass and fail. A gap says the case could not observe, not that the product misbehaved, and burning it as a product failure trains people to ignore red. Where only two states exist, fail it - but put the gap in the first line of the failure message so triage takes seconds rather than an afternoon.

It is the difference between a camera that recorded all night and one that lost twenty minutes. It can still prove someone was in the room; it can no longer prove nobody was.

saying these in an interview costs you the question

  • Reconnects and re-asserts as though nothing happened
  • Keeps an absence check valid across a dropped connection
  • Assumes the transport replayed everything missed
  • Counts deliveries across a reconnect
  • Passes the case because the record looks plausible
  • Hides the gap from the failure message