skip to content

questions

4

Why must an automated case open a push connection before triggering the action that feeds it?

level: middleimportance: must knowfreq 62%

answer

  1. Two steps whose order is not fixed
  2. Listening starts when you attach, not before
  3. Lengthening the deadline never changes this
  4. A returned client call proves very little

basics

~20 s

Open first, then act. A push connection delivers only what is emitted after it is established, so a case that triggers first can miss the update it exists to check and then passes or fails on timing alone.

solid answer

~50 s

A long-lived push connection is not a history; it carries what the producing side emits while the connection is established and the subscription registered. So the case must **subscribe before it acts**: open the connection, attach the collector that captures deliveries, confirm the far side has registered the subscription rather than trusting that the client call returned, and only then perform the action that produces the update. Capture a correlation identifier from that action and match on it. The classic misdiagnosis is worth naming. When the case subscribes too late the symptom is an expired wait, so someone lengthens the wait - which cannot help, because the update was never delivered to this connection at all. A useful discriminator: a genuinely slow arrival gets rarer as the deadline grows, while a subscribe-too-late race holds its failure rate however long you wait.

code

pseudocode · 12 lines
pseudocode
# listen, confirm registration, then act
channel  = open_push_connection(watching = "orders")
received = []                                   # buffer owned by this case
channel.on_message(m -> received.append(m))     # attach BEFORE confirming

await channel.subscription_acknowledged()       # far side has registered it

order_id = place_order(item = "widget").id      # the triggering action

await until(deadline_policy, () ->
    any(m in received where m.order_id == order_id
                        and m.type == "order_confirmed"))

go deeper

for a junior

Be ready to say that a push connection only carries what is sent while you are connected, and that the case therefore opens it before performing the action that produces the update.

for a middle

Explain the race concretely: which two steps are unsynchronised, why the symptom looks like an expired wait, and why lengthening that wait cannot possibly fix it.

for a senior

Show that you separate the connection being established from the subscription being registered, and describe how you prove the second before acting when the channel offers no acknowledgement.

for a principal

Own the suite-wide rule: a per-case connection opened before the trigger, or a shared one demultiplexed by injected identifiers, and the cost each choice places on concurrent connection count and on diagnosis.

## The race in one sentence A long-lived push connection carries only what the producing side emits **after** that connection is established and its subscription registered. A case that performs the triggering action first and opens the connection second is not receiving the update late - it is not receiving it at all. Whether such a case passes depends on which of two unsynchronised things happened first, and that is the working definition of a flaky test. The window is small, which is why the design survives review: it is the interval between the action reaching the producing side and the subscription becoming live. That is often single-digit milliseconds on a quiet machine, and far longer on a loaded parallel run. ## Why the symptom misleads The failure almost never announces itself as an ordering bug. - On a loaded machine the subscription happens to win the race and the case is green. - On an idle machine the producer wins and the case reports an expired wait. - Under a wide parallel run the winner changes from run to run and from worker to worker. Because the observed symptom is *a wait that expired*, the reflex fix is to lengthen the wait. That cannot work: no deadline recovers an update that was never delivered to this connection, and a longer one only makes the failing run slower. One discriminator is worth memorising - a genuine slow-arrival problem gets **less** frequent as the deadline grows, while a subscribe-too-late race holds the same failure rate however long you wait. If raising the budget from ten seconds to sixty changes nothing at all, stop tuning the budget and look at the ordering. ## The ordering that works 1. Open the push connection and let it reach an established state. 2. Attach the collector, so the very first delivery is captured rather than dropped. 3. Confirm the producing side has registered the subscription - not merely that the client call returned. 4. Perform the triggering action and capture the correlation identifier it returns. 5. Assert that what the collector accumulated eventually contains an item matching that identifier. Steps 2 and 3 are the ones that get skipped, and each has its own failure mode. ### Established is not registered The client-side handle can be constructed, connected and returning cheerfully while the far side has not yet attached it to the entity or the category the case wants to watch. Subscription is frequently itself an asynchronous request travelling over the connection that was just opened. Most channels emit something observable when that completes: an acknowledgement item, an initial snapshot, a welcome frame. Wait for that observable thing, not for the constructor. Where a channel offers no such signal, the honest fallback is application-level. Have the case provoke a cheap, harmless change it can recognise - a no-op touch on a record the case owns - and treat receiving it as proof that deliveries reach this connection. It costs one round trip and converts an invisible registration race into an explicit, loudly failing step whenever the pipe is dead. ### Attach the collector before you confirm If the collector is attached only after the acknowledgement has been awaited, anything delivered between the two is discarded - and the acknowledgement is precisely the moment traffic starts flowing. Attach first, then await. ## Three orderings, side by side | Order the case uses | What it can observe | Verdict | |---|---|---| | Act, then open | Only what is emitted after the open; the wanted update is usually already gone | Broken; green by luck | | Open, then act | Everything after the connection is established, minus whatever registration lag swallows | Right idea, still racy under load | | Open, attach, confirm registration, then act | Every update the action causes | Correct and stable | ## What this does not license - **Do not insert a pause before the action** to give registration time. That is the same race with a wider window and a fixed cost added to every case. - **Do not assume the channel replays what you missed.** Some transports can resume from a last-seen position; whether a given deployment enables that, and how far back it reaches, is a property of that system and not something a case may lean on unless it set the resume up itself and asserts that it happened. - **Do not open one connection for the whole suite.** That removes the per-case race and installs a worse one: every case sees every other case's traffic, so a loose match can be satisfied by a neighbour, and the resulting false green is indistinguishable from a real pass. Where a shared connection is unavoidable, demultiplex it by a correlation identifier each case injects into its own data. ## The cost you accept Opening first means the connection is held for the whole case rather than for its last few seconds, so a wide parallel run holds roughly one connection per active case. That is a real capacity number worth checking against the system under test, and it is the price of an assertion that does not depend on timing. It is almost always the right trade: a hundred concurrent connections poses a capacity question you can answer, while a suite that is green by luck poses one you cannot.

  • The channel gives you no subscription acknowledgement at all. How do you still prove the connection is live before acting?
    Provoke a cheap, harmless change you can recognise - a no-op touch on a record the case owns - and treat receiving it as proof that deliveries reach this connection. Then perform the real action. It costs one round trip and turns an invisible registration race into an explicit step that fails loudly when the pipe is dead.
  • Your case subscribes first and still misses the update roughly one run in fifty. What is left to suspect?
    Registration on the producing side is itself asynchronous, so the acknowledgement you waited on may confirm the connection rather than the subscription. Check what it actually means, and add an application-level proof if it only confirms the connection. Also check the collector is attached before that wait, so nothing delivered in between is silently discarded.
  • Why not open one push connection in suite setup and share it across every case?
    It removes the per-case race and installs a worse one: every case sees every other case's deliveries, so a loose match can be satisfied by a neighbour's traffic and the false green is indistinguishable from a real pass. If sharing is unavoidable, demultiplex by a correlation identifier each case injects into its own data.

It is the difference between standing at the arrivals gate before the flight lands and turning up an hour later to ask who came through. The second question has no answer; the crowd has dispersed.

saying these in an interview costs you the question

  • Adds a pause before the action instead of subscribing first
  • Says a longer wait will recover the missed update
  • Treats a returned client call as proof of subscription
  • Assumes the channel replays everything sent before connecting
  • Shares one connection across cases and matches the newest delivery
open as a page

What goes wrong when an automated case leaves the push connection it opened still open?

level: juniorimportance: should knowfreq 45%

basics

~20 s

Leaked connections accumulate until the run exhausts its limits, and their receivers keep collecting, so a later case can match an update an earlier one caused. Close in a cleanup step the harness runs even when the case fails.

open as a page

How should an automated case assert on updates that a push connection delivers at unpredictable moments?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Collect, then assert over the collection. Attach a receiver at subscription time that appends every delivery to a buffer the case owns, then assert that the buffer eventually holds one matching a correlation identifier the case injected.

open as a page

After a push connection drops and reopens mid-case, which assertions are still valid?

level: seniorimportance: nice to knowfreq 22%

basics

~20 s

Only assertions that survive missing evidence. A gap leaves the case holding an incomplete record, so existence checks on items received after the reopen can stand, while counts, ordering and any claim that nothing was sent are void.

open as a page