skip to content

What goes wrong when an automated case leaves the push connection it opened still open?

level: juniorimportance: should knowfreq 45%

answer

  1. Something opened was never given back
  2. The next case pays the bill
  3. A failing assertion skips the line after it
  4. Cleanup belongs where failures cannot skip it

basics

~20 s

Leaked connections accumulate until the run exhausts its limits, and their receivers keep collecting, so a later case can match an update an earlier one caused. Close in a cleanup step the harness runs even when the case fails.

solid answer

~50 s

Two costs, one of them silent. The loud one is exhaustion: every leaked connection holds a handle on the runner and a registration on the system under test, so after enough cases the run starts failing for reasons unrelated to what those cases test - or refuses to exit at all, because a receiving loop is still alive. The quiet one is contamination: the leaked receiver keeps appending deliveries, so a later case that matches loosely can be satisfied by traffic an earlier case caused, and that false green is indistinguishable from a real pass. The fix is placement, not diligence. Register the close **at the moment of opening**, through the per-case cleanup the harness always runs, so a failed assertion cannot skip it. Detach the receiver before closing, wait for the closed state with a bound, and make the close idempotent.

code

pseudocode · 11 lines
pseudocode
open_watched_channel():
    channel = open_push_connection(watching = "orders")
    case.on_cleanup(() -> close_quietly(channel))   # runs on pass, fail and error
    return channel

close_quietly(channel):
    if channel.already_closed: return               # idempotent
    channel.detach_receiver()                       # stop collecting first
    channel.close()
    await_or_give_up(channel.is_closed, bound = short)
    # never rethrows: a failed close is logged, not a case verdict

go deeper

for a junior

Be ready to say that whatever a case opens the case gives back, and that the close belongs in a cleanup the harness runs even when the case fails, rather than on the last line of the body.

for a middle

Explain both costs - handles accumulating until unrelated cases fail, and a still-live receiver feeding a later case's match - and why the failing cases are precisely the ones that leak.

for a senior

Show how you make the close deterministic and idempotent, why the receiver is detached before the connection is closed, and how you diagnose a run that finishes its cases but will not exit.

for a principal

Own the rule that nothing is opened without its release registered in the same step, and decide how live-connection counts are monitored so the suite catches the drift instead of a timed-out job.

## Four things a leaked connection does A push connection is not just a client object; it is a pair of allocations, one on each side, plus something running that consumes deliveries. Leaving it open after the case ends has four distinct consequences, and only the first is the one people expect. 1. **Handles accumulate.** Every leaked connection holds a handle the operating system counts on the runner and a registration on the system under test. Long before a suite finishes, a run that leaks one connection per case starts failing on limits - and it fails in cases that have nothing to do with the leak, which sends triage to the wrong place entirely. 2. **The leaked receiver keeps collecting.** It is still appending deliveries into a buffer that outlived its case. A later case that matches loosely - on the newest item, on a type, on anything but its own injected identifier - can be satisfied by traffic an earlier case caused. That is a false green, and a false green looks exactly like a real pass. 3. **The system under test behaves differently.** Fan-out cost grows with the number of live subscriptions. A run that accumulates hundreds of them is measuring a configuration nobody deploys, so its timing observations stop meaning anything. 4. **The run will not exit.** A receiving loop that is still alive holds the process open. The tell is unmistakable once you have seen it: all case results are reported, and then the job simply runs out its clock with nothing on the output. ## Why the close ends up in the wrong place Almost always because it was written as the last line of the case body. That placement is correct only on the happy path, and it fails on precisely the runs where it matters: **a failed assertion leaves the body early**, so the cases that fail are the ones that leak. The leak therefore grows fastest on a bad day, when the run is already red and hardest to read. The same is true of any close guarded by a condition that a failure can skip, and of a close that lives after an early return. ## Register the release beside the acquisition The fix is placement, not diligence. At the moment the case opens the connection, register its release with the per-case cleanup the harness runs on pass, failure and error alike. The two lines sit together, so nobody can add an early return that skips one, and a reviewer can see the pairing without reading the rest of the case. Where a helper opens the connection, the helper registers the cleanup - never the caller. A helper that hands back an open connection and trusts every future caller to release it is a leak waiting for its first hurried change. ## Closing deterministically Fire-and-forget is not closing. | Close style | What it leaves behind | |---|---| | Last line of the case body | Every failing case's connection, which is the worst possible selection | | Fire the close, return immediately | A window where deliveries still land in the old buffer and the far side still holds the registration | | Detach the receiver, close, wait for closed with a bound | Nothing, and the wait is bounded so cleanup cannot hang the run | Order matters inside that last row. **Detach the receiver first**, then close. Detaching first stops deliveries reaching a buffer whose case has ended; closing first leaves a brief window in which they still arrive. Then wait for the connection to report closed, with a bound, so a far side that never answers costs a few seconds rather than the whole job. Make the close idempotent too. A case that closes explicitly and then hits the registered cleanup should not throw on the second attempt; check for already-closed and return. ## Cleanup that cannot itself fail the run A close that throws after the assertions have run tells you nothing about the behaviour under test. Turning it into a case failure produces red runs nobody trusts, and people start ignoring the colour. Swallow the error, log it with enough detail to identify which connection it was, and let a separate systemic check catch a real problem: - Count live connections at the end of the run and assert the count returned to its baseline. - Count how often the bounded close wait expired; a rising number means the far side is struggling, which is a finding about the system rather than about the suite. That split keeps the case's verdict about the product and moves housekeeping problems into a place where they are visible without being confused with defects. ## The rule in one line Whatever a case opens, the case gives back - through a path the harness runs whatever the outcome, in an order that stops deliveries before it releases the connection, with a bound so cleanup can never be the thing that hangs the run.

  • Why is closing at the end of the case body not good enough?
    A failed assertion leaves the body early, so the close never runs - and the failing cases are exactly the ones whose connections leak, which is why the leak grows fastest on a bad day. Registering the cleanup at open time moves the close onto a path the harness runs whatever the outcome.
  • The suite reports every case and then hangs. What would you look at first?
    A receiving loop still alive. A run only exits when nothing is holding it open, and a leaked push connection with a live receiver does exactly that. The tell is that the results are complete and the job still times out. Find the connection nobody closed and fix where its cleanup is registered, rather than adding a forced exit.
  • Should the cleanup fail the case if the close itself throws?
    No. A close that fails after the assertions have run says nothing about the behaviour under test, and turning it into a failure produces red runs nobody trusts. Swallow it, log enough detail to identify the connection, and let a separate check on live-connection counts at the end of the run catch a systemic problem.

saying these in an interview costs you the question

  • Closes the connection on the last line of the case body
  • Says leaked connections are harmless because the run ends
  • Ignores that a leaked receiver keeps collecting deliveries
  • Fires a close and never confirms it happened
  • Fails the case when the close itself throws
  • Adds a forced exit to stop the run hanging