skip to content

How does a suite that leaks driver sessions show itself before the run host runs out of resources?

level: seniorimportance: nice to knowfreq 24%

answer

  1. The signature is monotonic, not random
  2. A staircase where a sawtooth belongs
  3. Late cases fail, whichever they are
  4. Count acquires against releases per scope
  5. Tag each session with its acquiring case

basics

~20 s

Run duration climbs within a single run, host resource use rises and never falls back between cases, orphan processes survive the run, and late cases fail with resource errors that name the machine rather than the feature.

solid answer

~50 s

The signature is **monotonic**, which is what distinguishes it from ordinary flakiness. Cases run slower as the run proceeds on unchanged code; the host's memory, handle and process counts climb during each case and never sawtooth back down; processes started by the suite survive after it reports finished; and the failures cluster at the end of the run, name a resource rather than a behaviour, and move to whichever cases run last when you reorder. Re-running the failing cases alone passes, which is why leaks get mislabelled intermittent. Catch it before the machine does with **balance counting** — increment on acquire, decrement on release, assert zero at the end of each owning scope — and by tagging each session with the case that acquired it so the report names the culprit instead of the victim.

code

pseudocode · 13 lines
pseudocode
open_sessions = {}   # session_id -> acquiring case name

def acquire_owned_session(scope, case_name):
    session = acquire_session()
    open_sessions[session.id] = case_name
    scope.on_exit(lambda: open_sessions.pop(session.id, None))
    scope.on_exit(release_quietly, session)
    return session

def assert_balanced(scope_name):
    if open_sessions:
        fail(scope_name + " leaked sessions acquired by: "
             + list(open_sessions.values()))

go deeper

for a junior

Know that a session which is never closed keeps holding memory and a connection, and that the machine running the suite eventually refuses to give out more. The case that fails is usually not the case at fault.

for a middle

Describe the observable signature: duration climbing within one run, resource use that staircases instead of sawtoothing, orphan processes after the run, and failures naming a resource rather than a behaviour.

for a senior

Demonstrate the diagnosis. Separate a leak from an undersized host by shape rather than level, use balance counting and owner-tagged sessions to name the acquiring case, and confirm with a repeat run of one case while sampling host counters.

for a principal

Decide what the harness guarantees by default, so a leak fails the run that introduces it instead of the pipeline six months later. Weigh the small permanent cost of balance checks and boundary sampling against the cost of diagnosing this in production hours.

## The outward signature of a leak A suite that never releases its driver sessions does not announce itself. It degrades, and the degradation has a recognisable shape: - **Run duration climbs within a single run.** The last hundred cases take visibly longer than the first hundred, on unchanged code. Each abandoned session is still holding memory, still holding a connection, sometimes still driving a process that keeps doing work. - **Resource use on the run host rises and never falls back.** Between cases the gauge should sawtooth — up during a case, down at its end. A leak turns the sawtooth into a staircase. - **Orphan processes survive the run.** After the suite reports finished, processes it started are still alive. Counting them before and after a run is often the whole diagnosis. - **Failures cluster at the end of the run and move when you reorder.** Whichever cases run last fail, regardless of what they test. - **The failures name a resource, not a behaviour.** Cannot allocate, cannot connect, no handles available, timed out starting a session — messages about the machine rather than about the feature. - **A re-run of the failing cases alone passes immediately**, which makes the whole thing look intermittent and invites a re-run policy instead of a fix. ## Why it hides until it is expensive Three properties conspire to delay the discovery: 1. **The budget is generous.** Machines allow many concurrent handles and connections, so a leak of one session per case needs hundreds of cases before anything breaks. A small suite leaks happily for months. 2. **The symptom appears far from the cause.** The case that fails is the one unlucky enough to run when the budget ran out. The case with the missing release passed. 3. **It grows with the suite, not with the change.** Nobody changed the leaking code; the suite simply got big enough. There is no suspicious commit to bisect to, which defeats the first instinct. | Observation | Innocent explanation | What makes it a leak | |---|---|---| | Late cases slow | heavier cases scheduled late | reordering moves the slowness to whatever runs last | | Memory high at run end | a large data set held for the run | it never falls back between cases either | | Resource-exhaustion errors | an undersized machine | opened and released counts do not match | | Passes when re-run alone | genuine intermittency | the run host was restarted between the two attempts | ## Find it before the machine does Detection should be built into the harness, not left to the day the pipeline dies. In rough order of cost: - **Balance counting.** Increment a counter on acquire, decrement on release, and assert it is zero at the end of each owning scope. This is a handful of lines and it catches the defect on the run that introduces it, naming the exact case. - **Tag every session with its owner.** Record the acquiring case's identifier when a session is created, and print the identifiers still open at the end of the run. The report then reads *"three sessions still open, all acquired by the checkout cases"*, which is a fix, not an investigation. - **Watch the host between cases.** Sample handles, processes and memory at each case boundary and record the value in the run's results document. A staircase in that column is unambiguous, and it is visible long before anything fails. - **Repeat one case many times in a single process.** Running the same case a few hundred times in one run is the fastest way to turn a slow leak into a fast one; if the resource curve climbs monotonically across identical iterations, the case leaks. ## Where the leaks actually come from Four sources account for most of them, and all four are shapes rather than mistakes: - **A session acquired inside a helper that returns it**, leaving the caller with an obligation nothing enforces. - **A release reachable only on the success path**, so every failing case leaks by definition — which is why leaks accelerate exactly when a suite is unhealthy. - **A second session acquired mid-case** — to check something as another user, or to observe a side effect — outside whatever mechanism owns the first one. - **A release that throws and is swallowed.** The call was made, the resource was not freed, and the log line that said so was suppressed to keep the run green. The last one deserves the most suspicion, because it looks like correct code in review. A release that can fail needs its failure recorded somewhere the suite reads, or it is indistinguishable from no release at all.

  • How do you tell a leaked session apart from an undersized run host?
    Compare the shape, not the level. An undersized machine is at a high but flat level from the first case; a leak climbs monotonically and never returns to baseline between cases. The decisive check is the balance count: if acquisitions and releases match at each scope's end, the machine is simply too small.
  • Why do leaks tend to appear as a sudden problem rather than a gradual one?
    The resource budget is generous, so a leak of one session per case runs fine for hundreds of cases. The threshold is crossed when the suite grows past it, not when the leaking code changes, so there is no suspicious commit to bisect to and the first instinct fails.
  • What is the fastest way to confirm a suspected leak in one case?
    Repeat that single case a few hundred times in one run and sample the host's memory, handle and process counts at each iteration boundary. Identical iterations should hold a flat baseline; a monotonic climb across them confirms the leak and gives you a fast loop for verifying the fix.

It reads like a slow leak in a tyre rather than a puncture: nothing fails at the moment of the fault, and the thing that eventually goes wrong happens far from where the hole is.

saying these in an interview costs you the question

  • Calls late-run resource failures intermittent and adds a re-run policy
  • Raises machine limits instead of finding the missing release
  • Expects a leak to fail at the case that causes it
  • Never checks whether processes survive after the run reports finished
  • Bisects recent commits for a leak that grew with suite size