skip to content

A CI runner is killed holding open hosted browser sessions. What keeps consuming, and what ends it?

level: seniorimportance: should knowfreq 54%

answer

  1. the far side cannot see your process
  2. no close command was ever sent
  3. orphans hold a slot as well as time
  4. finally does not survive a kill
  5. silence is what the reaper measures

basics

~20 s

A killed runner sends no close command, so its sessions stay open on the provider's side, consuming and holding their slots. They usually end when you delete them later, or when the far side's inactivity reaper notices the silence.

solid answer

~50 s

Nothing in the remote protocol tells a fleet that your process died. A session is state on the far side keyed by an identifier, and it ends when something deletes it — so a runner killed mid-run leaves live browsers behind it. Those orphans keep consuming under a time-based meter and keep holding their share of the account's concurrency ceiling, which is why the next job queues behind sessions nobody is using. Two things you can plan around end them: your own code, if the identifiers were written somewhere that outlived the process, or the fleet's inactivity reaper. A `finally` block does not help — it runs when a case throws, not when a process is killed — so the defences that matter are a cancellation handler on the runner and a sweep that closes whatever the live run does not claim.

code

python · 16 lines
python
# Every session goes through this helper; no case opens its own.
def run_case(case):
    handle = fleet.open_session(case.name)
    remember_open(handle.id)           # written where a dying process cannot take it
    try:
        case.execute(handle)
    finally:
        fleet.close_session(handle.id)  # ordinary exits: pass, fail, error
        forget_open(handle.id)

def on_cancel_signal(_sig, _frame):     # cancellation, not a hard kill
    for held in list_open():
        fleet.close_session(held)
    raise SystemExit("cancelled")

# Neither path runs on a hard kill. The sweep that reads list_open() later does.

go deeper

for a junior

Know that closing a session is a command your test sends, not something that happens when your process exits. If the job is cancelled, that command is never sent at all.

for a middle

Be able to explain why a finally block covers a thrown assertion but not a killed process, and to name the two things that can still end an orphaned session.

for a senior

Expect a scenario. Say what the orphans cost while they live, what bounds the exposure, and which of your defences survives cancellation, a killed container and a vanished host.

for a principal

You will be asked whether orphan sweeping is worth building. Frame it as bounded exposure per orphan against a measured cancellation rate, and argue from recorded rows rather than from the worst week anyone remembers.

## Why the far side never notices A remote session is state held on the far side. Your client holds an identifier; the fleet holds the browser and the slot that session occupies. The protocol has exactly one way for you to say *I am finished*: a delete for that identifier. There is no heartbeat from your side that the fleet watches, and the transport offers nothing either: commands come and go over connections that open and close, so a dropped connection is indistinguishable from a client that is simply thinking. You can watch this in an open implementation. Selenoid — unmaintained by its own README — logs a client disconnection and answers with a bad-gateway status without removing the session or releasing the slot; its proxy releases a slot only in the branch that handles a delete. Your process going away is not an event on the far side. So when a CI runner is cancelled, killed for memory, or loses the host underneath it, the sessions it opened do not notice. They stay open, they stay yours, and they keep consuming under any meter that counts time a session is open. The usual cause is not a crash either, but somebody pressing cancel on a run that was already behaving badly — which means orphans arrive in clusters, on the days the suite is at its worst. ## What an orphan costs while it lives Two costs, and teams usually see only the first: - **Consumption.** Under a time-based meter an orphan consumes at full rate while doing nothing. - **A held slot.** The orphan still counts against whatever ceiling your account enforces, so the next run finds less capacity than it expected and waits for browsers nobody is using. Which ceiling binds in that situation is its own subject; what matters here is that a cancelled job degrades the runs that follow it. A quieter third cost: any capture the provider started for those sessions keeps running, producing artefacts nobody chose to keep. ## What actually ends an orphan Something has to delete the session, and these are the candidates worth planning around. 1. **You, later.** This works only if the identifiers were written somewhere that outlived the process — a file on a mounted volume, or a label the fleet can list back to you. Identifiers held only in the dead process's memory died with it. 2. **The fleet's own inactivity reaper.** The open implementations show the shape plainly. Selenium Grid's node takes a `--session-timeout`, documented as automatically killing a session that has had no activity for that long and releasing the slot for other tests. `docker-selenium` exposes that same setting to its images as the environment variable `SE_NODE_SESSION_TIMEOUT`. Selenoid, unmaintained per its own README, carries a `-timeout` for the session idle timeout and a `-max-timeout` bounding the idle timeout a client may ask for. Other things can end it too — a vendor operator, a fleet restart, or a maximum total session length where a fleet enforces such a cap — but none of those is something you can plan around. A hosted provider has something equivalent — but its window and its rule are facts to establish with that vendor, not carried over from an open project. What you may say with confidence is the shape: an orphan left by a killed runner is **silent**, and silence is exactly what an inactivity reaper is built to catch. That bounds your exposure per orphan. It does not remove it, and does nothing for the held slot until the reaper fires. ## The rings of teardown, and which one fails here 1. **Per-case teardown.** A `finally` around the case body closes the session on a pass, a failure, an error and a thrown assertion. It runs when the *case* ends — not when the *process* does. 2. **A cancellation handler on the runner.** Catching the termination signal a CI platform sends before it gives up buys a short window in which to close what you hold. This covers cancellation, the common cause. It loses to a hard kill, to an out-of-memory kill, and to a host that simply disappears. 3. **A sweep outside the run.** A separate job that lists what your account has open and closes anything the live run does not claim. It is the only ring of yours that survives a process which never executed another instruction. The instructive part is that the first ring — the one everyone names first — is precisely the one that does not help here. Exception safety unwinds a stack; a killed process does not unwind. ## Bounding the exposure - **Label sessions with the run that created them**, wherever the fleet lets you name a session, so a sweep can tell a live run's from a dead one's. - **Record every open** somewhere the process cannot take it with it, so orphans are counted rather than sensed. - **Prefer a cancellation path that lets your handler run**, with a budget small enough to finish before the platform stops waiting. - **Decide whether the sweep is worth building** by comparing a measured orphan rate against the reaping window you established. - **Do not claim an orphan runs forever.** Where a reaper exists, exposure per orphan is bounded by its window; the honest worry is how often orphans happen and what their slots block meanwhile.

  • Your suite's teardown is exception-safe and orphans still appear. Where are they coming from?
    From the exits an exception handler never sees: a cancelled job, an out-of-memory kill, a runner host reclaimed underneath you, and a hard pipeline timeout that stops the process rather than the case. None of those unwind the stack, so no teardown runs. They also cluster, because cancelling is what people do to a run that is already misbehaving.
  • How would you find out how long an orphaned session survives on a provider whose source you cannot read?
    Measure it. Open a session deliberately, stop talking to it, and watch the vendor's own record of that session until it ends. Repeat with a session that keeps polling, because an inactivity reaper measures silence rather than elapsed time and the two experiments can give different answers. Then take the result to the vendor and have them confirm the rule rather than generalising from your own probe.
  • Why is a held concurrency slot sometimes worse than the consumption itself?
    Because consumption is a bill and a held slot is a stoppage. An orphan keeping its share of the account ceiling makes the next run wait for capacity nothing is using, so a cancelled job quietly degrades the runs after it. Which ceiling binds in that situation, and how pipelines end up starving each other, is a subject of its own.

An open session is a taxi with the meter running, not a phone call that ends when you hang up. Your end going away changes nothing; the meter stops when someone tells the driver to stop, or when the driver notices nobody has spoken for a while.

saying these in an interview costs you the question

  • Claims a dropped connection closes the remote session
  • Says an exception-safe teardown covers a cancelled job
  • Assumes an orphaned session runs until a human notices it
  • Thinks the only cost is money, not the held slot
  • Guesses the provider's reaping window instead of establishing it
  • Treats cancellation as rarer than crashes