On a long-lived shared runner, how far do you go beyond driver.quit() to stop Selenium sessions accumulating?
answer
- Start from the only clean path
- Name what quit() cannot cover
- A backstop, never a replacement
- Blast radius on a shared machine
- Count how often the sweeper fires
basics
~20 sMake quit() unconditional first, then add one bounded backstop for the runs where it cannot execute: a scoped pre-run sweep, a disposable run environment, or moving browsers to a remote end that reaps its own abandoned sessions.
solid answer
~40 s`quit()` is the only clean path: it deletes the session, terminates the browser, removes the generated profile directory and stops the local driver executable. But it is code, and code needs a live process, so a job timeout, a reboot or an out-of-memory kill leaves residue no teardown can prevent. Above that baseline I choose by blast radius. A disposable run environment removes residue unconditionally and costs provisioning time. A scoped pre-run sweep is cheap but must match on process ancestry, account or age rather than on the browser's name, or it eventually kills a colleague's live session. Pushing browsers to a remote end moves the residue somewhere that reaps its own abandoned sessions. Whatever I add, I count what it reaps: a backstop that fires regularly is a defect list, not a success.
go deeper
Know that quit releases the browser, the driver process and the temporary profile, and that a run killed by the agent releases none of them. The strategy layered above that is not yours to own yet.
Be able to explain why a kill signal never runs teardown code, and to list precisely what is therefore left on the machine afterwards: a browser, a driver executable and a profile directory.
Describe a backstop you have actually operated: what it reaps, how it is scoped so it cannot touch another job's work, and how you know whether it fires often enough to be hiding a bug.
Argue the layers and their costs - cleanup script, disposable environment, or pushing browsers to a remote end - and say how you keep a backstop from quietly becoming the primary mechanism.
## What quit() covers, and where its guarantee stops `driver.quit()` is the only clean path out of a session. It sends **Delete Session**, so the remote end closes every window the session owns, terminates the browser, and discards the session's state including the **temporary profile directory** the driver generated; on a Grid the node frees the slot; on a local run the client stops the driver executable it launched. When `quit()` runs, nothing is left behind. The whole question at this level is what happens on the runs where it does not run - because `quit()` is code, and code needs a live process to execute it. On a laptop that is a nuisance. On a runner that stays up for weeks and is shared between jobs, the residue is cumulative, and eventually it is the thing that fails your builds. ## The paths that skip it 1. The job hits its wall-clock limit and the agent kills the client. A kill signal does not run a `finally` block. 2. The host reboots, the run environment is stopped, or the kernel's out-of-memory killer picks your client. 3. The connection to the driver or to a remote end drops before the delete command lands, orphaning the session on the far side. 4. The driver executable itself crashes, leaving the browser it started with no parent to shut it down. None of those are fixable by writing better teardown code, and that is the honest framing: `quit()` covers the expected paths, and a long-lived shared host needs an answer for the rest. ## The layers, and what each costs | layer | what it removes | what it costs | |---|---|---| | `quit()` on every session | everything, cleanly | nothing - this is the baseline, not an option | | scoped pre-run sweep of stale processes and temp directories | residue from earlier runs | can match work that is still live on a shared host | | a disposable run environment per run | all residue, unconditionally | provisioning time and infrastructure to own | | remote sessions instead of local browsers | residue leaves the runner entirely | the remote end now carries abandoned sessions until its own timeout | | monitoring process counts, live sessions and temp size | nothing - it only tells you | someone has to own the signal and act on it | ## How I would decide 1. **Baseline first.** Make `quit()` unconditional before buying anything above it. Every layer in that table is an expensive workaround for a cheap bug if the suite still forgets to quit. 2. **Choose by blast radius.** On a runner rebuilt for each job, disposal already solves it and a sweep is redundant. On a machine several jobs share, a broad process killer is genuinely dangerous - the residue of one job and the live browser of another look identical from the outside. 3. **Prefer disposal to cleanup where you can afford it.** Throwing the environment away is the only option that needs no rules about which processes are dead, which is why it survives the arrival of a second browser, a second suite and a new team. 4. **Consider pushing browsers off the runner.** With remote sessions the client owns no browser and no driver executable, so an unbounded local leak becomes a bounded remote one: an abandoned session holds its slot until the remote end's own session timeout reaps it. That timeout is a Grid configuration decision owned elsewhere; what you own is the choice to depend on it. 5. **Instrument whatever you add.** A backstop with no counter tells you nothing about whether the underlying bug is still there. ## Making a backstop honest - Scope it. Match on process ancestry, on a dedicated account, or on an age threshold longer than your longest run - never on the browser's name alone. - Run it **before** a run rather than after, so it can never race the run it is protecting. - Publish what it reaps per run. Zero is healthy; a steady non-zero count is a list of tests whose sessions never received Delete Session. - Name it for what it is. A script called "cleanup" acquires a reputation as infrastructure; one that reports three leaked sessions reaped stays visibly a symptom. ## What goes wrong when the backstop becomes the mechanism - Missing `quit()` calls stop being visible, so the number of them grows and the suite quietly loses the ability to run anywhere without the sweeper. - The killer's pattern eventually matches something live - a browser another job is using, or a developer's own window on a shared box. - Killing a driver skips its own cleanup, so the temporary profile directory it would have deleted survives; you convert a process leak into a slower disk leak. - Capacity planning drifts. The fleet gets sized for the leak instead of the workload, and the extra headroom hides the next leak too.
- How do you tell whether your backstop is masking missing quit() calls?Instrument it. Have the sweep log and count what it reaps on every run; a healthy suite reaps nothing. A steady non-zero count is a defect list - each entry is a session that never received Delete Session - rather than evidence that the backstop is earning its place.
- Why does moving to remote sessions change the risk rather than remove it?The client stops owning a browser and a driver executable, so nothing accumulates on the runner itself. The residue moves to the remote end, where an abandoned session holds its browser and its slot until that end's own session timeout reaps it. Bounded and centrally observable, but not free.
- Is a per-run disposable environment always the right answer?No. It removes residue unconditionally, but you pay provisioning time on every run and you need the infrastructure to rebuild environments reliably. On a small team with one dedicated runner, an unconditional `quit()` plus a scoped pre-run sweep is usually the cheaper equilibrium.
saying these in an interview costs you the question
- Treats a scheduled process killer as the primary cleanup mechanism
- Assumes quit() always runs, so no residue is ever possible
- Kills every browser process on a machine other jobs share
- Ignores that an abandoned remote session still holds its slot
- Measures success by free disk space rather than by leaks fixed