skip to content

Nightly Selenium runs leave chromedriver processes alive on the build host between builds. How do you find and stop that?

level: seniorimportance: must knowfreq 57%

answer

  1. Look at the host, not the heap
  2. Whoever started it owns stopping it
  3. One per run versus a burst
  4. Some kills run no Java at all
  5. A sweep is a net, not a fix

basics

~20 s

A driver service you build and start yourself is yours to stop. Orphans pile up when stop() is skipped on an exception path or the JVM is killed, so pair every start() with a finally, and add a host-level sweep for the kills.

solid answer

~50 s

Start by looking at the host rather than the code: list the driver processes between runs, note how many there are and which ports they hold, and check their start times against the nightly schedule. One leftover per run points at a service that is built and started explicitly and never stopped; a burst per run points at a service started per case. The code fix is ownership — whoever calls `start()` calls `stop()` in a `finally` that wraps the whole run, and only one place in the suite does either. That still leaves the case no Java code can cover: if a build timeout kills the JVM outright, nothing in the process runs, so the job needs a cleanup step of its own. Finally, make the leak loud instead of silent — a pinned port fails fast when yesterday's driver still holds it.

code

java · 17 lines
java
ChromeDriverService service = new ChromeDriverService.Builder()
    .usingPort(9515)
    .build();
service.start();
try {
    WebDriver driver = new ChromeDriver(service, new ChromeOptions());
    try {
        driver.get("https://lims.internal.example/samples");
        driver.findElement(By.cssSelector("#sample-log tbody tr:first-child")).click();
    } finally {
        driver.quit();
    }
} finally {
    if (service.isRunning()) {
        service.stop();
    }
}

go deeper

for a junior

Know that the driver is a real operating-system process that can outlive your test run, and that a service you started yourself has to be stopped yourself.

for a middle

Explain the exception path and the shared-service path that produce orphans, and show the finally block scoped to match where the service was started.

for a senior

Separate leaks your code causes from leaks an external kill causes, use process listings and start times to tell them apart, and justify where a host-level sweep is legitimate.

for a principal

Own the policy across teams: who may build services explicitly, how agents are checked for leftovers before a run, and whether a leaked process fails the build rather than being cleaned up quietly.

## What an orphan actually is An orphaned driver is an operating-system process — `chromedriver`, `geckodriver` or `msedgedriver` — that is still running and still holding its port after the run that launched it has gone. It is not a Java object, so no amount of reasoning about references reaches it. It consumes memory, it may still be holding a browser process beneath it, and on a shared build agent a handful of them per night becomes a slow resource leak that surfaces weeks later as an agent that cannot start anything at all. The reason this is a **service** problem rather than a session problem is ownership. The process was started by a `DriverService`, and only something holding that service can shut it down. ## Who owns stop() | | Implicit service (`new ChromeDriver()`) | Explicit service (`build()` then `start()`) | |---|---|---| | Who starts the process | Selenium, during construction | your `start()` call | | Who holds the reference | Selenium, internally | your code | | Who can call `stop()` | only Selenium, which owns it | only you — nothing else will | | Port | free ephemeral by default | whatever you configured | | Typical leak | rarer, since Selenium owns both ends | the common case | That table is the whole diagnosis in miniature. A suite that never builds a service explicitly leaks driver processes far less often, because the object that started the process also owns shutting it down. The moment a suite builds its own service — usually to pin a port, capture the driver's log, or share one process across cases — the duty moves to your code and nothing puts it back. ## How the leak actually happens 1. **The exception path.** `service.start()` runs, a step throws, and the `stop()` two lines further down is never reached because there is no `finally` around it. 2. **The forgotten shared service.** One service is started once for the whole sample-tracking suite, and the shutdown is written in a per-case hook that quietly does nothing, or in no hook at all. 3. **Two owners.** Setup code starts a service, a helper starts another because the first was not visible to it, and only one gets stopped. 4. **The killed JVM.** A build timeout or an agent restart sends the process a signal it cannot handle. **No Java code runs at all** — not a `finally`, not a shutdown hook — and whatever the JVM had spawned is simply inherited by the operating system. ## Finding them - List the driver processes on the agent **between** runs, when the count should be zero. Anything there is an orphan by definition. - Read the **start times**. Processes whose ages line up with the nightly schedule tell you it is one per run; a cluster from a single timestamp tells you one run leaked many. - Check **which ports** they hold. A leftover on a pinned port explains tomorrow's bind failure; leftovers on scattered ephemeral ports mean each run picked a new one and never cleaned up. - Check whether a **browser process** is still parented beneath each driver. That distinguishes a driver left running with nothing to do from one still holding a whole browser and its temporary profile. ## Fixing it, in the order that matters 1. **Single ownership.** Exactly one place in the suite builds and starts the service, and the same place stops it. Anything else that needs the driver receives it rather than starting its own. 2. **A `finally` that wraps the whole run**, not one per case, matching the scope the service was started at. Whatever throws in between, `stop()` still runs. 3. **A host-level sweep** in the job itself, before or after the run, for the killed-JVM case that no Java code can reach. Treat it as a safety net, never as the primary fix — a sweep that quietly cleans up ten orphans a night hides a real defect. 4. **Make it loud.** A pinned port turns an orphan into an immediate bind failure at the start of the next run; a check that fails the build when driver processes are found before it starts turns a slow leak into a red build with an obvious cause. ## The judgment an interviewer is listening for Weak answers reach straight for "kill the processes in a shell script". Strong answers separate the two populations: leaks caused by code that skipped its own shutdown, which are fixable and should be fixed, and leaks caused by the run being killed from outside, which code cannot fix and which need an environment-level answer. The sweep is legitimate only for the second population, and only if the first has actually been closed — otherwise you have built a machine for hiding your own defect, and the next symptom will be a build agent that runs out of memory with no explanation in any test log.

  • How do you tell a leaked driver process from a leaked browser process?
    By name and by parentage. The driver is chromedriver, geckodriver or msedgedriver; the browser is Chrome or Firefox, usually parented beneath the driver that launched it. A driver alone means the process outlived its shutdown; a driver with a browser still under it means the whole stack survived, and that one is the expensive kind on a shared agent.
  • Would a host-level sweep before every run be enough on its own?
    It stops the accumulation but hides the cause. A sweep that removes several orphans nightly is evidence of a real defect in ownership, and it also risks killing a driver belonging to a concurrent job on a shared agent. Use it as a safety net for the killed-JVM case after the code path is fixed, not instead of fixing it.
  • Why can a pinned port be preferable when you suspect a leak?
    Because it converts a silent leak into an immediate failure. A free ephemeral port routes around yesterday's leftover driver and the count grows unnoticed; a pinned port cannot bind while the old process holds it, so the very next run fails loudly with an address already in use and points straight at the leak.

saying these in an interview costs you the question

  • Assuming a leftover driver process disappears when the JVM exits
  • Expecting a finally block to run after the process is killed
  • Treating a nightly kill script as the fix rather than a net
  • Starting a service in several places and stopping it in one
  • Blaming the browser when the leaked process is the driver