skip to content

A Selenium suite is green at two workers and flaky at six on one host - how do you tell a per-worker collision from memory pressure?

level: seniorimportance: must knowfreq 65%

answer

  1. Two problems, one red report
  2. When in the run does it fail
  3. Which workers die, and how many
  4. Does a bigger machine change anything
  5. Session-start error versus a lost session

basics

~20 s

A collision fails early and asymmetrically: the first worker survives and the rest die at session start or read each other's files. Memory pressure fails late and everywhere, worsens smoothly with worker count, and kills browsers mid-test.

solid answer

~40 s

Separate them by shape before changing anything. A collision means two workers share something each should own, so it lands at session start, hits every worker but one, produces `SessionNotCreatedException` or cross-contaminated artefacts, and is unchanged by a bigger machine. Host pressure lands mid-run once several browsers have loaded real pages, spreads randomly across cases, kills browser processes so the session reports itself lost, scales smoothly with worker count, and disappears on an agent with more memory. The cheap discriminators are bisecting the worker count - a collision usually shows at exactly two - and re-running the same count on a larger host. Then audit the four resources a Selenium worker must own: its driver session, its profile directory, its driver port, and its download directory.

code

java · 21 lines
java
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
import org.openqa.selenium.chrome.ChromeDriver;
import org.openqa.selenium.chrome.ChromeOptions;

public class CatalogueExportWorker {

  public static ChromeDriver create() throws Exception {
    Path downloads = Files.createTempDirectory("catalogue-export-");

    ChromeOptions options = new ChromeOptions();
    options.setExperimentalOption(
        "prefs",
        Map.of(
            "download.default_directory", downloads.toAbsolutePath().toString(),
            "download.prompt_for_download", false));

    return new ChromeDriver(options);
  }
}

go deeper

for a junior

Know that a suite failing only in parallel is a real defect with a cause, not random flake, and that re-running it at one worker is the first thing to try.

for a middle

Explain the two mechanisms: shared state between workers, and a host that cannot hold that many browsers. Be able to name what each worker must own separately in a Selenium run.

for a senior

An interviewer expects a diagnostic sequence, not a guess: reproduce at one worker, bisect the count, audit per-worker resources, watch host memory and process counts, and re-run on a larger machine to decide.

for a principal

Own the policy that keeps this rare - how the harness derives per-worker identity by construction, what the measured worker ceiling is on each agent, and who is accountable when it is raised.

## Two failures that look identical in the report A museum exhibit catalogue suite that is green at two workers and flaky at six has one of two very different problems, and both surface as red cases with screenshots of the wrong page. **A collision** means two workers are sharing something each was supposed to own. **Host pressure** means the machine cannot hold six browsers at once. The fixes have nothing in common - one is a configuration change, the other is a capacity decision - so the first job is telling them apart rather than adding retries. ## The collision signature - Failures cluster **at session start**. The Java binding raises `SessionNotCreatedException`, whose message opens with "Could not start a new session", before any navigation happens. - The **first worker in survives** and the rest die, or the failure count tracks worker count exactly: two workers, one failure; six workers, five. - Failures are **asymmetric and repeatable in kind**: the same step fails, or a worker asserts on an artefact another worker produced - a catalogue export from a different gallery filter, a screenshot of somebody else's exhibit page. - Running the same suite at one worker is **reliably green**, and it stays green on a much larger machine only because the timing changed, not because the sharing did. ## The host-pressure signature - Failures start **late in the run**, once several browsers have loaded real pages and grown their heaps, not during startup. - They are **spread across cases** with no pattern, and a case that failed in one run passes in the next. - Browser processes **disappear mid-test**. The session then answers with a `WebDriverException` about a lost or disconnected session, and the driver log stops abruptly rather than reporting an error. - The failure rate rises **smoothly with worker count** and with page weight; six workers on a larger agent are green with no code change at all. ## Side by side | Signal | Per-worker collision | Host memory pressure | |---|---|---| | When it fails | at session start | mid-run, under load | | Which workers | all but one, deterministically | any of them, at random | | Effect of N = 1 | always green | always green | | Effect of a bigger machine | unchanged | goes green | | Typical exception | `SessionNotCreatedException` | lost session, timeouts | | The fix | make a resource unique | fewer workers, or move browsers off-host | Two rows carry most of the diagnostic weight: **which workers fail**, and **whether a bigger machine changes anything**. A collision does not care how much memory you have. ## A procedure that separates them 1. Re-run at **one worker**. If it fails there too, this is not a parallel problem at all and the rest of this is wasted effort. 2. **Bisect the worker count** - two, then four, then six. A collision usually appears at exactly two. Pressure degrades gradually. 3. Audit the four Selenium-side resources a worker must own, listed below, and confirm each one is derived per worker rather than from a constant. 4. Watch the host while the run proceeds: resident memory, the count of browser and driver processes, and whether the kernel logged an out-of-memory kill. Browsers that vanish without an error in the driver log were killed, not crashed. 5. Re-run the failing worker count on a machine with materially more memory. Green means pressure; identical failures mean a collision. ## The Selenium-side resources to check first - **The driver session.** Each worker constructs and owns its own `WebDriver`; nothing hands one worker's instance to another. - **The profile directory.** `--user-data-dir` for Chrome, or Firefox's `-profile` argument, is derived per worker rather than pinned to one path. Leaving it unset is safest, because the driver then creates a temporary profile per session. - **The driver port.** Left to `DriverService.Builder`'s default of 0, so `PortProber.findFreePort()` picks one, rather than pinned with `usingPort`. - **The download directory.** Chrome and Edge take `download.default_directory` and `download.prompt_for_download` through `setExperimentalOption("prefs", ...)`; Firefox takes `browser.download.folderList` set to `2` plus `browser.download.dir`. Six workers exporting the exhibit catalogue into one folder will each assert on whichever file happened to land last. Selenium's own Grid solves this by creating a temporary downloads directory per session and rewriting exactly these preferences before the browser starts. ## What you change once you know A collision is fixed by deriving the shared value per worker - never by a retry, a sleep, or a lock that serialises the run and gives back the parallelism you paid for. Host pressure is fixed by running fewer workers, giving the agent more memory, or moving the browsers off the test host entirely with `RemoteWebDriver` pointed at a grid endpoint, so the test process stays cheap and the browser footprint lands elsewhere.

  • A worker's browser vanishes mid-test and the driver log just stops. What does that tell you?
    That nothing inside the browser reported an error, so it was killed rather than crashing. On a memory-constrained host the kernel's out-of-memory killer is the usual cause, and the run is above the machine's ceiling. Check the system log for the kill, and correlate it with the worker count and with which pages were loaded at that moment.
  • Why is adding a retry the wrong first response to a parallel-only failure?
    Because it hides both causes without fixing either. A collision retried is still a collision, and it will now fail somewhere less obvious; host pressure retried adds another browser to an already saturated machine. A retry is defensible once you know which cause you have and have decided to live with the residual risk, never as the diagnosis.
  • Six workers each download the exhibit catalogue and assert on the newest file. Why is that unsafe even with plenty of memory?
    Because the browsers share one download directory, so newest-file is whichever worker finished last. Give each worker its own directory through the browser preferences - download.default_directory for Chrome and Edge, browser.download.folderList set to 2 with browser.download.dir for Firefox - and assert within that directory. Selenium's Grid does the same thing per session when managed downloads are enabled.

saying these in an interview costs you the question

  • Treats every parallel-only failure as flake and adds a retry
  • Assumes more memory will fix a collision between workers
  • Serialises the suite instead of making the shared resource unique
  • Cannot distinguish a session-start failure from a session lost mid-test
  • Never reproduces at one worker before changing the parallel configuration