skip to content

How do you choose how many parallel Selenium browser workers a single machine should run?

level: principalimportance: must knowfreq 50%

answer

  1. Count browsers, not threads
  2. Headless is still a whole browser
  3. A baseline before a ramp
  4. Watch per-case time, not just total
  5. The number expires when the machine does

basics

~20 s

Measure it on the real hardware; there is no formula. Each worker runs a whole browser process tree, so memory and CPU cap the count long before threads do. Ramp the workers up until per-case time degrades, then back off.

solid answer

~50 s

Start from what a worker actually costs: a browser is several operating-system processes plus a driver process, and headless changes what is painted, not what is spawned. So the ceiling on one host is memory and CPU, not thread count. Take a single-worker baseline of median and p95 per-case duration plus peak resident memory for one case, then ramp - two, four, six - recording the same numbers and the failure count at each step. The ceiling is the last count where p95 is still near baseline; the tell that you passed it is per-case duration rising while total throughput stays flat. Treat the result as a property of that machine and that suite, re-measured when either changes. Beyond that ceiling the real choice is to move the browsers off the test host with `RemoteWebDriver` at a grid endpoint.

go deeper

for a junior

Know that each parallel worker starts a real browser, and that a machine can only hold so many of them. You are not expected to pick the number, only to understand what it is limited by.

for a middle

Explain why memory and CPU bind before threads or ports do, and why headless does not make a browser cheap. Be able to describe what you would measure to find the limit.

for a senior

Show the measurement discipline: a single-worker baseline, a ramp, per-case p95 alongside failure counts, and the host-side evidence that a browser was killed rather than crashing.

for a principal

Own the decision and its expiry. Justify the number with evidence, state where the browsers should live and what a grid costs in latency and dependency, and be explicit about which limits belong to other owners.

## A worker is a browser, not a thread The instinct carried over from unit testing - set the worker count to the core count and move on - misses what a Selenium worker costs. Each one launches a **real browser**, and a modern browser is not one process: a parent process, one or more renderer processes, a GPU process and assorted utility processes, plus the driver process (`chromedriver`, `geckodriver`) that fronts it. Headless changes what is painted, not what is spawned; `--headless=new` still runs the same binary. Whatever number your runner is configured with, the machine is being asked for that many browser process trees at once. That is why the ceiling on a single host is set by **memory and CPU**, and it is reached long before you run out of threads, ports or file handles. ## What actually runs out first - **Memory.** Each browser's footprint scales with the pages it loads. A museum exhibit catalogue's gallery grid, full of high-resolution images, costs several times what a text-only form costs, so the same worker count that is comfortable on one suite thrashes on another. - **CPU during page load.** Layout, script execution and image decode are bursty. Workers do not consume evenly; they spike together whenever the suite's cases happen to align. - **Disk.** Every worker writes a temporary profile directory and, if the suite downloads catalogue exports, a download directory. Long runs accumulate both. - **Shared memory and file descriptors**, which the browser uses heavily and which containerised agents often constrain more tightly than the host would. The signal that you have crossed the line is not a crash. It is **per-case duration climbing** while total throughput stays flat, followed by browsers being killed mid-test. ## There is no formula - measure on your own hardware 1. Start with a hypothesis, not an answer: roughly one worker per available core, or fewer if the agent is a small container. 2. Record a baseline at one worker: **median and p95 per-case duration**, and peak resident memory of the whole browser process tree for a single case. 3. Ramp - two, four, six, eight - and at each step record the same numbers plus the failure count. 4. Stop at the last worker count where per-case p95 is still close to the single-worker baseline. That is your ceiling on that machine, for that suite. 5. Re-measure when the application changes materially, when the agent image changes, or when the suite starts exercising heavier pages. The number is a property of a machine and a suite, not of Selenium. Publishing the measured number and the date beside the pipeline configuration is what stops the next person from doubling it because a run felt slow. ## Local browsers versus RemoteWebDriver at a grid | | Browsers on the test host | `RemoteWebDriver` at a grid | |---|---|---| | What the test process holds | a full browser tree per worker | an HTTP client per worker | | What caps the worker count | that one machine's memory and CPU | capacity wherever the browsers run | | Startup cost | process launch only | process launch plus a queue wait | | Failure modes added | none | network hop, unreachable endpoint, no matching slot | | Browser and version choice | whatever the host has installed | whatever the remote end offers | The strategic point is that pointing `RemoteWebDriver` at a grid URL **moves the browser footprint off the machine running the tests**. The test process becomes cheap - it is holding sockets, not browsers - so the worker count stops being a property of the CI agent. What you buy in exchange is a network hop on every command, a queue in front of session creation, and a second system that can be down. Sizing and operating that remote capacity is a separate piece of work with its own owners; the decision here is only whether the browsers belong on this host. ## Signals you have gone one worker too far - Per-case p95 duration rises while the run's total wall-clock time stops improving. - The same cases fail in different runs, always late, never at startup. - Browser processes vanish without an error in the driver log, which means something killed them. - The suite is green on a developer laptop at the same worker count and red only on the agent. ## What this decision does not include Choosing the worker count is not the same as choosing **how the work is split across those workers**, nor how a long suite is divided so wall-clock time actually drops - those are separate concerns with separate owners. Nor does it cover the capacity of the system under test, which imposes its own limit independent of how many browsers you can afford to run. Own the number, own how it was measured, and be explicit that it expires.

  • Why is one worker per CPU core a starting hypothesis rather than an answer?
    Because it models the wrong resource. A browser worker is usually bounded by memory and by bursty page-load work, not by steady CPU occupancy, and its footprint scales with how heavy the pages are. The heuristic gives you somewhere to start ramping from; only measurement on that machine, with that suite, gives you the number.
  • What does pointing RemoteWebDriver at a grid actually change about the ceiling?
    It moves the browser footprint off the machine running the tests. The test process then holds an HTTP client per worker instead of a browser process tree, so the worker count stops being a property of the CI agent and becomes a property of wherever the browsers run. In exchange you take on a network hop per command, a queue in front of session creation, and a second system that can fail.
  • How would you keep the chosen worker count from silently going stale?
    Record the number, the machine it was measured on, the suite it was measured with, and the date, next to the pipeline configuration. Re-measure when the agent image changes, when the application starts serving materially heavier pages, or when the suite's shape changes. Without that, the next person doubles it because a run felt slow.

saying these in an interview costs you the question

  • Sets worker count to core count without measuring anything
  • Believes headless browsers are cheap enough that memory stops mattering
  • Judges the ceiling by total wall-clock time and ignores per-case duration
  • Raises the worker count in response to failures rather than lowering it
  • Treats a measured worker count as permanent across agents and suites