skip to content

Grid Deployment

Selenium Grid 4 as software you run yourself: the roles that make up a grid, how it declares capacity, and how a remote client reaches it. Interviewers probe where a shared grid breaks.

on this pageshow

explore

questions

28

In Selenium's Docker images, why is --shm-size=2g recommended for browser containers?

level: juniorimportance: must knowfreq 66%

answer

  1. The browser needs room the container withholds
  2. Look at a shared-memory filesystem
  3. The browser crashes rather than slowing down
  4. A docker run flag sized in gigabytes
  5. shm_size: 2gb in compose

basics

~20 s

Browsers put rendered content in the shared-memory filesystem /dev/shm, and Selenium's image docs say Chrome crashes on the small 64 MB allocation a container gets by default. Passing --shm-size=2g, or shm_size: 2gb in compose, gives it room.

solid answer

~40 s

Selenium's browser images run a real Chrome, Edge or Firefox inside the container, and those browsers pass rendered content between their processes through `/dev/shm`. Selenium's own image documentation states that Chrome will crash on the default 64 MB `/dev/shm`, which is why every `docker run` example for `selenium/standalone-chrome` and `selenium/node-chrome` carries `--shm-size="2g"` and every browser service in the project's compose files sets `shm_size: 2gb`. Selenium calls 2 GB arbitrary but known to work well and recommends tuning it — the all-browsers images are documented at `3g`. The alternative lever the images ship is the `SE_BROWSER_ARGS_` prefix, which appends a browser argument such as `--disable-dev-shm-usage`; Selenium's own Helm chart injects exactly that when a node has no shared-memory volume size configured.

code

bash · 4 lines
bash
docker run -d --name loan-renewal-browser \
  -p 4444:4444 -p 7900:7900 \
  --shm-size="2g" \
  selenium/standalone-chrome:4.48.0-20260905

go deeper

for a junior

Be ready to name the symptom and the fix: browsers inside Selenium's images crash when the container's /dev/shm is too small, and --shm-size=2g, or shm_size: 2gb in a compose file, is the documented remedy.

for a middle

Explain why a multi-process browser needs shared memory at all, why a container's allocation is small, and what the second lever is — the SE_BROWSER_ARGS_ prefix that appends --disable-dev-shm-usage so the browser writes to disk instead.

for a senior

Show that you would recognise this from the outside: sessions dying mid-run on the heaviest pages, worse on the most loaded containers. Then size shared memory per browser rather than copying 2 GB everywhere in the fleet.

for a principal

Own it as fleet policy. Decide whether shared memory is sized in the container specification or worked around with a browser argument, apply one choice everywhere, and make sure nobody on the team debugs the same crash twice.

## What `/dev/shm` is doing for a browser `/dev/shm` is a shared-memory filesystem: a directory backed by RAM that separate processes can use to hand each other large blocks of data without copying them through a pipe. Modern browsers are multi-process by design — a browser process, one or more renderer processes, a GPU process — and they use exactly that filesystem to pass rendered content between those processes. On a desktop the allocation is generous and nobody thinks about it. Inside a container it is a small fixed size, and Selenium's own image documentation is blunt about the consequence: **Chrome will crash on the default 64 MB `/dev/shm`**. That is why every browser-image example the project publishes carries the size flag, and why the recommendation appears at the very top of the Quick Start rather than buried in a troubleshooting note. ## The symptom, and why it reads as flake - The session **starts cleanly**. The remote end accepts the new session, the client gets a driver back, and the first navigations work. - The browser dies **part-way through a page**, so the client sees the browser exiting or the connection dropping mid-command, not a capacity error. - It hits **heavier pages first**. A loan-renewal page with a long borrowed-items table, an embedded date picker and a cover-image grid is far more likely to trip it than the catalogue search that precedes it. - It looks **intermittent**, because whether the allocation runs out depends on what the page happened to render, so the first instinct is usually to add a wait rather than to look at the container. Selenium's FAQ describes the size flag as a known workaround for the browser crashing inside a container, and links the upstream Chrome and Firefox reports that document it. Treating this as a test-code problem is the classic wrong turn. ## The documented fix, in each place a container starts | Where you start it | What you set | Example value | |---|---|---| | `docker run` | `--shm-size` | `--shm-size="2g"` | | The project's compose files | `shm_size` on each browser service | `shm_size: 2gb` | | Selenium's `selenium-grid` Helm chart | `dshmVolumeSizeLimit` on a node | `"2Gi"`, mounted as a memory volume at `/dev/shm` | The value is not magic. Selenium calls 2 GB **arbitrary but known to work well** and recommends tuning it to your use case — and applies the same reasoning upward, documenting `--shm-size="3g"` for the `selenium/standalone-all-browsers` and `selenium/node-all-browsers` images because one container is holding more browsers. ## The other lever: keeping the browser off shared memory Chromium-family browsers accept an argument that makes them write to a temporary directory on disk instead of `/dev/shm`. Selenium's Chrome, Chromium and Edge node images expose a way to append arbitrary browser arguments: **any environment variable whose name begins with `SE_BROWSER_ARGS_`** has its value appended to the browser's command line, with the suffix chosen by you. The project's own Helm chart uses this as the fallback path — when a node has no `dshmVolumeSizeLimit` configured, the chart injects `SE_BROWSER_ARGS_DISABLE_DSHM` carrying `--disable-dev-shm-usage` so the browser never depends on the small allocation in the first place. The two levers are alternatives, and worth being deliberate about: 1. **Enlarge `/dev/shm`** when you want the browser running the way it does on a desktop and can afford the RAM. 2. **Move the browser off `/dev/shm`** when you cannot control the container's shared-memory size — the browser then pays disk I/O instead of memory. ## Sizing it for a real fleet - Size **per browser**, not per fleet. A container holding one Chrome and a container holding three need different numbers, and copying 2 GB everywhere is how a fleet ends up both wasteful and still crashing. - Remember the allocation is **shared inside the container**. Raising the sessions a node offers without raising this number moves the crash rather than fixing it. - Firefox images get the same treatment in Selenium's own compose files and `docker run` examples, so do not treat this as a Chrome-only ritual. - Make it **uniform**. Whichever lever you pick, apply it in one place for every browser container you run, so nobody on the team debugs the same crash twice. The short version to carry into an interview: browsers put rendered content in shared memory, containers hand out very little of it, Selenium's images document the crash and the fix, and the fix lives in the container specification — `--shm-size`, `shm_size`, or the chart's `dshmVolumeSizeLimit` — not in your test code.

  • What does the failure actually look like from the test's side?
    Not a clean error. The session starts, the first navigations work, and then the browser dies part-way through a page, so the client reports the browser exiting or the connection dropping mid-command. It shows up on heavier pages first, which makes it look like an intermittent test rather than a container that was sized wrong.
  • Does Firefox need the same treatment as Chrome?
    Selenium's own compose files set `shm_size: 2gb` on the Firefox node service and its `docker run` examples pass `--shm-size="2g"` for `selenium/standalone-firefox`, so the project applies it to Firefox too. The upstream crash reports the FAQ links cover both browsers, so this is not a Chrome-only ritual.
  • Is 2 GB a required value?
    No. Selenium's documentation calls 2 GB arbitrary but known to work well and recommends tuning it to your use case. The same reasoning is applied upward: the `selenium/standalone-all-browsers` and `selenium/node-all-browsers` images are documented with `--shm-size="3g"` because one container holds more browsers.

Shared memory is the browser's scratch bench. On a desktop it is a full workshop table, but a container hands the browser a bench barely wide enough for one page, so it runs out of room mid-render and falls over rather than tidying up.

saying these in an interview costs you the question

  • Says the browser only slows down when shared memory runs short
  • Looks for an SE_ environment variable that sets the shared-memory size
  • Blames flaky waits when the browser is dying from shared memory
  • Assumes headless mode removes the browser's need for shared memory
  • Treats 2 GB as a fixed requirement rather than a tunable starting point
open as a page

In Selenium Grid, what is the difference between the hub role and the node role?

level: juniorimportance: must knowfreq 82%

basics

~20 s

The hub is the grid's single entry point: it takes session requests from test clients and hosts the event bus that nodes register on. A node owns a machine's browsers and runs the sessions the hub places on it.

open as a page

In Selenium Grid 4, what does the /status endpoint return and what does its ready flag mean?

level: juniorimportance: must knowfreq 66%

basics

~20 s

Selenium Grid's /status returns JSON holding a boolean ready flag, a message, and a nodes array describing every registered node and its slots. Ready is true only when at least one node is up and has a free slot.

open as a page

In Selenium 4, how do you point a test at a Grid endpoint instead of starting a local browser?

level: juniorimportance: must knowfreq 76%

basics

~20 s

Build a RemoteWebDriver with two arguments: the Grid's URL and an options object such as ChromeOptions. Those options become the session request the Grid routes, and every line of test code after that is unchanged.

open as a page

In Selenium Grid, what does a Node's --max-sessions flag default to, and what silently caps it?

level: middleimportance: must knowfreq 60%

basics

~20 s

A Selenium 4 Node's --max-sessions defaults to the number of processors its JVM reports. Selenium clamps any higher value back to that number unless --override-max-sessions true is passed, and a Node never runs more sessions than it declared slots.

open as a page

In Selenium 4, what does Grid's standalone command run inside its single process?

level: middleimportance: must knowfreq 64%

basics

~20 s

All of it. Selenium 4's standalone command builds the router, distributor, session map, new session queue, event bus and node in one JVM, serving WebDriver on port 4444 at both the root path and /wd/hub.

open as a page

How do you tell an unreachable Selenium Grid apart from a Grid with no slot matching your request?

level: seniorimportance: must knowfreq 62%

basics

~20 s

Ask who produced the error. A transport failure has a network cause such as connection refused or unknown host and no WebDriver error body; a Grid that answered returns a session-not-created error whose message echoes the capabilities you asked for.

open as a page

When is Selenium Grid's standalone mode the right topology for a train timetable board test suite?

level: principalimportance: must knowfreq 46%

basics

~20 s

Standalone is right whenever one machine can hold the whole run. Selenium 4's standalone command starts every Grid component in one process on port 4444, so a laptop, a CI job or a container needs nothing else.

open as a page

In Selenium 4, what port does Grid's standalone server listen on by default?

level: juniorimportance: should knowfreq 74%

basics

~10 s

Port 4444. A Selenium 4 server started with the standalone command serves WebDriver on http://localhost:4444, at both the root path and /wd/hub, and the -p or --port flag moves it.

open as a page

Why can raising SE_NODE_MAX_SESSIONS on a Selenium node container leave its concurrency unchanged?

level: middleimportance: should knowfreq 42%

basics

~20 s

Selenium's Node caps its session count at the processors it detects, so a larger SE_NODE_MAX_SESSIONS is silently reduced. The higher number applies only when SE_NODE_OVERRIDE_MAX_SESSIONS is also true and the container really has that CPU.

open as a page

In Selenium 4's Grid, how does a node register with the hub and stay registered?

level: middleimportance: should knowfreq 56%

basics

~20 s

The node publishes a status event onto the hub's event bus and repeats it until the hub answers, then sends a heartbeat on a fixed period. A clean stop publishes a removal event so the hub drops it.

open as a page

How does a client query Selenium Grid's /graphql endpoint to read live session and queue state?

level: middleimportance: should knowfreq 44%

basics

~10 s

Selenium 4's Grid exposes a read-only GraphQL API at /graphql. You POST a JSON body holding a query string, and the schema offers four roots: grid, nodesInfo, sessionsInfo and session by id.

open as a page

In Selenium Grid, what do the --session-request-timeout and --session-retry-interval flags each control?

level: middleimportance: should knowfreq 47%

basics

~20 s

One bounds the wait, the other paces the poll. In Selenium 4, --session-request-timeout is how many seconds a request may sit in Grid's queue before failing, default 300. --session-retry-interval is how often in milliseconds the Distributor rechecks, default 15.

open as a page

In Selenium Grid, what does the se: capability namespace hold, and what do se:name and se:recordVideo do?

level: middleimportance: should knowfreq 38%

basics

~20 s

se: is Selenium's own capability prefix, read by the Grid and its nodes rather than by the browser. se:name labels the session in the Grid console and in any recording's file name; se:recordVideo asks a node that can record to capture that session.

open as a page

In a fully distributed Selenium Grid 4, why is the distributor started with --bind-bus false?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Because the distributor defaults to binding the event bus itself, the way a hub does. When a separate event-bus process already owns ports 4442 and 4443, --bind-bus false makes the distributor join that bus rather than start a second.

open as a page

Which Selenium Grid 4 components talk only over HTTP and never join the event bus?

level: seniorimportance: should knowfreq 35%

basics

~20 s

The Router and the New Session Queue. Neither declares the event bus role, so neither accepts the bus flags. The Distributor, the Session Map and the Node are the bus members, alongside the event-bus process itself.

open as a page

In Selenium Grid 4, what does the fully distributed topology buy that a single hub cannot?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Independent sizing, restart and replacement. A hub is one JVM holding the event bus, session map, queue, distributor and router together; split apart, each can be given its own machine, restarted alone, or swapped for a different implementation.

open as a page

A Selenium Grid node registers with the hub, then is marked down minutes later. Why?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Almost always the node advertised an address the hub cannot reach. Registration travels node to hub over the event bus, but the hub then calls the node back over HTTP, and it is that second direction that fails.

open as a page

A Selenium Grid Node runs Chrome, yet requests pinning browserVersion never match its slot. Why, and what fixes it?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The Node auto-detected Chrome, so its stereotype declares only a browser name and a platform, never a version. Selenium 4 treats an undeclared version as no match. Declare the version explicitly in a driver-configuration stereotype instead.

open as a page

A Selenium Grid request stays queued: how do you tell a busy slot from a missing stereotype?

level: seniorimportance: should knowfreq 54%

basics

~20 s

Compare what was asked for with what is advertised. Read the pending capabilities from Selenium Grid's GraphQL sessionQueueRequests, then list every node's stereotypes and occupancy: a full match means saturation, no match means the request can never be served.

open as a page

When a queued Selenium Grid session request hits its timeout, what does the client actually receive?

level: seniorimportance: should knowfreq 41%

basics

~20 s

An HTTP 500 carrying the WebDriver error session not created, which the Java client raises as SessionNotCreatedException. With defaults, though, the client's own 180-second read timeout fires before Grid's 300-second queue timeout, so a transport error appears instead.

open as a page

A Selenium Grid suite stopped getting sessions after browserVersion was pinned in the client's options. Why?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Everything the client sends that the Grid matches on narrows the set of slots that may serve it. A pinned version string no node advertises can never match, so the request becomes unroutable and comes back as a session-not-created failure.

open as a page

How do you decide between many single-session Selenium node containers and fewer multi-session ones?

level: principalimportance: should knowfreq 36%

basics

~20 s

Size the container to one browser first. One session per container gives clean isolation, its own shared memory and one video per run; packing sessions in saves scheduling overhead but shares CPU, /dev/shm and the blast radius of a crash.

open as a page

How would you size a Selenium Grid Node's slot layout for a mixed-browser suite, and what does over-declaring cost?

level: principalimportance: should knowfreq 35%

basics

~20 s

Start from one browser session per processor, declare one stereotype block per browser and version the suite must reach, and hold the concurrent ceiling at the core count. Extra slots widen what a Node serves, never how fast it serves.

open as a page

What are the six components of a fully distributed Selenium Grid 4 deployment?

level: juniorimportance: nice to knowfreq 34%

basics

~20 s

Router, Distributor, Session Map, New Session Queue, Event Bus and Node. In Selenium 4's fully distributed topology each is a separate process with its own subcommand and port, so each can be sized and restarted on its own.

open as a page

In Selenium Grid, what is a Node's stereotype and which session requests will match that slot?

level: juniorimportance: nice to knowfreq 30%

basics

~20 s

A stereotype is the fixed set of capabilities one Node slot advertises, such as browser name chrome and browser version 131. Grid gives a request that slot only when the request's capabilities match the declared values.

open as a page

In Selenium Grid 4, why does the event bus bind two separate ports, 4442 and 4443?

level: middleimportance: nice to knowfreq 31%

basics

~20 s

The Grid 4 event bus is a ZeroMQ proxy with two ends. Port 4442 is the socket the bus publishes from and every component subscribes to; port 4443 is the socket the bus subscribes on and every component publishes to.

open as a page

When sizing a Selenium Grid, why is new-session queue depth a better capacity signal than Node count?

level: principalimportance: nice to knowfreq 27%

basics

~20 s

Node count is a static ceiling; queue depth measures whether that ceiling is ever reached. A grid with sixty idle slots is no busier than one with six, so depth is what tells you when to add capacity.

open as a page