skip to content

How do you decide between many single-session Selenium node containers and fewer multi-session ones?

level: principalimportance: should knowfreq 36%

answer

  1. Start from what one browser needs
  2. Isolation and overhead pull opposite ways
  3. Shared memory and CPU are per container
  4. One recorder per container, not per session
  5. Scale out before overriding the ceiling

basics

~20 s

Size the container to one browser first. One session per container gives clean isolation, its own shared memory and one video per run; packing sessions in saves scheduling overhead but shares CPU, /dev/shm and the blast radius of a crash.

solid answer

~40 s

Selenium 4's node images default to one session per container, and that default is the right starting shape: each `selenium/node-chrome` gets its own `/dev/shm`, its own CPU allocation and its own recorder, so a crash on the loan-renewal page costs one session. Packing more in means setting `SE_NODE_MAX_SESSIONS` together with `SE_NODE_OVERRIDE_MAX_SESSIONS=true`, and it buys fewer containers to schedule and less per-container overhead — at the cost of a shared `--shm-size` budget, shared CPU, and a `selenium/video` sidecar that can capture several sessions into one file. I scale out rather than up unless container churn is a measured bottleneck, and I treat `SE_DRAIN_AFTER_SESSION_COUNT` as a separate, state-hygiene decision about how long a container's browser state may accumulate.

go deeper

for a junior

Know that Selenium's browser node images run one session per container by default, and that a second session in the same container is a deliberate configuration change rather than something the grid arranges for you.

for a middle

Explain what a second session in the same container actually shares: the CPU allocation, the /dev/shm you sized, and the display a recorder attached to that container is watching.

for a senior

Demonstrate the diagnosis behind the choice. Say which containers were CPU-starved, whether crashes cluster on the most loaded ones, and that you would change one dimension at a time rather than resizing the whole fleet.

for a principal

Own the policy. Pick a default container shape for the organisation, state what evidence would move it, and be explicit that overriding the per-container session ceiling is a measured exception rather than a growth strategy.

## The default shape, and why it is the default Selenium 4's browser node images set `SE_NODE_MAX_SESSIONS="1"`, so out of the box **one container serves one session**. That default is not timidity — it is the shape the rest of the image is built around. Each container gets its own `/dev/shm` allocation, its own CPU allocation, its own virtual display sized by `SE_SCREEN_WIDTH` and `SE_SCREEN_HEIGHT`, and, if you attach one, its own `selenium/video` recorder watching that display. Every one of those resources is drawn at the container boundary, which means the container boundary is also your isolation boundary. Packing more sessions into a container is a deliberate departure. It requires `SE_NODE_MAX_SESSIONS` raised **and** `SE_NODE_OVERRIDE_MAX_SESSIONS=true`, because the Node otherwise clamps the request to the processor count it detects. Selenium logs a warning when you take that route. ## What the two shapes actually trade | Dimension | Many single-session containers | Fewer multi-session containers | |---|---|---| | `/dev/shm` | One allocation per browser | One allocation divided among browsers | | CPU | Sized per session, easy to reason about | Shared; contention appears as slow pages | | Failure | A dead container costs one session | A dead container costs every session in it | | Recording | One `selenium/video` sidecar per session | One recorder may capture several sessions in one file | | Overhead | More containers to schedule, pull and start | Fewer images to place, less per-container overhead | | Recycling | `SE_DRAIN_AFTER_SESSION_COUNT` retires a container after one run | The same setting retires several sessions' worth of state at once | Read that table as a single sentence: scaling **out** buys isolation and predictability; scaling **up** buys scheduling economy. Neither is free, and the loan-renewal suite does not care which you pick — the operators of the fleet do. ## Deciding, in order 1. **Start from what one browser needs.** Whole processors, a shared-memory allocation you have actually sized, enough RAM. If a single-session container is not comfortable, a multi-session one will not be. 2. **Default to scaling out.** More single-session containers is the shape the images assume and the shape that keeps a crash local. 3. **Measure before you pack.** The case for multi-session containers is container churn — image pulls, start-up time, scheduling pressure — not a desire for a bigger number. If churn is not measurably hurting you, it is not the problem. 4. **If you pack, pack modestly.** Stay at or below one session per processor unless you have evidence, and raise the shared-memory allocation in the same change. 5. **Decide recycling separately.** `SE_DRAIN_AFTER_SESSION_COUNT` makes a node drain and detach after N sessions so a fresh container replaces it; that is a state-hygiene decision, not a capacity one. ## Levers that only exist at the container layer - `SE_NODE_MAX_SESSIONS` with `SE_NODE_OVERRIDE_MAX_SESSIONS` — the only way a node advertises more slots than it detects processors. - `--shm-size` or `shm_size`, and `dshmVolumeSizeLimit` in Selenium's `selenium-grid` Helm chart — the shared memory the browsers in that container divide. - `SE_DRAIN_AFTER_SESSION_COUNT` — how long a container's browser state is allowed to accumulate before the container is replaced. - `SE_SCREEN_WIDTH` and `SE_SCREEN_HEIGHT` — the display every session in that container renders into, so a mixed-viewport fleet is a container-shape decision, not a per-session one. - `SE_EVENT_BUS_HOST`, with `SE_EVENT_BUS_PUBLISH_PORT` and `SE_EVENT_BUS_SUBSCRIBE_PORT` defaulting to `4442` and `4443` — what makes a node container part of the grid at all, whatever shape you chose. ## Moving the same shape from compose to Kubernetes The images and the `SE_` variables are identical in both places; that portability is the point of configuring nodes this way. What changes is how the two non-variable pieces are expressed: - **Shared memory.** Compose takes `shm_size` on the service. Selenium's Helm chart takes a per-node `dshmVolumeSizeLimit` and mounts a memory-backed volume at `/dev/shm`; leave it empty and the chart instead injects `SE_BROWSER_ARGS_DISABLE_DSHM` so the browser avoids shared memory. - **Replication.** Compose gives you one service per node shape; the chart gives you one deployment per browser. In both, "more capacity" means more containers of your chosen shape, which is exactly why the shape is worth deciding once. A principal-level answer names the default, gives the evidence that would move it, and is explicit that overriding the per-container session ceiling is a measured exception rather than a growth strategy.

  • What makes you recycle node containers instead of leaving them up?
    State that leaks between sessions: profile leftovers, downloaded files, a browser process that survived a bad teardown. `SE_DRAIN_AFTER_SESSION_COUNT` makes a node drain and detach after a set number of sessions so a fresh container replaces it, trading a little start-up cost for a clean browser every N runs.
  • How does the shape change when you move from compose to Kubernetes?
    The images and the `SE_` variables do not change — `SE_EVENT_BUS_HOST` still points a node at the bus, `SE_NODE_MAX_SESSIONS` still sets its slots. What changes is how shared memory is expressed: compose takes `shm_size` on the service, while Selenium's `selenium-grid` Helm chart takes a per-node `dshmVolumeSizeLimit` and mounts a memory-backed volume at `/dev/shm`.
  • How do you size a node container's CPU allocation?
    Start from one browser session per processor, which is the ratio Selenium's Node assumes when it clamps its session count to the processors it detects. Give whole processors rather than fractions: a browser starved of CPU surfaces as slow page loads and timeouts on the loan-renewal form, not as an obvious capacity signal.

saying these in an interview costs you the question

  • Packs many sessions into one container to make the fleet look bigger
  • Assumes containers are free, so extra sessions cost nothing
  • Ignores that sessions in one container share /dev/shm and CPU
  • Picks a container shape from a target number without measuring anything
  • Keeps node containers running forever and blames flake on the tests