skip to content

In Selenium Grid 4, what does the fully distributed topology buy that a single hub cannot?

level: seniorimportance: should knowfreq 46%

answer

  1. A hub is five roles in one
  2. Ask what you can restart alone
  3. Different roles run out of different things
  4. One of them can be swapped entirely
  5. Session state need not live in the JVM

basics

~20 s

Independent sizing, restart and replacement. A hub is one JVM holding the event bus, session map, queue, distributor and router together; split apart, each can be given its own machine, restarted alone, or swapped for a different implementation.

solid answer

~40 s

A `hub` is not a distinct piece of software. It is one JVM that constructs a bound event bus, a `LocalSessionMap`, a `LocalNewSessionQueue`, a `LocalDistributor` and a `Router` wired to all three, so five of the six roles share one process, one heap and one restart. Splitting them changes nothing about the roles and everything about their lifecycles: the Distributor, which creates sessions on a thread pool sized from its machine's processor count, can be given its own host; the Session Map can be pointed at a JDBC or Redis backed implementation so session state survives a restart; and any one role can be reconfigured without dropping the Grid. The cost is six processes and a wiring surface that fails silently when it is wrong.

go deeper

for a junior

Know that a hub packs several roles into one process and that Grid 4 lets you run those roles separately. You will not be asked to design the split yourself.

for a middle

Say which roles a hub bundles and what each becomes as its own process, including which one holds session state in memory and therefore loses it on restart.

for a senior

Show the diagnosis before the change: name the saturated component, and explain what a JDBC or Redis backed session map actually buys a long-running grid.

for a principal

Own the tradeoff between six operable processes and one. State when the operational surface is not worth it, and what evidence would change your mind.

## What a hub actually is The clearest way to answer this is to say what a `hub` really contains, because most candidates treat it as a black box. In Selenium 4 the `hub` subcommand starts one JVM that constructs, in order: an event bus with `bind` defaulting to `true`, a `LocalSessionMap`, a `LocalNewSessionQueue`, a `LocalDistributor`, and a `Router` wired to all three. That is five of the six roles in one process, listening on `4444` for clients and on `4442`/`4443` for its Nodes. The only role a hub does not hold is the **Node** itself. Fully distributed mode does not add anything to that list. It takes the same five objects and gives each its own process, its own port and its own lifecycle. ## The three things the split actually buys 1. **Independent sizing.** The five roles are limited by completely different resources, so putting them in one JVM means sizing the machine for the sum of them and tuning nothing individually. 2. **Independent restart.** A configuration change to one role no longer takes the whole Grid down. Restarting the Router does not evict Node registrations; restarting the Distributor does not lose the session map. 3. **Independent replacement.** One role has real drop-in alternatives, and you cannot use them without splitting: the Session Map's implementation is selectable with `--sessions-implementation`, and the tree ships `JdbcBackedSessionMap` and `RedisBackedSessionMap` alongside the in-memory `LocalSessionMap` and a `NullSessionMap`. ## Where the pressure actually lands | Component | What limits it | The knob, or the replacement | |---|---|---| | Router | inbound client connections and forwarded command traffic | stateless; it holds no Grid state of its own | | Distributor | CPU while creating sessions, one worker per concurrent creation | `--newsession-threadpool-size`, defaulted from the machine's processor count | | Session Map | one lookup per forwarded command, plus durability of the mapping | `--sessions-implementation` pointing at a JDBC or Redis backed map | | New Session Queue | how long requests wait against how many slots exist | its own request timeout and retry settings | | Event Bus | message fan-out as the Node count grows | `--eventbus-heartbeat-period`, and giving the bus its own host | | Node | the browsers the machine can actually run | Node configuration, which is a subject of its own | The row that matters most in a real interview is the Distributor. Session creation runs on a fixed thread pool sized from `availableProcessors()` on the **Distributor's** machine, so a Grid that is slow to hand out car-rental pickup-form sessions while Node slots sit free is usually a Distributor that is out of threads, not a Grid that is out of browsers. Inside a hub that pool competes with the Router's request handling on the same box; split out, you can give it a machine of its own. ## What the split costs - **Six processes instead of two.** Six sets of flags, six log streams, six things to health-check. - **A wiring surface that can be silently wrong.** Every bus member needs the identical `--publish-events`/`--subscribe-events` pair, the Distributor needs `--bind-bus false`, the Router needs three URLs, and a mistake in any of them produces a Grid that starts cleanly and never runs anything. - **A shared `--registration-secret`**, which now has to be distributed to six places rather than two. - **More network exposure.** Ports `4442`, `4443`, `5553`, `5556`, `5557` and `5559` all have to be reachable between components, and none of them should be reachable from where the tests run. ## Deciding for a car-rental pickup-form suite Split when you can name the component that is saturated, and not before. Selenium's own sizing guidance points at Hub/Node for grids in the tens of Nodes and reserves the distributed shape for grids past roughly a hundred, precisely because the wiring cost is fixed while the benefit scales with size. - If the suite queues because every Chrome slot is busy, add Nodes. Nothing about the topology helps. - If sessions are handed out slowly while slots are free, the Distributor is the candidate, and it is the first role worth its own machine. - If losing the hub mid-run orphans browsers and strands session ids, moving the Session Map to a JDBC or Redis backed store is the change that actually helps, and it requires a separate `sessions` process to host it. - If none of those describe your Grid, a hub is the right answer and six processes is operational cost with no return.

  • Which single component would you move off the hub first?
    The Session Map, because it is the one with real drop-in alternatives. Pointing its implementation at a JDBC or Redis backed map moves session-to-Node state out of the JVM entirely, so restarting the Router or the Distributor no longer strands the session ids of car-rental pickup-form runs already in flight.
  • What does the split not fix?
    Browser capacity. Throughput is still bounded by the slots Nodes offer, so a car-rental suite that waits because every Chrome slot is busy gains nothing from six processes. That is a Node count and Node configuration question, and splitting the control plane will not move it.
  • What must every component share once they are split?
    The bus members need the identical publish and subscribe address pair, and every component needs the same registration secret where one is set, because that secret guards the HTTP calls components make to each other. A mismatch leaves processes healthy on their own and rejecting one another.

saying these in an interview costs you the question

  • Says distributed mode makes browsers start faster
  • Thinks splitting the roles adds browser capacity
  • Calls the hub just a router under another name
  • Assumes the session map survives a hub restart
  • Recommends six processes for a five-node grid