skip to content

In Selenium 4's Grid, how does a node register with the hub and stay registered?

level: middleimportance: should knowfreq 56%

answer

  1. The node speaks first, not the hub
  2. It keeps saying that it is alive
  3. A bounded retry window, then it stops
  4. Status event, node-added reply, then heartbeats
  5. A shutdown hook announces the departure

basics

~20 s

The node publishes a status event onto the hub's event bus and repeats it until the hub answers, then sends a heartbeat on a fixed period. A clean stop publishes a removal event so the hub drops it.

solid answer

~40 s

In Selenium 4 registration is event-driven, not an HTTP call to the hub. On start-up the node health-checks itself, then fires a node-status event onto the hub's event bus and retries every `--register-cycle` seconds (10 by default) for up to `--register-period` seconds (120) until it is accepted; `--register-shutdown-on-failure true` makes it exit rather than sit there unregistered. The hub's distributor picks that event up, builds a remote handle from the URI inside it, and answers with a node-added event carrying the node's id. From then on the node fires a heartbeat event every `--heartbeat-period` seconds, 60 by default, which refreshes the hub's last-seen time. A clean shutdown fires a node-removed event and the hub drops the node immediately; a killed node is detected instead through missing heartbeats.

go deeper

for a junior

Know that the node announces itself to the hub rather than the other way round, and that it keeps sending a periodic heartbeat so the hub can tell it is still alive.

for a middle

Be ready to walk the sequence: self health check, status event on the bus, node-added reply, then heartbeats. Name the register-cycle, register-period and heartbeat-period defaults and say what each one bounds.

for a senior

An interviewer expects you to reason about start-up races in an orchestrated fleet: nodes coming up before the hub, the register window expiring silently, and when shutting down on registration failure beats an unbounded retry.

for a principal

Own how fast the fleet's picture of itself converges. The heartbeat period, health-check interval and purge interval interact, and choosing them badly costs you either stale slots handed real work or a constant probe load.

## Registration is a conversation, not a call In **Selenium 4** a node does not POST itself to an endpoint on the hub, and the hub does not scan the network looking for nodes. Membership is entirely event-driven over the hub's **event bus**, which listens on `4442` for published messages and `4443` for subscriptions. The exchange runs like this: 1. The node starts its own HTTP server and runs a **self health check**. If it reports itself down at that moment, it does not announce anything yet and tries again on the next cycle. 2. It publishes a **node-status event** onto the bus, carrying its id, the URI it is reachable at, and the slots it offers. 3. The hub's **distributor** receives that event, builds a remote handle for the node from the URI in it, and publishes a **node-added event** naming the node's id. 4. The node sees its own id in that node-added event and marks itself registered, which stops the retry loop. For a fleet of nodes running an **energy-tariff switcher** suite, this is why nodes can be started in any order relative to the hub: a node that starts first simply keeps announcing until something is listening. ## The start-up retry window The retry loop is bounded, and two flags define it: - `--register-cycle` — how often, in seconds, the node re-announces itself while it is still unregistered. Default **10**. - `--register-period` — how long, in seconds, it keeps trying at all. Default **120**. When this expires the node logs that registration failed and **stops trying** — it does not retry forever. A node that has burned its window sits there running, serving nothing, in a grid that has never heard of it. That is the wrong shape for an orchestrated fleet, which is what the third flag is for: ```bash java -jar selenium-server.jar node \ --hub http://tariff-grid:4444 \ --register-period 60 \ --register-shutdown-on-failure true ``` With `--register-shutdown-on-failure true` the node exits instead of lingering, so whatever supervises the process can restart it and let it try a fresh window. ## Heartbeats and what they refresh Once registered, the node publishes a **heartbeat event** carrying its current status every `--heartbeat-period` seconds, **60** by default. On the hub side this is deliberately forgiving: - If the node id is already known, the heartbeat **refreshes** the hub's last-seen time for it and, if the hub had it marked down while the node reports itself up, restores it. - If the node id is **not** known, the hub treats the heartbeat as a registration and adds the node. That is why a hub restarted underneath a running fleet repopulates itself within about one heartbeat period, with no action on the nodes. Heartbeats carry status only. Session commands never travel on the bus; they go hub-to-node over HTTP. ## Leaving the grid: the clean path and the unclean one | | Clean stop (the process is asked to exit) | Unclean loss (killed, or the network drops) | |---|---|---| | What the node does | a shutdown hook fires a **node-removed event** | nothing; no event is ever published | | How the hub learns | immediately, from that event | from missing heartbeats and its own failing probe | | Time to disappear | effectively at once | tens of seconds to a few minutes | | Sessions on it | already finished or abandoned by the exit | held in the hub's picture until the node is removed | The unclean path is governed by heartbeat freshness on a sweep that runs every `--purge-nodes-interval` seconds, **30** by default. A node last heard from longer ago than twice its heartbeat period is switched from up to **down**; one that has been down for four times its heartbeat period is removed and a node-removed event is fired for it. With the defaults that is roughly two minutes to down and four to gone. Setting the purge interval to zero disables the sweep entirely, so dead nodes stay listed. A third, deliberate path exists as well: **draining**. A drained node stops accepting new sessions, finishes the ones it has, then publishes a completion event and exits — so it leaves the grid without abandoning a tariff-switch test mid-run. ## Choosing the periods - Shortening `--heartbeat-period` makes the hub notice a lost node sooner, because both purge thresholds are multiples of it — at the cost of more bus traffic per node. - Shortening `--register-cycle` speeds up the join for nodes racing a slow hub start. - Lengthening `--register-period` helps only if the hub genuinely takes that long to come up; otherwise it just delays the failure signal. - Leaving `--purge-nodes-interval` at zero keeps every node that ever registered in the hub's picture, including ones that will never answer again.

  • What happens when a heartbeat arrives for a node the hub has never seen?
    The hub treats it as a registration. Its listener checks whether the node id is already known: if it is, the heartbeat just refreshes the record, and if it is not, that same heartbeat is handled as a fresh registration. This is why a hub restarted underneath a running fleet repopulates itself within about one heartbeat period, with nothing done on the nodes.
  • A node was killed rather than stopped cleanly. How long before the hub notices?
    No removal event is fired, so the hub falls back on heartbeat freshness. The purge sweep runs every `--purge-nodes-interval` seconds, 30 by default: a node not heard from for twice its heartbeat period is switched to down, and one down for four times that period is removed. With the defaults that is roughly two minutes to down and four to gone.

saying these in an interview costs you the question

  • Saying the hub discovers nodes by scanning, rather than nodes announcing themselves
  • Claiming a node registers by POSTing to an HTTP endpoint on the hub
  • Thinking heartbeat events carry session commands rather than status
  • Assuming a node that cannot register keeps retrying forever
  • Expecting a killed node to de-register itself the way a stopped one does