skip to content

A Selenium Grid node registers with the hub, then is marked down minutes later. Why?

level: seniorimportance: should knowfreq 48%

answer

  1. Registration succeeded; something else did not
  2. Two directions, only one was proven
  3. What address did the node report?
  4. The hub dials the advertised URI
  5. --host, --external-url and --bind-host

basics

~20 s

Almost always the node advertised an address the hub cannot reach. Registration travels node to hub over the event bus, but the hub then calls the node back over HTTP, and it is that second direction that fails.

solid answer

~50 s

Registration and health checking use opposite directions. The node publishes its status onto the hub's event bus, so it appears in the grid; the hub's distributor then builds a remote handle from the URI inside that status and health-checks it over HTTP every 120 seconds by default. If the node picked a URI that only means something on its own machine, the first probe fails and its availability flips to DOWN while heartbeats keep arriving. In Selenium 4 a node derives that URI from `--external-url` if set, otherwise the server `--host` value, otherwise the first non-loopback address it finds, and it logs the result as `Reporting self as:`. Fix it by telling the node what it is reachable as: `--host` for a routable name, `--external-url` when the reachable address is not derivable locally, `--bind-host false` when the name cannot be bound.

code

bash · 5 lines
bash
java -jar selenium-server.jar node \
  --hub http://tariff-grid:4444 \
  --host tariff-node-chrome.internal \
  --bind-host false \
  --port 5555

go deeper

for a junior

Know that a node has to tell the hub an address, and that the hub uses it to call back. If a node appears and then goes down, suspect that address before you suspect the browser.

for a middle

Explain how a node derives the URI it advertises, in order: external URL, then host, then the first non-loopback address. Say why the hub's periodic health check is what finally exposes a wrong one.

for a senior

An interviewer expects a diagnosis order: read the node's reported URI from its own log, try it from the hub's machine, then reach for --host, --external-url or --bind-host. Explain why a restart changes nothing.

for a principal

Own the convention that stops this recurring: every node in the fleet advertises a name resolvable from the hub, fixed once in the launch template the platform owns rather than rediscovered per host by each team.

## Registration and health checking run in opposite directions This is the failure the hub-and-node shape is famous for, and it is a direction problem rather than a browser problem. In **Selenium 4** the two halves of a node's relationship with the hub travel opposite ways: - **Registration and heartbeats go node to hub**, over the hub's **event bus** on ports `4442` and `4443`. The node opens those connections outbound. - **Health checks and session commands go hub to node**, over plain HTTP to the node's own port, `5555` by default. The hub opens that connection outbound. A node that appears in the grid has proved only the first direction. Nothing about a successful registration says the hub can get back. So when a node for the **energy-tariff switcher** suite registers cleanly at start-up and is marked down a minute or two later, the first hypothesis is always that it advertised an address the hub cannot reach. ## What the node advertises, and how it decides The node does not tell the hub "call me back at the socket this message came from". It computes a URI for itself and puts that in the status it publishes. The order it uses is: 1. `--external-url`, if you set it — used verbatim, whatever else is configured; 2. otherwise the server's `--host` value, combined with the node's port; 3. otherwise the first non-loopback address of the machine, discovered at start-up. The scheme is `https` when the server is configured with TLS and `http` otherwise. Step 3 is where the trouble starts: an address that is genuinely non-loopback on the node's own machine can still be meaningless from where the hub sits — a second interface, a segment the hub does not route to, or a name only that host resolves. The node prints its choice. Its start-up log carries a `Reporting self as:` line, and that URI is exactly what the hub will dial: ```text INFO [NodeServer.createHandlers] - Reporting self as: http://10.42.7.19:5555 ``` Two things follow from that line being computed rather than observed: - **The node never learns it is wrong.** Nothing on the node side tests the URI, so its own logs stay clean while the hub gives up on it. - **The value is fixed at start-up.** It is not renegotiated per session, so a session placed on that node is driven over the same unreachable address. ## Why it looks like a success first When the status event arrives, the hub's **distributor** builds a remote handle for the node from the URI in that status and logs `Added node <id> at <uri>. Health check every 120s`. The node is now in the grid, visible, apparently healthy — because nothing has tested the URI yet. Then the periodic probe runs. `--healthcheck-interval` defaults to **120 seconds**; each cycle the distributor makes an HTTP request to the advertised URI. An unreachable node yields an unreachable-node result and its availability flips to **DOWN**. Heartbeats keep arriving over the bus the whole time, which is why the node can look alive and unusable at once. ## A diagnosis order that settles it quickly 1. Read the node's own start-up log and note the URI on the `Reporting self as:` line. 2. From the **hub's** machine, not the node's, make a plain HTTP request to that exact URI. 3. If it fails there and succeeds from the node's own host, the address is the defect — stop looking at drivers, browsers and capabilities. 4. Decide which of the three address flags applies, restart the node with it, and re-read the `Reporting self as:` line to confirm the new value. ## The three flags, and which one to reach for | Flag | What it controls | Reach for it when | |---|---|---| | `--host` | the address the node both binds to and reports | the machine has several interfaces and the routable one has a usable name | | `--external-url` | the URI the node reports, regardless of what it bound | the reachable address is not one the node can derive locally | | `--bind-host false` | makes the host value reporting-only, not a bind target | the name you must advertise cannot be bound on that machine | `--port` is not on this list on purpose: it changes what the node binds and the port it reports, but it cannot fix a wrong host part. ## What does not fix it - **Restarting the node.** It will recompute the same address from the same inputs and fail the same way. - **Raising `--heartbeat-period`.** Heartbeats are already arriving; they are not the failing direction. - **Shortening `--healthcheck-interval`.** That only reaches the wrong verdict sooner. - **Adding capacity.** More slots on an unreachable node are still unreachable slots.

  • The node vanishes from the grid entirely rather than sitting there marked down. What changed?
    That is the purge sweep, which runs every `--purge-nodes-interval` seconds, 30 by default, and works off heartbeat freshness. A node not heard from for twice its heartbeat period is switched to down; one that has been down for four times that period is removed outright. Setting the interval to zero disables the sweep, so dead nodes stay listed indefinitely.
  • How would you prove the address, rather than the node itself, is the problem?
    Read the node's start-up log: it prints a `Reporting self as:` line carrying the exact URI it will advertise. Then make a plain HTTP request to that URI from the hub's machine. If it fails from there but succeeds from the node's own host, the advertised address is the defect and the browser stack is irrelevant.
  • Does raising the health-check interval make the problem go away?
    No. `--healthcheck-interval` only sets how often the distributor probes; a longer interval delays the verdict rather than changing it, and the node stays unusable in between because sessions placed on it still have to be driven over the same unreachable URI. It is a diagnosis knob, not a fix.

It is like a contractor who phones the dispatcher to say they are free but leaves a desk extension that only rings inside their own building. The dispatcher writes it down, tries it once, and crosses them off the list.

saying these in an interview costs you the question

  • Blaming the browser or driver when the hub simply cannot reach the node
  • Assuming a successful registration proves the hub can call the node back
  • Restarting the node repeatedly instead of checking the address it advertises
  • Thinking --port changes the address a node reports, not just what it binds
  • Treating a node marked down as proof that the node process has crashed