In a fully distributed Selenium Grid 4, why is the distributor started with --bind-bus false?
answer
- A default inherited from the hub
- Only one process owns the sockets
- Bound versus connected
- Two brokers, no shared events
- Its config ships events bind true
basics
~20 sBecause the distributor defaults to binding the event bus itself, the way a hub does. When a separate event-bus process already owns ports 4442 and 4443, --bind-bus false makes the distributor join that bus rather than start a second.
solid answer
~40 s`--bind-bus` chooses between hosting the bus and joining it: true gives a `BoundZmqEventBus` that binds `XPUB` and `XSUB` sockets, false gives an `UnboundZmqEventBus` that connects to them. Most components default to false, but `DefaultDistributorConfig` ships `bind` as `true`, inherited from the fact that a `hub` is a distributor that also hosts the bus. In a distributed Grid that default is wrong. Leaving it true makes the Distributor try to bind an address belonging to another machine, which fails, or bind a second broker that no Node is connected to, which is worse because nothing logs an error. The symptom is a Grid where car-rental pickup-form requests queue and expire while the Distributor reports no Nodes at all.
code
bash · 7 linesjava -jar selenium-server.jar distributor \
--publish-events tcp://10.0.0.5:4442 \
--subscribe-events tcp://10.0.0.5:4443 \
--sessions http://10.0.0.6:5556 \
--sessionqueue http://10.0.0.8:5559 \
--bind-bus false \
--port 5553go deeper
Know that one process hosts the event bus and the rest connect to it. You will not be asked to wire a distributed Grid, only to recognise that the roles differ.
Explain what the flag switches between, and why the Distributor's default differs from the Session Map's and the Node's. Naming the hub as the reason is the point.
Walk the failure: two brokers, no error, a Grid that queues forever. Show the log lines and the health check you would use to prove which process is hosting the bus.
Own the configuration standard so this cannot recur: one declared bus host, the same address pair everywhere, and a startup check that fails loudly rather than silently.
## The default that surprises people `--bind-bus` is a boolean that decides whether a component **hosts** the event bus or merely **joins** one. Internally, `ZeroMqEventBus.create` reads it and returns either a `BoundZmqEventBus` — which binds an `XPUB` socket and an `XSUB` socket and runs a proxy between them — or an `UnboundZmqEventBus`, which connects a `SUB` and a `PUB` socket to addresses somebody else has bound. The surprise is the default. Most components leave `events.bind` unset, and the code falls back to `false`. The **Distributor does not**: `DefaultDistributorConfig` explicitly ships `"bind", true`, exactly as `DefaultHubConfig` does. That is deliberate — a `hub` is a distributor that also hosts the bus, and running `distributor` on its own is meant to work the same way without extra flags. In a fully distributed Grid that default is wrong, because a separate `event-bus` process is already the host. ## Bound versus unbound, side by side | | `--bind-bus true` | `--bind-bus false` | |---|---|---| | Class used | `BoundZmqEventBus` | `UnboundZmqEventBus` | | Sockets | `XPUB` and `XSUB`, both bound | `SUB` and `PUB`, both connected | | Startup log | `XPUB binding to ..., XSUB binding to ...` | `Connecting to <publish> and <subscribe>` | | Who should use it | the `event-bus` process, a `hub`, a `standalone` | the Distributor, the Session Map, every Node | | Failure if wrong | address unavailable, or a second silent broker | none; it is the safe setting | ## What actually goes wrong if you leave it true 1. The Distributor is given `--publish-events tcp://10.0.0.5:4442`, where `10.0.0.5` is the bus host, and tries to **bind** that address. A process can only bind an address on an interface it owns, so on a different machine the bind fails outright and the Distributor does not come up. 2. If the Distributor happens to share a host with the bus, the bind collides with the sockets the `event-bus` process already holds, and you get an address-in-use failure instead. 3. Worst case, the addresses differ just enough that both binds succeed — say the bus was started on `tcp://*:4442` and the Distributor on a second interface. Now there are **two brokers**. Every Node that connects to the first one is invisible to a Distributor sitting on the second, and no error is logged anywhere, because ZeroMQ connect succeeds against an address nobody is publishing on. 4. A fourth variant is quieter still. If the Distributor comes up **before** the `event-bus` process, its bind can succeed on a free port, and the bus process then fails to bind moments later. Start order is not decoration in a distributed Grid; the bus host goes first precisely so nothing else can claim those sockets. The visible symptom of case 3 is the one that reaches an interview: the car-rental pickup-form suite gets no sessions at all. The Router accepts requests, the New Session Queue accepts them, and they expire there, because the Distributor never learned that any Node exists. Nothing in the Distributor's own log looks wrong, which is why candidates who have only read the architecture diagram tend to go hunting in the queue's timeouts instead. ## The correct wiring Give the Distributor the same two bus addresses as everybody else and add `--bind-bus false`. It still needs the HTTP URLs of the Session Map and the New Session Queue, because `LocalDistributor.create` builds a `SessionMap` and a `NewSessionQueue` client from configuration; only the Node picture comes over the bus. ```bash java -jar selenium-server.jar distributor \ --publish-events tcp://10.0.0.5:4442 \ --subscribe-events tcp://10.0.0.5:4443 \ --sessions http://10.0.0.6:5556 \ --sessionqueue http://10.0.0.8:5559 \ --bind-bus false \ --port 5553 ``` ## How to confirm you got it right - **Read the Distributor's startup log.** `Connecting to tcp://10.0.0.5:4442 and tcp://10.0.0.5:4443` is what you want. An `XPUB binding to` line means it is hosting a bus of its own. - **Read the bus process's log too.** Exactly one process in the whole Grid should print the binding line. - **Ask the bus for its health.** `GET /status` on port `5557` pushes a `healthcheck` event through the proxy and answers `Event bus running` only if the round trip completed. - **Watch for a Node that never shows up.** A Node whose bus addresses are right but whose Distributor is on a different broker looks perfectly healthy on its own machine and simply never appears. - **Treat a silent Grid as a bus problem first.** Requests that queue and expire while every Node process is running and reachable point at the message channel, not at capacity. ## The generalisable point Every distributed Grid has exactly one bus host, and every other bus member connects. `--bind-bus` is how you say which is which, and the only component whose default fights you is the Distributor, because it inherited that default from the hub topology it can also serve.
- Which other components ship with the bind setting defaulting to true?The `event-bus` process itself and the `hub`, because each is meant to host the broker in its own topology. `sessions` and `node` leave it unset and fall back to false, and `router` never touches the bus at all, so the Distributor is the only role whose default fights a distributed layout.
- How would you tell from the logs that the distributor bound its own bus?A component hosting the bus logs an XPUB and XSUB binding line at startup, while one joining logs a connecting line naming the two addresses. Seeing the binding line on both the event-bus process and the Distributor means you have two brokers and no shared events.
- Why does the Distributor still need the sessions and sessionqueue URLs?Only the Node picture arrives over the bus. The Distributor polls the New Session Queue and writes into the Session Map over plain HTTP, so it builds a queue client and a session map client from configuration. Omitting either URL leaves it unable to take work or record a created session.
saying these in an interview costs you the question
- Thinks --bind-bus takes an address rather than true or false
- Assumes every component defaults to connecting to the bus
- Believes the distributor never touches the event bus
- Says the flag only matters when components share a host
- Blames the queue timeout instead of the split bus