How do you tell an unreachable Selenium Grid apart from a Grid with no slot matching your request?
answer
- who answered you, if anyone
- plumbing error or protocol error
- instant failure versus a wait
- follow the cause chain to its root
- a network exception or a session-not-created body
basics
~20 sAsk who produced the error. A transport failure has a network cause such as connection refused or unknown host and no WebDriver error body; a Grid that answered returns a session-not-created error whose message echoes the capabilities you asked for.
solid answer
~40 sIn Selenium 4 the two failures end a run at the same point but come from different layers. If the client could not reach the endpoint, the exception's cause chain ends in a network failure such as `java.net.ConnectException` for a refused port, `java.net.UnknownHostException` for a name that did not resolve, or a connect timeout for traffic being dropped; no session id exists and the Grid has no record of the attempt. If the Grid answered and refused, you get `SessionNotCreatedException`, a `session not created` error carried in an HTTP response body, usually after a wait, with a message naming the requested capabilities. The quick discriminators are the root of the cause chain, the timing, and whether the Grid saw the request at all. Catch the subclass first if you branch, since `SessionNotCreatedException` extends `WebDriverException`.
code
java · 24 linesimport java.net.URI;
import java.net.URL;
import org.openqa.selenium.SessionNotCreatedException;
import org.openqa.selenium.WebDriverException;
import org.openqa.selenium.chrome.ChromeOptions;
import org.openqa.selenium.remote.RemoteWebDriver;
public final class CanvasGridSession {
public static RemoteWebDriver start(String gridEndpoint) throws Exception {
URL endpoint = URI.create(gridEndpoint).toURL();
try {
return new RemoteWebDriver(endpoint, new ChromeOptions());
} catch (SessionNotCreatedException refused) {
throw new IllegalStateException("Grid answered and refused the canvas session: " + refused.getMessage(), refused);
} catch (WebDriverException unreachable) {
Throwable root = unreachable;
while (root.getCause() != null) {
root = root.getCause();
}
throw new IllegalStateException("Grid at " + endpoint + " was not reached: " + root, unreachable);
}
}
}go deeper
Recognise that a run can die before the first click for two different reasons, and that the exception's cause is where you look first rather than the test code.
Explain the two layers: an HTTP request that must arrive, and a WebDriver answer that may be a refusal. Name what each failure looks like in the exception.
Diagnose under pressure: check the cause chain, the timing, and whether the Grid saw the request, then hand it to networking or to whoever owns the fleet accordingly.
Own the diagnosability of the wiring, so every environment logs its endpoint, its requested capabilities and the root cause, and neither failure ever reaches a team as an unexplained flake.
## Two layers, two failures Starting a session against a Grid crosses two layers, and each can fail on its own. The **transport** layer has to carry an HTTP request from wherever your test runs to the endpoint you configured. Only if that succeeds does the **WebDriver** layer get a chance to answer, and its answer may be a refusal. In a test report both look the same: the run dies before the first interaction with the survey-builder canvas. In the exception they look nothing alike, and in Selenium 4 the difference is readable in seconds. ## Failure one: nothing answered The client never got a WebDriver-level answer at all, so the error is a plumbing error wearing a Selenium wrapper. Walk the cause chain to the root and you find a network exception: - `java.net.ConnectException` — the host is there and the port is not open, the classic "connection refused"; the Grid is down, or you are on the wrong port. - `java.net.UnknownHostException` — the name did not resolve from where the test ran. A container-network name used from a laptop is the usual version of this. - a connect timeout — packets are being dropped rather than refused, which is what a firewall or a security group looks like from the client side. The distinguishing evidence is negative as much as positive: **no session id exists**, and the Grid has no record of the attempt, because the request never arrived. Reachability must also be checked from where the client actually runs — a host that answers from your machine proves nothing about a pipeline container. ## Failure two: the Grid answered and refused Here the request arrived, the Grid processed it, and the response body carries a W3C `session not created` error. The Java client surfaces that as `SessionNotCreatedException`, and the message echoes the capabilities that were requested. Because the Grid was involved, the attempt is part of its record and the failure usually arrives after a delay rather than instantly. Note the class hierarchy before you branch on it: `SessionNotCreatedException` **extends** `WebDriverException`, so a `catch (WebDriverException)` written first will swallow both and destroy the distinction you were trying to make. Catch the subclass first. ## The thirty-second discriminator | | unreachable endpoint | refused session | |---|---|---| | who produced the error | the client's HTTP layer | the Grid | | root of the cause chain | a `java.net` exception | none; a WebDriver error body | | Java exception you see | `WebDriverException` | `SessionNotCreatedException` | | timing | immediate, or a connect timeout | typically after a wait | | visible on the Grid | no, the request never arrived | yes | | what to change | the address, DNS, ports, network path | what the request asks for, or the fleet | ## The look-alikes in between Two cases muddy the split and are worth recognising by sight. 1. **Something in the middle answered.** A reverse proxy or load balancer in front of the Grid can return a 502 or a 404 with an HTML body. The client reports a response it could not interpret as a WebDriver answer. The endpoint resolves and accepts connections, so it is not "unreachable" — but the Grid never saw the request either, so it is an ingress problem, not a capability problem. 2. **A refusal that is not about matching.** `SessionNotCreatedException` is also what you get when a slot was found and the browser failed to start on the node. The class is identical; only the message separates "nothing could serve this" from "something tried and the browser died". Read the message before concluding anything about capabilities. ## What to log at the point of construction 1. The endpoint string you are about to use, resolved from configuration rather than assumed — a surprising share of these incidents end with someone discovering the value was empty or stale. 2. The capabilities you are sending, so a refusal can be compared against what the Grid offers without re-running anything. 3. The root cause of any failure, not just the top-level message: the wrapper says `WebDriverException` in both worlds, and the root says which world you are in. 4. On success, the session id, so the run can be found on the Grid afterwards. ## Why the distinction is worth the effort The two failures belong to different people and different fixes. An unreachable endpoint is configuration and network: an address, a port, a DNS name, a route out of a CI container. A refused session is a conversation about what the canvas suite asks for and what the fleet provides. Reporting one as the other sends the wrong team looking, and both failure modes are common enough that a suite which cannot tell them apart will eventually waste a day on each.
- The client reports an HTTP 502 with an HTML body instead of either failure. What does that tell you?Something between you and the Grid answered, typically a reverse proxy or load balancer. The endpoint resolves and accepts connections, so it is not unreachable, but the Grid never processed the request either. Treat it as an ingress problem and check what is terminating the connection in front of the Grid.
- Is every session-not-created failure a capability-matching problem?No. The same error and the same Java exception come back when a slot was found and the browser failed to start on the node. The class does not separate the two; only the message does. Read it before concluding anything about capabilities or fleet composition.
- Your client hangs for a long time and then fails. Which side stalled?A connect timeout means the TCP handshake never completed, so nothing accepted you. A read timeout means the connection was accepted and no answer came back in time, which points at the Grid rather than the network. Selenium 4's client configuration sets those two timeouts separately, so this can be a fact rather than a guess.
It is the difference between a phone that never rings and a phone that is answered by someone who says they cannot help you. One is a wiring problem, the other is an inventory problem, and they need different people.
saying these in an interview costs you the question
- Calls every failed session start a network problem without reading the cause
- Assumes a session-not-created error always means no matching slot
- Checks reachability from a laptop when the test runs inside CI
- Expects an unreachable endpoint to fail slowly like a queued request
- Retries construction instead of reading the message the Grid returned