skip to content

When a queued Selenium Grid session request hits its timeout, what does the client actually receive?

level: seniorimportance: should knowfreq 41%

answer

  1. Two clocks, only one belongs to Grid
  2. The client can give up first
  3. A five hundred with a W3C error body
  4. One hundred eighty against three hundred
  5. SessionNotCreatedException, or no Selenium error at all

basics

~20 s

An HTTP 500 carrying the WebDriver error session not created, which the Java client raises as SessionNotCreatedException. With defaults, though, the client's own 180-second read timeout fires before Grid's 300-second queue timeout, so a transport error appears instead.

solid answer

~40 s

When a request outlives `--session-request-timeout`, Selenium Grid completes it with a `SessionNotCreatedException` and answers the client with HTTP 500 and the W3C body carrying `"error": "session not created"` plus a message and stacktrace. The Java client rethrows that as `SessionNotCreatedException` from the `RemoteWebDriver` constructor. The trap is that the client rarely waits long enough to see it: in Selenium 4 `ClientConfig`'s read timeout defaults to 180 seconds while `--session-request-timeout` defaults to 300, so on a congested Grid the socket read expires first and you get a transport failure with no WebDriver error code at all. Line the two numbers up - raise the client's `readTimeout` above the Grid's timeout, or lower the Grid's below 180 seconds - so congestion always reports as `session not created`.

code

java · 24 lines
java
import java.net.URI;
import java.time.Duration;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.chrome.ChromeOptions;
import org.openqa.selenium.remote.RemoteWebDriver;
import org.openqa.selenium.remote.http.ClientConfig;

public class AlbumUploadSession {

  public static void main(String[] args) throws Exception {
    ClientConfig config = ClientConfig.defaultConfig().readTimeout(Duration.ofSeconds(330));

    WebDriver driver =
        new RemoteWebDriver(
            URI.create("http://grid.internal:4444").toURL(), new ChromeOptions(), config);

    try {
      driver.get("https://albums.example.com/upload");
      System.out.println(driver.getTitle());
    } finally {
      driver.quit();
    }
  }
}

go deeper

for a junior

Know that a new-session call against Selenium Grid can fail because no slot ever became free, not only because the browser or driver is broken. The error reported in that case is session not created.

for a middle

Explain the mechanics: the request sits in the queue until --session-request-timeout expires, then Grid answers HTTP 500 with the W3C session not created error, which the Java client raises as SessionNotCreatedException.

for a senior

Demonstrate the diagnosis. Notice that the client's own read timeout can fire first and mask the real error, compare the two configured numbers before blaming the Grid, and align them so congestion reports honestly.

for a principal

Own the contract between suite and platform. Decide which side owns the timeout budget, make the client and Grid numbers deliberate rather than accidental defaults, and ensure saturation surfaces as one recognisable failure everywhere.

## What the Grid puts on the wire When a new-session request outlives `--session-request-timeout` (default **300 seconds**), Selenium Grid completes it with a `SessionNotCreatedException` and answers the waiting client with **HTTP 500** and a standard WebDriver error body: ```json {"value": {"error": "session not created", "message": "Timed out creating session", "stacktrace": "..."}} ``` The W3C WebDriver specification maps the `session not created` error to status 500, so this is the ordinary, specified shape of "your session could not be started" - not a Grid-specific invention. The Java client turns it back into a `SessionNotCreatedException` thrown from the `RemoteWebDriver` constructor, so the photo-album uploader suite sees the failure at the point it asked for a browser. ## Two internal paths, one wire error There are two ways a queued request reaches that outcome, and they produce slightly different messages: - The thread serving the client waits on the request for `--session-request-timeout`; if that wait expires first, the failure message is `New session request timed out`. - A background sweep, running on `--session-request-timeout-period` (default **10 seconds**), reaps requests whose deadline has already passed and fails them with `Timed out creating session`. Both are `SessionNotCreatedException`, both become the same `session not created` error code, and neither distinguishes "everything was busy" from "nothing could ever match". Because the sweep runs on its own cadence, a request can linger a few seconds past its nominal deadline before it is removed. ## The clock the Grid does not control The trap is that the client usually gives up before the Grid does. The Java client's `ClientConfig` sets a default HTTP **read timeout of 180 seconds** (overridable with the `webdriver.httpclient.readTimeout` system property), while `--session-request-timeout` defaults to 300. With both left at their defaults, the socket read expires while the request is still sitting in the queue, and the caller sees a transport-level failure with no WebDriver error code in it at all. | | Client read timeout | Grid request timeout | |---|---|---| | Setting | `ClientConfig.readTimeout` | `--session-request-timeout` | | Default | 180 seconds | 300 seconds | | Owned by | the test process | the Grid deployment | | What fires | a transport read failure | HTTP 500, `session not created` | | Diagnosable as saturation | poorly | directly | That mismatch is why the same congestion produces two entirely different bug reports: a suite with a raised read timeout reports `session not created`, and a suite on the defaults reports a hung or broken connection to the Grid and gets misfiled as a network problem. ## Aligning the two numbers 1. Decide the longest queue wait you are willing to absorb - that is your `--session-request-timeout`. 2. Set the client's `readTimeout` **above** it, with a margin, so the Grid always answers first. 3. Or, if you cannot change every client, lower `--session-request-timeout` **below** 180 seconds so the default client outlasts it. 4. Whichever direction you choose, make both numbers explicit in configuration; two accidental defaults that disagree is how this stays invisible. ## Why no browser is left stranded A reasonable worry is that the client walks away, a slot then frees, and a browser starts with nobody driving it. Grid handles this. When the Distributor finishes creating the session it calls back into the queue to complete the request; if the request has already been reaped or the connection dropped, that call reports the session as no longer valid, and the Distributor issues a `DELETE` for the session it just created rather than leaving it running on the Node. ## Reading the failure correctly - `session not created` after a long pause on a busy Grid is a **capacity** report, not a driver or browser-version report. - The same error arrives instantly, with a message about no Nodes supporting the requested capabilities, only when the Distributor is started with `--reject-unsupported-caps`; that flag defaults to **false**, so by default an unmatchable request burns the full deadline first. - A hung new-session call that never returns an error code is usually the client's own read timeout, not an unreachable Grid. - Raising `--session-request-timeout` after a wave of these failures changes when they happen, never whether the Grid has enough slots to avoid them. Treat the timeout as the honest edge of a promise: the Grid will hide a queue wait from the photo-album uploader suite up to that many seconds, and past it will tell the truth in the protocol's own vocabulary.

  • How do you stop a mistyped browserName from burning the whole request timeout?
    Start the Distributor with `--reject-unsupported-caps`, which defaults to false. With it enabled, a request whose capabilities no registered Node supports is failed immediately with `session not created` and a message saying no nodes support the requested capabilities, instead of sitting in the queue until `--session-request-timeout` expires.
  • If the client disconnects while waiting and a slot then frees, is a browser left running?
    No. When the Distributor finishes creating the session it calls back into the queue to complete the request, finds it already removed, logs that the request timed out or the connection dropped, and sends a delete for the session it just created. The browser is stopped rather than stranded on the Node.
  • Does raising --session-request-timeout fix a Grid that keeps timing out?
    Only when the wait was genuinely almost long enough. Raising it converts a fast failure into a slower one and does nothing about slot supply. If requests time out because arrival rate exceeds capacity, the queue simply grows deeper against the higher bound.

saying these in an interview costs you the question

  • Assumes the client always sees session not created when the Grid queue times out
  • Thinks a hung new-session call means the Grid is unreachable rather than saturated
  • Never compares the client's own read timeout against the Grid's request timeout
  • Believes a timed-out queued request leaves a browser running on the Node forever
  • Reads session not created on timeout as a driver or browser version mismatch