A hosted browser provider is holding your new-session request open — what bounds that wait?
answer
- nobody promised you a deadline
- read the handler, not the marketing
- a slot frees, or the client hangs up
- the socket is yours, so is the patience
- size it under the job budget
basics
~20 sNothing on the far side necessarily bounds it. Selenoid, unmaintained per its own README, waits until a slot frees or the client disconnects, with no server deadline, so your client's read timeout is the bound you actually control.
solid answer
~40 sRead the far side's control flow before assuming a deadline exists. Selenoid — Aerokube's unmaintained session-container grid — gates new sessions through a handler that blocks on exactly two things: a freed run slot, or the request context being cancelled because the client gave up. Its HTTP server sets no read, write or idle timeout, and its own documentation says queued requests `just block and continue to wait`. Selenium Grid does bound its hold with a server-side request deadline, so the behaviour is a per-product choice you must check, never a protocol guarantee. The practical consequence: put the deadline in your client. In Selenium's Java bindings that is `ClientConfig.readTimeout(Duration)`, handed to `RemoteWebDriver.builder().config(...)`, sized under the job budget so a saturated account fails the run loudly instead of silently eating it.
code
java · 23 linesimport java.net.URI;
import java.time.Duration;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.chrome.ChromeOptions;
import org.openqa.selenium.remote.RemoteWebDriver;
import org.openqa.selenium.remote.http.ClientConfig;
// Selenium's own client. The far side may never give up on a queued
// new-session request, so the deadline that matters lives here.
ClientConfig config =
ClientConfig.defaultConfig()
.baseUri(URI.create(System.getenv("TICKETS_GRID_URL")))
.connectionTimeout(
Duration.ofSeconds(Long.parseLong(System.getenv("TICKETS_CONNECT_SECONDS"))))
.readTimeout(
Duration.ofSeconds(Long.parseLong(System.getenv("TICKETS_SLOT_WAIT_SECONDS"))));
// baseUri on the config supplies the address, so do not also call address(...)
WebDriver driver =
RemoteWebDriver.builder()
.oneOf(new ChromeOptions())
.config(config)
.build();go deeper
Know that a hanging test run can be a session that never got a slot, not a broken test. Learn where your client sets its own timeout for talking to a remote grid.
Be able to say why the client, not the service, is the reliable place for the deadline, and to name the knob in the client library you use.
Expect a scenario where a job dies with no error at all, and be ready to explain the unbounded-wait plus unbounded-client combination and how you would instrument it out of existence.
Be ready to set the policy: which lanes get which patience, how construction latency is monitored across the estate, and when chronic waiting stops being a timeout question and becomes a capacity one.
## Read the control flow, do not assume a deadline The instinct on hearing "the request is queued" is that something on the far side is counting down and will eventually give up. That is a product-by-product fact, and it is checkable in the open implementations. Selenoid, unmaintained per its own README, routes a new-session request through a handler that blocks on a `select` with exactly two cases: a run slot becoming free, or the request's context being cancelled. There is no third case, no timer, and its HTTP server is constructed with no read, write or idle timeout of its own. Selenoid's own usage-statistics documentation states it directly: queued requests "just block and continue to wait". The only thing on that side that ends the wait, other than capacity appearing, is the client hanging up — which the handler handles explicitly, draining its queued counter and logging the disconnect rather than leaking it. Selenium Grid takes the other route: it bounds the hold with a server-side request deadline and answers `session not created` when the deadline passes. Both are legitimate; the point is that the behaviour is a per-product choice, never a protocol guarantee. ## What is actually being consumed While a request is held: - **No browser exists yet.** The session has not been created, so a per-session meter has no session to attach to; whether a provider nevertheless counts from admission is something to establish with a deliberately over-concurrent probe rather than to assume. - **Your runner is fully occupied.** The CI machine sits in a blocking call, holding its container, its cache and its place in your own pipeline's concurrency — money on a rented machine, forgone work on one you own. - **The job budget is being spent.** A merge-gating smoke suite for a ticketing site that waits for a slot is consuming the same minutes it would have spent testing. - **The signal is suppressed.** Because the wait ends in success, saturation never surfaces as a failure; it surfaces as a slow week. That asymmetry is the reason a wait feels free and is not. It is cheap where you are watching and expensive where you are not. ## Put the deadline where you control it Since the far side may never give up, the client must. In Selenium's Java bindings the knob is `ClientConfig.readTimeout(Duration)`, handed to the builder: ```java ClientConfig config = ClientConfig.defaultConfig() .baseUri(URI.create(System.getenv("TICKETS_GRID_URL"))) .readTimeout(Duration.ofSeconds(Long.parseLong(System.getenv("TICKETS_SLOT_WAIT_SECONDS")))); ``` A few rules make that deadline useful rather than decorative: 1. **Size it under the job's budget, not over it.** A read timeout longer than the pipeline step's own limit is never reached; the step is killed first and you lose the diagnosis. 2. **Make it a separate value from the command timeout.** Waiting for admission and waiting for a slow page are different failures and want different limits. 3. **Log the elapsed construction time on every session, successful or not.** The distribution of that number is the saturation signal; the timeout is only the safety net. 4. **Do not set it to something enormous "to be safe".** An unbounded client against an unbounded server is a hang, and a hang in CI is the most expensive failure mode there is. ## Instrumenting the wait The measurement you want is the gap between "the suite asked for a session" and "the session existed", per lane, over time. A rising floor is saturation; a fat tail is contention with another consumer of the same account. Where the grid is yours to query, read its own counters instead of inferring: Selenoid publishes a state document with `queued` alongside `total`, `used` and `pending`, so the depth of the queue is a fact rather than a guess. Two further observations that come out of the same code and matter operationally: - A client that gives up is **accounted for**, not leaked — the disconnect path drains the queued counter. So aggressive client-side deadlines do not corrupt the service's own view of demand. - The queue depth itself is unbounded in that implementation; only the simultaneous-run limit is a cap. A service can therefore accumulate an arbitrarily deep backlog of patient clients while reporting a perfectly healthy set of running sessions. ## The interview answer Say plainly that nothing obliges a remote service to bound the hold, cite that at least one readable implementation bounds it only by client disconnect while another bounds it by a server deadline, and then move the responsibility to where it always sits: the client owns its own patience because it owns the socket. Add that the wait certainly costs runner wall-clock, whatever the far side's meter does with it, which is why it hides so well, and that the fix for chronic waiting is admission control — asking for fewer sessions at once — rather than a longer timeout. A candidate who answers "the provider will time it out" has assumed a feature.
- If a client gives up while queued, does the service leak the abandoned request?Not in the implementation you can read. Selenoid, unmaintained per its own README, blocks in a handler that selects on the request context being cancelled as well as on a slot freeing; on cancellation it drains its queued counter and logs the disconnect. So an aggressive client-side deadline does not corrupt the service's own view of demand. Do not assume the same of every product, but the pattern is the sane one.
- How should the session-wait deadline relate to the pipeline step's own limit?Strictly under it. A read timeout longer than the step's limit is never reached: the step is killed first and you lose the diagnosis entirely, ending up with a dead job and no error to read. Sized under the step's budget, a saturated account produces a clear client-side failure with a message and a timestamp you can act on.
- Why is a very large client timeout worse than none at all in CI?Both are hangs, but a large one is a hang you believed you had handled. An unbounded client against a service that also never gives up produces a job that occupies a runner until an outer kill, with no diagnosis and the whole wait paid for out of your own capacity. Choose a value you would actually be willing to spend, and record the construction latency separately so the number can be revisited.
saying these in an interview costs you the question
- Assuming the provider will eventually time the request out
- Setting an enormous client timeout so nothing ever fails
- Using one timeout for session creation and for page commands
- Assuming what the far side's meter does with a queued request instead of probing it
- Reading total run time instead of session-construction latency