When a hosted browser provider has no free slot, what can it do with your new-session request?
answer
- a ceiling you did not set
- hold it, or hand it back
- a queued request has no browser yet
- the wait burns your runner's clock
- unmaintained Selenoid: -disable-queue refuses
basics
~20 sA provider with every slot busy either holds the new-session request until a slot frees or refuses it at once. Holding burns runner wall-clock and hides the ceiling; refusing surfaces it immediately and forces the suite to decide.
solid answer
~50 sOver an account ceiling a service must either hold the request or refuse it, and the choices are not interchangeable. Holding means the request sits open until capacity frees: no browser has started, so a per-session meter has no session to attach to — whether a provider nevertheless counts from admission is to be probed, not assumed — while your runner burns wall-clock every second it waits and the queue is invisible unless you look. Refusing means an error arrives at session construction, before any command runs, so the suite learns the truth immediately and must own the retry. Selenoid — an unmaintained session-container grid, per its own README — implements both: by default its own docs say queued requests `just block and continue to wait`, while `-disable-queue` makes it answer `unknown error` with the message `queue is full`. Choose deliberately: each trades visibility against a chance of success.
code
go · 22 lines// aerokube/selenoid (unmaintained per its own README), main.go:
// the new-session route chains two fail-fast gates in front of the blocking one
mux.HandleFunc(seleniumPaths.CreateSession,
post(queue.Try(queue.Check(queue.Protect(create)))))
// protect/queue.go - Try refuses only when the client asked it to
func (q *Queue) Try(next http.HandlerFunc) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
_, noWait := r.Header["X-Selenoid-No-Wait"]
select {
case q.limit <- struct{}{}:
<-q.limit // room right now; peek only, nothing reserved
default:
if noWait { // full, and the client asked not to wait
err := errors.New(http.StatusText(http.StatusTooManyRequests))
jsonerror.UnknownError(err).Encode(w)
return
}
}
next.ServeHTTP(w, r) // otherwise Check, then Protect blocks
}
}go deeper
Know that a remote provider can make your session request wait, and that the wait happens before any test code runs. If a run is slow but every test looks normal, say out loud that session start-up is where to look.
Be ready to describe both behaviours and what each costs, and to name an open implementation that does each. Explain why a held request burns your runner's time while no browser exists yet on the far side.
Expect to be asked how you would detect chronic queueing in a suite nobody has complained about, and to answer with the session-construction latency distribution rather than with total run time.
Be prepared to argue which behaviour you want per lane and to justify buying it — a merge-gating suite and an overnight sweep have opposite answers, and the wait each one can afford is the thing to settle.
## The moment the ceiling bites A hosted browser provider sells capacity as an account-level ceiling: a limit on how many sessions your credential may hold open at the same time. Your test runner neither knows that ceiling nor enforces it; it opens as many sessions as its own worker count tells it to, and the ceiling is discovered on the request that would have crossed it. What the service does with **that** request is the whole subject here, and it is a design choice rather than anything the protocol settles. A service over its ceiling has two options, and they sit at opposite ends of one trade: - **Hold the request.** The HTTP connection stays open, no browser is started, and the client blocks inside its own new-session call until capacity appears. - **Refuse the request.** An error comes back immediately, before any command has been sent, and the caller has to decide what happens next. | | holding | refusing | |---|---|---| | what the client sees | a call that has not returned | an error at session construction | | when you learn the ceiling | after the wait, or never | immediately | | what is consumed | runner wall-clock and a job's budget | almost nothing | | who absorbs the overflow | the service | your suite | | failure mode | a green run that took far too long | a red run that was really a capacity signal | ## Holding: the wait that looks like nothing A held request is harder to reason about because it produces no artefact. Your ticketing site's smoke suite still passes on every merge; only its wall-clock creeps upward, and nobody can point at a slower test. The symptoms are indirect: - The time disappears *before* the first command, so per-test timing, screenshots and traces all look normal — the case did not start late, the session did. - A queued request has no browser behind it yet, so a per-session meter has no session to attach to; whether a provider nevertheless counts from admission is something to establish with a deliberately over-concurrent probe, not to assume. What is certain is your own side: the runner holds the connection open and burns wall-clock for every second of it. - Because the wait ends in success, a permanently saturated account looks like a merely slow suite, and teams optimise tests instead of looking at the ceiling at all. ## Refusing: the error that arrives before the first command A refusal is louder and, for that reason, often more useful. The request comes back with an error at construction time, so: - The failure is unambiguously about admission, not about the page under test, and the suite rather than the service decides whether to back off, stagger, shed load or fail the run. - Nothing is spent waiting, so a saturated account costs a fast red instead of a slow green. - The cost is that a transient burst — one lane overlapping another — becomes a hard failure that a short wait would have absorbed. ## What the open implementation shows Selenoid, an unmaintained session-container grid by its own README's admission, is worth reading because it implements both behaviours and chains them explicitly. Its new-session route is wrapped in three middlewares, outermost first: 1. `Try` refuses at once **only if** the client sent the `X-Selenoid-No-Wait` request header. 2. `Check` refuses at once **only if** the server was started with `-disable-queue`. 3. `Protect` otherwise takes a run slot, blocking until it can. Both `Try` and `Check` merely peek at the run limit with a non-blocking send that they immediately take back, so neither reserves anything; `Protect` is the handler that actually holds the slot. And Selenoid's own usage-statistics documentation describes the default plainly: queued requests "just block and continue to wait". The refusal, when it happens, is a W3C `unknown error` whose message carries the cause — `queue is full` on the `-disable-queue` path. Selenium Grid, by contrast, holds the request but bounds the hold with a server-side request deadline and ends it with `session not created`. Same situation, three different observable behaviours across two open products. ## Writing the suite for each, and what to watch The suite's job is to make whichever behaviour it faces legible. - **If the service holds:** put a deadline on the client's own new-session call, so a saturated account fails inside your job budget rather than consuming all of it. - **If the service refuses:** catch the construction failure, tell it from a real defect by its message and its timing, and apply admission control on your side — stagger lane starts, or ask for fewer sessions at once so the refusal stops arising. - **Either way:** treat saturation as a fact about the account, not as evidence about the code under test. The response is to change how many sessions you ask for, or when you ask for them. The single most useful instrument is the elapsed time between "the suite asked for a session" and "the session existed". Plot it; a rising floor there is saturation, and it shows up long before anyone notices total run time. And be honest about the trade: holding buys a chance of success at the price of visibility, and refusing buys visibility at the price of that chance.
- If holding hides saturation, what is the one metric that exposes it?The elapsed time between asking for a session and having one, recorded per lane and separately from test duration. A rising floor in that distribution is a saturated account, and it moves long before total run time becomes suspicious. Where the grid exposes its own counters — Selenoid, unmaintained by its own README's admission, publishes `queued` beside `total`, `used` and `pending` — read the depth directly rather than inferring it.
- Why does a refusal cost less than a hold even though it fails the run?Because a refusal spends almost no wall-clock and produces an unambiguous signal at session construction, before any command has been sent. A hold spends the job's budget to reach the same outcome later, or reaches success while concealing that the account is permanently at its ceiling. Fast and red is usually cheaper to act on than slow and green.
- Does the client have any say in which behaviour it gets?Sometimes. Selenoid, unmaintained per its own README, honours an `X-Selenoid-No-Wait` request header, which makes it answer immediately instead of queueing — a mechanism that lives in its source and in none of its own documentation. But that is a per-product courtesy, and an intermediary in front of the service may strip or rewrite the header, so treat a client-side hint as a preference rather than a guarantee.
saying these in an interview costs you the question
- Assuming every provider queues, so a request always eventually succeeds
- Asserting what a provider's meter does with a queued request instead of probing it
- Treating a slow suite as slow tests when the time is spent before any command
- Thinking the WebDriver protocol dictates what a full service must do
- Claiming a refusal means the account was misconfigured rather than saturated