How does a suite tell a capacity refusal from a genuine session failure at a remote grid?
answer
- the reply looks like any other failure
- the protocol has no word for full
- generic code, specific message
- it fails before the first command
- read your own queue depth
basics
~20 sNot from the error code. WebDriver defines no capacity error, so implementations reuse the generic ones and only the message string names the cause. Read the message text, the failure's timing, and your own queue counters instead.
solid answer
~40 sThe W3C WebDriver error table has no entry for *out of capacity*, so every implementation improvises inside the generic codes. Selenoid — unmaintained per its own README — answers a refusal with `unknown error` and puts the cause in the message: `queue is full` when `-disable-queue` is set, `Too Many Requests` when the client sent `X-Selenoid-No-Wait`. It reserves `session not created` for an attempt that genuinely failed, and Selenium Grid uses that same code when its bounded hold expires — so the code alone cannot separate saturation from a broken node. Three signals do the work instead: the message string, the fact that the failure lands at session construction before any command runs, and your own instrumentation. Selenoid's state endpoint reports `queued` beside `total`, `used` and `pending`, which turns a guess into a measurement.
code
go · 17 lines// aerokube/selenoid (unmaintained per its own README), jsonerror/jsonerror.go
func SessionNotCreated(err error) *SeleniumError {
return newSeleniumError("session not created", err, http.StatusInternalServerError)
}
func UnknownError(err error) *SeleniumError {
return newSeleniumError("unknown error", err, http.StatusInternalServerError)
}
// protect/queue.go - both capacity refusals go through UnknownError,
// so only the message distinguishes them:
jsonerror.UnknownError(errors.New("queue is full")).Encode(w) // -disable-queue
jsonerror.UnknownError(
errors.New(http.StatusText(http.StatusTooManyRequests)), // X-Selenoid-No-Wait
).Encode(w)
// the encoded body:
// {"value": {"error": "unknown error", "message": "queue is full"}}go deeper
Know that a failure which happens before any test step ran is about getting a browser, not about the page. Read the message in the error rather than only its type.
Be able to explain that the protocol has no capacity error, so implementations reuse generic codes and put the real cause in free text.
Expect to design the classification: message matching plus failure timing plus your own queue metric, with an unrecognised failure staying unrecognised rather than being forced into a bucket.
Be ready to set the standard across teams — capacity signals reported as account facts, verified against each provider with a real probe, and never inferred from another product's behaviour.
## There is no capacity error The W3C WebDriver specification defines a fixed table of error codes — `session not created`, `unknown error`, `invalid session id`, `stale element reference` and the rest. **None of them means "the remote end is at capacity".** Nothing in the protocol carries "we are full" as a distinct, machine-readable outcome. Work through what the reply actually offers you: - **The HTTP status** is no help: both codes a full service is likely to reach for map to the same server-error status in that same table. - **The error name** is generic by construction, because the only candidates are a catch-all and a code that also covers a browser that failed to start. - **The message** is free text, so it is the only field that can carry the cause — and it is prose, not API. That is the whole reason this question is asked at all. A suite that branches on the error code to detect saturation is branching on a value that was never designed to carry the information. ## What the open implementations actually send Reading them side by side is the fastest way to see the improvisation. Selenoid and Ggr below are Aerokube projects that declare themselves unmaintained in their own READMEs, and are read here only as a legible model of the mechanism: | situation | error name | message | |---|---|---| | Selenoid, `-disable-queue`, limit reached | `unknown error` | `queue is full` | | Selenoid, `X-Selenoid-No-Wait` sent, limit reached | `unknown error` | `Too Many Requests` | | Selenoid, a start attempt that genuinely failed | `session not created` | the underlying cause | | Selenium Grid, bounded hold expires | `session not created` | a timed-out message | | Ggr, every candidate host exhausted | a legacy-shaped body | names how many attempts it made | Selenoid — unmaintained by its own README's admission — is instructive because it has a `SessionNotCreated` helper and does **not** use it for either capacity refusal. It sends `unknown error` and puts the cause in the message string, keeping `session not created` for an attempt that was admitted and then failed. Selenium Grid uses `session not created` for its expired hold. So the same code means "we were busy" in one product and "the browser did not start" in another. ## The three signals that do work Since the code cannot separate them, use the signals that can: - **The message string.** It is the only field that names the cause, and it is therefore the only thing worth matching on — while knowing that it is prose, not API, and can change without notice. - **The timing.** A capacity refusal lands at session construction, before a single command has been sent. A defect in the page under test cannot fail there. That one bit of information separates most cases without reading any text at all. - **Your own counters.** The elapsed time from "asked for a session" to "session existed", plotted per lane, plus the service's own queue depth where it exposes one. Selenoid publishes `queued` beside `total`, `used` and `pending`; that turns "I think we are saturated" into a measurement. A fourth, weaker signal is correlation: refusals that arrive in a burst across unrelated cases, at the moment another consumer of the same credential starts, are about capacity and not about the tests. ## Writing the suite for it Some practical shape for a merge-gating smoke suite on a ticketing site: 1. **Catch failures at construction separately from failures during a test.** They are different events and should not land in the same bucket in the report. 2. **Classify by message, but never assume the message.** Treat an unrecognised construction failure as unknown rather than forcing it into the capacity bucket. 3. **Report saturation as a capacity fact about the account**, not as a property of the code under test — the response is to ask for fewer sessions, or to ask at a different time. 4. **Verify the classification against the real service.** Run a deliberately over-concurrent probe and capture the exact body you get back; that is the only way to know what your provider sends, and it is checkable in a way that a remembered answer is not. ## Why this is a senior question It looks like trivia about error codes and is really about not trusting a field to carry a meaning it was never given. The candidate who says "catch `session not created`" has an answer that is right for one product and wrong for the next. The candidate who says "the code is generic, so I read the message, the timing, and my own queue metric, and I verify all three against the actual service" has described something that survives changing providers. The same reasoning generalises past this leaf. Wherever a protocol has no code for a condition, implementations put the condition in free text, and free text is a fragile contract. Match it if you must, but build the detection on the structural facts around it — when the failure occurred, and what your own instruments say — because those do not change when someone rewords a string.
- If matching on message text is brittle, why recommend it at all?Because it is the only field that names the cause, and brittle is not the same as useless. Match it, but never rely on it alone: treat an unrecognised construction failure as unknown rather than forcing it into the capacity bucket, and back the classification with the structural signals — when the failure occurred, and what your own queue metric said at that moment.
- What single structural fact separates a capacity refusal from a product defect?When it happens. A capacity refusal lands at session construction, before a single WebDriver command has been sent, so nothing about the page under test can have caused it. Recording construction failures in a different bucket from in-test failures gives you that separation for free, with no string matching at all, and it survives a provider changing its wording.
- How would you find out what your own provider actually sends when it is full?Run a deliberately over-concurrent probe against it and capture the exact response body, status and timing. That is a cheap experiment and the only trustworthy source, since the behaviour is a product choice rather than a protocol rule. Keep the captured sample in the repository beside the classification code so the next person can see what the matching was written against.
A full restaurant can seat you late or turn you away at the door, but the slip it hands you is the same one it uses when the kitchen catches fire. Only the line scribbled on it tells you whether to wait or go somewhere else.
saying these in an interview costs you the question
- Branching on the WebDriver error code to detect saturation
- Assuming session not created always means the remote end was full
- Expecting a dedicated capacity error code in the specification
- Reading the HTTP status as the signal that the grid is busy
- Classifying every construction failure as a capacity refusal