skip to content

A Selenium Grid request stays queued: how do you tell a busy slot from a missing stereotype?

level: seniorimportance: should knowfreq 54%

answer

  1. One symptom, two very different causes
  2. Ask the grid what was requested
  3. Then ask what the nodes advertise
  4. Occupancy versus stereotype match
  5. sessionQueueRequests against nodesInfo stereotypes

basics

~20 s

Compare what was asked for with what is advertised. Read the pending capabilities from Selenium Grid's GraphQL sessionQueueRequests, then list every node's stereotypes and occupancy: a full match means saturation, no match means the request can never be served.

solid answer

~40 s

Two queries against a Selenium 4 Grid separate the cases. First `{ grid { sessionQueueSize } sessionsInfo { sessionQueueRequests } }` — `sessionQueueRequests` returns one JSON string per waiting request holding exactly the capabilities the client sent. Then `{ nodesInfo { nodes { uri status stereotypes slotCount sessionCount maxSession } } }`, or the `nodes` array of `GET /status`, for what the grid advertises and how full it is. If some node's `stereotypes` matches the request and its `sessionCount` equals `maxSession`, you have **saturation** — correct request, no room. If no node advertises a matching stereotype, the request can never be served however long it waits, and the near-miss case (right `browserName`, a `browserVersion` or `platformName` nobody offers) is the one that masquerades as saturation. Count `DRAINING` and `DOWN` nodes out of usable capacity.

code

bash · 7 lines
bash
curl -s -X POST -H "Content-Type: application/json" \
  --data '{"query":"{ grid { sessionQueueSize } sessionsInfo { sessionQueueRequests } }"}' \
  http://localhost:4444/graphql | jq -r '.data.sessionsInfo.sessionQueueRequests[]'

curl -s -X POST -H "Content-Type: application/json" \
  --data '{"query":"{ nodesInfo { nodes { uri status maxSession slotCount sessionCount stereotypes } } }"}' \
  http://localhost:4444/graphql | jq '.data.nodesInfo.nodes'

go deeper

for a junior

Know that a queued request means the grid has not matched it yet, and that the console's Sessions screen lists the waiting requests. Being able to find that panel is the expectation here.

for a middle

Explain the mechanics of the comparison: pending capabilities from sessionQueueRequests on one side, node stereotypes and occupancy from nodesInfo or the /status nodes array on the other.

for a senior

Demonstrate the diagnosis under pressure: separate saturation from a near-miss stereotype, discount draining nodes, spot a leaked session holding a slot, and say what you would change in each case.

for a principal

Own the systemic view. Decide what grid signals get exported and alerted on, how shared-grid contention between teams is arbitrated, and whether stereotype drift should be prevented rather than diagnosed.

## Two causes, one symptom A request that never becomes a session looks identical from the client side no matter why it is stuck. Inside a **Selenium 4 Grid** there are two quite different causes, and they need opposite responses: - **Saturation.** Some node genuinely advertises a matching **stereotype**, but every such slot is occupied. The request is correct and will start as soon as a slot frees. The fix is capacity or scheduling. - **No matching stereotype.** Nothing in the grid advertises what the request asked for — a wrong `browserName`, a `browserVersion` or `platformName` no node offers, a typo. The request will wait forever no matter how long you leave it. The fix is the request or the node configuration. The grid's own endpoints tell you which one you have in two queries. ## Step one: read what was actually requested Ask the grid what is in its queue: ```bash curl -s -X POST -H "Content-Type: application/json" \ --data '{"query":"{ grid { sessionQueueSize } sessionsInfo { sessionQueueRequests } }"}' \ http://localhost:4444/graphql ``` `sessionQueueRequests` is a list of strings, one per waiting request, each the JSON of that request's capabilities. This is the single most useful thing on the whole surface, because it shows what the **client** sent rather than what you believe it sent — after option classes, config files and environment variables have all had their say. The `/ui` console renders the same list in its **Sessions** screen, under the running sessions. ## Step two: read what the grid advertises ```bash curl -s -X POST -H "Content-Type: application/json" \ --data '{"query":"{ nodesInfo { nodes { uri status maxSession slotCount sessionCount stereotypes } } }"}' \ http://localhost:4444/graphql ``` `stereotypes` is a JSON string describing the slot shapes that node offers. `status` is `UP`, `DRAINING` or `DOWN`. `sessionCount` against `maxSession` is that node's occupancy. The same picture is in `GET /status`, where each node lists its `slots`, and every slot carries a `stereotype` and either a `session` object or `null`. ## Telling the two apart | Observation | Reading | |---|---| | A node advertises the stereotype, `sessionCount` equals `maxSession` | **saturated** — the request is fine, wait or add capacity | | A node advertises it and slots are free, yet the request still waits | look at node `status`: `DRAINING` and `DOWN` nodes take no new work | | No node's `stereotypes` mentions the requested browser at all | **no match** — this request can never be served as written | | `nodesInfo { nodes }` is empty, `/status` `nodes` is empty | nothing has registered; this is a grid problem, not a request problem | | Stereotype matches on `browserName` but not on a version or platform key | **no match**, and the near-miss is why it looks like saturation | The near-miss row is where most time is lost. Matching is over the whole requested capability set, so a request for a `browserVersion` no node advertises fails to match even though the browser name lines up perfectly — and the queue looks exactly like a busy grid. ## A worked permit-checklist run A nightly suite for a **construction-permit checklist** application starts sixteen parallel sessions against a grid of two nodes. Half the run proceeds; the rest sit in the queue. 1. `{ grid { sessionQueueSize sessionCount maxSession } }` returns a queue of eight, `sessionCount` 8, `maxSession` 8. The grid is exactly full. 2. `{ sessionsInfo { sessionQueueRequests } }` shows all eight waiting requests asking for `browserName: chrome`. 3. `{ nodesInfo { nodes { status stereotypes sessionCount maxSession } } }` shows both nodes `UP`, both advertising `chrome`, both at full occupancy. That is **saturation**: nothing is misconfigured and the run will finish, just serially. Now change one detail — the checklist team pins `browserVersion` for a print-layout check, and only those requests queue. Step 2 shows the pinned version in the waiting capabilities; step 3 shows no node advertising it. That is a **no-match**, and no amount of waiting resolves it. ## Traps worth naming - **A non-empty queue is not a fault.** Queue depth is a normal steady state on a shared grid; it only matters against occupancy and match. - **Don't scale a grid that has no matching stereotype.** Adding identical nodes multiplies capacity the request cannot use. - **Count `DRAINING` and `DOWN` nodes out.** They still appear in `nodesInfo` and in `/status`, so raw node counts overstate usable capacity. - **A node that never registered is invisible here.** The router only reports what the distributor knows, so hit that node's own `/status` on its port: the node document carries a `registered` boolean next to `ready`. - **Watch `sessionDurationMillis`.** A session that has been alive far longer than the suite's tests is usually a leaked driver holding a slot, and it makes healthy capacity look like a shortage.

  • Every node reports status DRAINING and the queue keeps growing — what does that mean?
    `DRAINING` in the Selenium 4 Grid GraphQL `Status` enum means the node is finishing its current sessions and will accept no new ones; the enum's other values are `UP` and `DOWN`. Draining nodes still appear in `nodesInfo` and in the `/status` nodes array, so raw node counts look healthy while usable capacity is zero. It clears only when replacement nodes register.
  • How do you spot a node that started but never joined the grid?
    It is simply absent: the router's `/status` and `nodesInfo` report only what the distributor has registered, so an unregistered node leaves no trace. Query that node directly on its own port — the Selenium 4 node's `/status` document carries a `registered` boolean beside `ready`, which separates 'running but not joined' from 'not running'.
  • Why can grid sessionCount look healthy while a specific suite starves?
    `grid { sessionCount }` and `maxSession` are grid-wide totals across every stereotype. A grid can sit at half occupancy while every slot of one browser is taken, so the aggregate looks fine and one suite waits. Break the picture down per node with `nodesInfo { nodes { stereotypes sessionCount maxSession } }`.

saying these in an interview costs you the question

  • Adds more nodes when nothing advertises the requested browser
  • Reads a non-empty queue as proof the grid is broken
  • Checks only the /status ready flag without reading the request
  • Ignores DRAINING and DOWN nodes when counting usable capacity
  • Assumes a capability typo surfaces differently from plain saturation