skip to content

When sizing a Selenium Grid, why is new-session queue depth a better capacity signal than Node count?

level: principalimportance: nice to knowfreq 27%

answer

  1. Supply versus demand, not supply alone
  2. A ceiling nobody reaches is invisible
  3. Where requests wait when slots are busy
  4. Depth and time queued, not slot totals
  5. Bounded by --session-request-timeout

basics

~20 s

Node count is a static ceiling; queue depth measures whether that ceiling is ever reached. A grid with sixty idle slots is no busier than one with six, so depth is what tells you when to add capacity.

solid answer

~40 s

Slot and Node counts describe supply - the maximum number of browsers that can run at once - and they read the same whether the Grid is idle or saturated. The New Session Queue is where a request waits when no matching slot is free, so its depth and the time requests spend in it describe demand: whether that supply ceiling actually binds. If the photo-album uploader suite's queue is empty, adding Nodes buys nothing; if it is persistently non-empty and waits climb toward `--session-request-timeout` (300 seconds by default), the ceiling binds and more slots will help. Depth must still be read against capability, because requests that no registered stereotype can ever match sit in the queue and inflate depth without representing a shortage of slots.

go deeper

for a junior

Be ready to say what the New Session Queue is: the place a new-session request waits when Selenium Grid has no free matching slot. Knowing that such a queue exists at all is the floor here.

for a middle

Explain why slot count and queue depth measure different things - one is the supply ceiling, the other is evidence that the ceiling binds - and name --session-request-timeout as the bound on any single wait.

for a senior

Show how you decide from evidence. A persistently non-empty queue with climbing waits justifies more slots, an empty queue does not, and an unmatchable capability inflates depth without meaning shortage.

for a principal

Own the tradeoff. Argue what queue depth you are willing to run at, how it trades against idle browser cost, and why arrival rate is a suite-side input you plan around rather than a Grid setting you control.

## Two numbers that answer different questions A **Selenium Grid** turns a new-session request into a running browser by finding a free **slot** - one browser's worth of capacity on a registered **Node**. Node count, and the slot count it implies, is a **supply** figure. It is fixed by how the Grid was provisioned, it reads the same at three in the morning as it does during the busiest run of the day, and on its own it says nothing about whether that supply was ever exhausted. The **New Session Queue** is the component a request enters when no matching slot is free. Its **depth** - how many requests are parked in it - and the time each request spends there are **demand** figures. They only become non-zero when supply ran out, which is exactly the event you are trying to size against. - Slot count answers *how many browsers could run at once*. - Queue depth answers *how often that ceiling was actually reached*. - Time in queue answers *how badly it was reached when it was*. - Only the last two move when the photo-album uploader suite gets busier. ## The photo-album uploader, sized two ways Suppose the photo-album uploader suite exercises album creation, batch upload, thumbnail rendering and share links. Sized on supply alone the conversation is "we have twelve Chrome slots and ninety cases, so twelve is obviously not enough." Sized on demand it becomes "the queue peaked at four requests for ninety seconds during the upload cases and was empty for the rest of the run." Those lead to opposite decisions. The first buys eight more Nodes; the second concludes that twelve slots absorb the suite with a short, bounded peak and that the budget belongs elsewhere. Arrival rate - how many cases the suite fires at once - is decided on the suite side and is an **input** to this decision rather than something the Grid controls. What the Grid gives back is the honest measurement of what that arrival rate cost. ## Reading the signals side by side | Signal | What it is | When it changes | What it justifies | |---|---|---|---| | Node and slot count | Static supply ceiling | Only when you reprovision | Nothing on its own | | Queue depth | Requests waiting for a slot | Whenever demand exceeds supply | More slots, if sustained | | Time in queue | How long each wait lasted | With contention and session length | Urgency, and the timeout budget | | `--session-request-timeout` | The bound on any single wait | Only when you change it | The failure point, not capacity | ## What a deep queue does not automatically mean Depth is evidence, not a verdict. Three different situations produce a non-empty queue, and only one of them is a shortage of slots: 1. **Genuine contention.** Every matching slot is busy and requests are waiting their turn. Depth rises and falls with the run and drains as sessions end. More slots help. 2. **An unmatchable request.** The requested capabilities match no registered **stereotype** at all, so the request waits out the whole `--session-request-timeout` and then fails with `session not created`. It inflates depth for five minutes and no amount of extra hardware drains it. 3. **Long sessions rather than many.** A few sessions hold their slots for a long time, so the queue is deep even though arrival rate is modest. Shortening the sessions is the cheaper fix. Confirm which of the three you are looking at before treating depth as a buy signal. That is why "queue depth is the honest signal" is a claim about *what to measure*, not a promise that one number decides the answer. ## Where the two queue flags fit the plan `--session-request-timeout` (seconds, **default 300**) is how long a request may sit before the Grid gives up on it. `--session-retry-interval` (milliseconds, **default 15**) is how often the **Distributor** rechecks the queue for requests it can now place. Neither creates capacity. What the timeout does give you is a **budget**: it states, as policy, the longest queue wait the platform is willing to hide from a suite. Sizing then reduces to a testable question - does observed time in queue stay comfortably inside that budget at peak? ## Planning rules that follow 1. Never justify new Nodes from a slot count alone; justify them from a queue that is repeatedly non-empty at peak. 2. Treat an empty queue as proof that extra slots would idle, whatever the slot count happens to be. 3. Read depth together with time in queue - a queue of ten that drains in a second is a different problem from a queue of two that sits for four minutes. 4. Rule out unmatchable capabilities before buying hardware, because on a depth graph they look identical to contention. 5. Decide the acceptable wait first, express it as `--session-request-timeout`, and size so that peak waits land well inside it.

  • The queue is empty but the photo-album suite still finishes slowly. What does that tell you?
    That new-session capacity is not the bottleneck. An empty queue means every request found a slot without waiting, so the time is going somewhere else - the pages under test, the steps themselves, or the shape of the run. Adding Nodes to a Grid whose queue never fills changes nothing except idle browser cost.
  • Queue depth is zero all day and spikes to forty for ten minutes each morning. Add slots?
    Only if the spike costs more than the idle capacity would. A ten-minute burst absorbed inside `--session-request-timeout` is the queue doing its job, smoothing a peak so you do not buy browsers that idle the rest of the day. Add slots when the spike breaches the timeout, or when the wait itself is unacceptable.
  • Why can queue depth mislead you when a requested capability is wrong?
    Because a request whose capabilities no registered stereotype can satisfy sits in the queue exactly like one that is merely waiting its turn. It inflates depth and then fails with `session not created` after the full `--session-request-timeout`. Depth answers how many are waiting, not how many could ever be served.

A photo lab with twelve printing machines and nobody at the counter is not busier than one with four. The line at the counter, not the machine count, is what tells you whether to buy a fifth.

saying these in an interview costs you the question

  • Claims more Nodes always shorten the wait without checking whether the queue is ever non-empty
  • Treats slot count as a utilisation figure rather than a static provisioning ceiling
  • Assumes a deep queue always means too few slots, never an unmatchable capability
  • Believes raising --session-request-timeout adds capacity instead of only lengthening the wait
  • Sizes the grid from Node count alone and never looks at time spent queued