skip to content

Every Selenium 4 command is a blocking HTTP round trip to the remote end - how does that shape a large suite?

level: principalimportance: must knowfreq 60%

answer

  1. One request per action
  2. Nothing arrives unless you ask
  3. Latency multiplied by command count
  4. Waits are request loops
  5. One execute/sync replaces many finds

basics

~20 s

Selenium's command protocol is strictly request-response, so every find, click and read costs a full round trip and nothing arrives unasked. Budget commands per case, keep the remote end close, and expect all waiting to be client-side polling.

solid answer

~50 s

Selenium's local end sends one HTTP request per command and blocks for the response; the remote end never pushes anything back on that contract. So a case is not measured in assertions but in round trips - a single screen check is easily thirty to sixty requests, and each wait poll is another one. That makes three things design decisions rather than details: how far the remote end sits from the runner, since latency is paid per command; whether a bulk read can be collapsed into one `POST /session/{session id}/execute/sync` instead of hundreds of finds, accepting that a scripted read no longer proves a user could see the data; and how many sessions run in parallel, since sessions are independent and concurrency moves wall clock far more than shaving commands. On a local driver with a short case, none of it matters.

go deeper

for a junior

Be ready to say that each Selenium call is its own request to the driver and that the test waits for each reply. Knowing that nothing comes back on its own is enough at this stage.

for a middle

Explain the mechanics: one command, one request, one response, and a wait implemented as repeated requests. Be able to walk a short case and count roughly how many requests it produces.

for a senior

Show that you have profiled a real suite. Say where the time actually went, which commands you removed, and how you proved the change helped rather than assuming that fewer calls must be faster.

for a principal

Own the tradeoff. Argue about where browsers run relative to the runners, how much bulk reading may be collapsed before a test stops proving anything, and why buying parallel sessions usually beats micro-optimising the command stream.

## The contract you are budgeting against In Selenium 4 the **local end** — the client library running inside your test process — talks to the **remote end**, the browser driver, over a strict request-response protocol carried on HTTP. The local end sends one request, blocks, and reads one response. On that contract there is no channel over which the remote end can announce that a page finished rendering, that a row appeared, or that the browser died. Every fact a test knows about the browser, it asked for by name. Three consequences fall out, and a lead is expected to plan around all three: 1. **Latency is charged per command, not per test.** A `findElement` call is one request; the `click()` on the element it returned is a second; reading the resulting text is a third. 2. **Waiting is polling.** A wait loop is a burst of repeated `POST /session/{session id}/element` requests, not a subscription, and every poll is a real round trip. 3. **The path multiplies everything.** Round-trip time is paid once per command, so a 40 ms path across 400 commands is 16 seconds of wall clock inside one case. ## Counting the round trips in a real case Take the fleet-maintenance scheduler's overdue-vehicles screen. A case that opens the depot list, filters to trucks past their service interval, opens the first work order and asserts three fields is not four operations on the wire. It is one `POST /session`, one `POST /session/{session id}/url`, a find plus a `click` per control, a find plus a text read per assertion, and a closing `DELETE /session/{session id}` — comfortably thirty to sixty requests before the case has done anything interesting. Multiply by the poll count of every wait and ten readable lines of Java become several hundred requests. That number, not the number of assertions, is the unit a suite's wall clock is denominated in, and it is the number worth putting on a slide when someone asks why the nightly run takes ninety minutes. ## The three levers, and what each is worth | Lever | What it removes | What it costs you | |---|---|---| | Shorten the path to the remote end | The per-command latency multiplier | Co-location constraints on where browsers run | | Collapse many reads into one `POST /session/{session id}/execute/sync` | N-1 round trips per batch | The read runs as page script, so it no longer proves a user could see the value | | Delete commands the case never needed | Whole requests, at no risk | Nothing, when the commands were incidental navigation | - **Shortening the path** is the only lever with no test-design cost. A remote end reached across a wide-area link, or through an extra intermediary hop, pays the same tax on every command of every case. - **Collapsing reads** is powerful and dangerous. Pulling three hundred vehicle due-dates back through one scripted call turns three hundred round trips into one, but a scripted read skips everything the browser would have made a real dispatcher do to see those dates. - **Deleting commands** usually means noticing that a case clicks through five screens to reach a state one `POST /session/{session id}/url` could open directly. - **Parallel sessions** move the wall clock further than any of these. Sessions are independent and each carries its own id in the path, so widening the browser fleet is the first answer and command-count surgery is the second. ## What a collapsed read looks like on the wire ``` POST /session/8f3c/execute/sync {"script":"return [...document.querySelectorAll('[data-due]')].map(e => e.dataset.due)","args":[]} {"value":["2026-03-01","2026-03-04","2026-03-11"]} ``` One request, one `value` array, three hundred elements. The same information gathered element by element is three hundred finds and three hundred text reads, and each one of those six hundred requests waits for the previous response before it is even written. ## Where the budget does not matter - A twenty-command case against a driver on `localhost` spends far more time inside the browser's own rendering than inside the protocol; tuning it is waste motion. - Batching and deep-linking earn their keep when a suite is counted in thousands of cases, or when the remote end deliberately lives somewhere else. - A single slow page load dwarfs dozens of round trips, so profile before you rewrite: measure command count per case first, then decide. ## What the interviewer is listening for The answer that lands treats command count as a first-class budget. You know a wait is a request loop rather than a subscription, you can estimate the round trips a representative case costs, you can say which reads may honestly be collapsed and which must stay as a user would perform them, and you would rather buy concurrency than shave individual commands. The answer that fails claims the protocol batches commands for you, or that the driver streams browser state back so the test can simply react to it.

  • Where does collapsing many reads into one scripted command stop being a legitimate optimisation?
    When the collapsed read is the thing under test. Pulling every vehicle due-date back through one `execute/sync` call proves the data is in the DOM, not that a dispatcher could reach it - scrolling, overlays and disabled controls are all skipped. Collapse reads that are setup or bulk assertion data; leave the user-visible path as real commands.
  • If nothing is pushed back, how does a test ever notice that the browser crashed mid-case?
    Only on the next command. The local end discovers it when a request comes back as a failure - typically an `invalid session id` or a no-such-window style error - or when the request itself cannot be completed. There is no crash notification on the command protocol, so a test hanging in a wait loop learns nothing until it times out or the next request returns.
  • Why does adding a network hop between the local end and the remote end hurt more than a slower browser?
    Because the hop is charged on every command while browser slowness is charged on the few commands that actually do work. An extra 20 ms of path across four hundred commands is eight seconds per case, and a suite pays it in every case, in every run, forever.

saying these in an interview costs you the question

  • Thinks the client batches several commands into one request
  • Believes the driver streams page state back as it changes
  • Treats a wait as a subscription rather than a request loop
  • Optimises command count before measuring where time goes
  • Assumes a remote end across a slow link costs the same as a local one