skip to content

A Selenium test reading 300 review-queue rows is slow, and swapping every XPath for CSS barely helped. Why?

level: seniorimportance: must knowfreq 57%

answer

  1. Check which term you are optimising
  2. The engine is the smaller of two costs
  3. Two commands per row, not one
  4. Six hundred round trips before any matching
  5. Count commands before rewriting locators

basics

~20 s

Because the time is going into the number of commands, not the matching engine. Three hundred finds plus three hundred text reads is six hundred round trips, and a faster in-page match saves only microseconds on each one.

solid answer

~40 s

The engine choice is the small term. Each `findElement` and each read such as `getText` is its own HTTP command to the remote end, so a loop over 300 rows of a scholarship-application review queue is 600 round trips before the browser matches anything at all. Rewriting `By.xpath` as `By.cssSelector` changes only what happens inside each of those commands, and on an ordinary page that match finishes far inside the request that carried it. The fix is to cut commands: one `findElements` for the row set instead of 300 singular finds, and read only the fields you actually assert on. Measure first - time the loop, change one thing, re-time - rather than rewriting 300 locators on a hunch and getting no signal back.

code

java · 13 lines
java
long start = System.nanoTime();
List<String> perRow = new ArrayList<>();
for (int i = 1; i <= 300; i++) {
    perRow.add(driver.findElement(
        By.cssSelector("#reviewQueue tbody tr:nth-child(" + i + ") td.applicant")).getText());
}
System.out.printf("600 commands: %d ms%n", (System.nanoTime() - start) / 1_000_000);

start = System.nanoTime();
List<String> batched = driver.findElements(By.cssSelector("#reviewQueue td.applicant"))
    .stream().map(WebElement::getText).toList();
System.out.printf("301 commands for %d rows: %d ms%n",
    batched.size(), (System.nanoTime() - start) / 1_000_000);

go deeper

for a junior

Be ready to say that a slow loop usually means too many calls to the browser rather than a bad selector. Counting how many calls the loop makes is the first thing to check.

for a middle

Explain the arithmetic out loud: one command per find, one per read, so a 300-row loop is 600 commands. Then show which of those you can collapse into a single plural find.

for a senior

Show a measurement habit. Change one thing at a time and demonstrate that selector syntax barely moves the number while command count moves it a lot, then name the cases where the engine really does matter.

for a principal

Set the team's mental model: the round trip is the unit of cost, and guidance framed around selector syntax will not survive a profile. Decide what evidence a performance claim must carry before it becomes a rule.

## Where the time in a 300-row loop actually goes The rewrite did not help because it addressed the smaller of the two terms. A Selenium test's cost for reading a list decomposes into: - **Command count** - every find and every read is its own HTTP request to the remote end, with JSON encoding on both sides and the driver's per-command bookkeeping around it. - **In-page evaluation** - the browser matching your selector once per command, using `querySelectorAll` for the CSS strategy or `evaluate()` for the XPath strategy. Swapping `By.xpath` for `By.cssSelector` changes the second term only. On a scholarship-application review queue with a few thousand nodes and reasonable expressions, that term is far smaller than the request that carried it, so the total barely moves. ## The two terms, side by side | | What it is | How it scales | What changes it | |---|---|---|---| | Command count | HTTP requests to the remote end | linear in finds plus reads plus interactions | collapsing N finds into one plural find | | In-page evaluation | one match per command | with candidate nodes and predicate work | anchoring a scan, dropping text predicates | For 300 rows read with one find and one `getText` each, the command count is 600 before the browser matches anything. The engine choice is being asked to compete with 600 round trips. ## How to prove it in ten minutes 1. Time the existing loop end to end and write the number down. 2. Replace the 300 per-row `findElement` calls with a single `findElements` for the whole set, leave the reads alone, and re-time. Command count drops from 600 to 301. 3. If the wall-clock time falls roughly in proportion to the commands removed, the round trip is your cost and you are done diagnosing. 4. Only then vary the locator syntax on its own, with the command count held constant, and see whether it moves anything. Varying one thing at a time is the whole method. The failed rewrite in the scenario changed 300 locators at once and produced no signal, which is exactly what you expect when you optimise the term that was not dominating. ## What actually removes commands - Replace a per-row singular find with one plural `findElements`; one request returns the whole set of references. - Do not re-find an element you already hold, unless the page has re-rendered it out from under you. - Do not read fields you are not asserting on. Each `getText`, `getDomAttribute` or `isDisplayed` is another request. - Narrow what the test needs: reading three fields from 300 rows to assert on five of them is 900 requests spent to use fifteen. ## When the engine choice does earn its rewrite The CSS-versus-XPath difference stops being noise in a small number of real situations: - The expression is doing pathological work - a wildcard descendant scan whose predicate builds each candidate's subtree text - so the in-page term grows past the round trip on its own. - The DOM is very large, so even a well-shaped expression walks a great deal of it. - The find sits inside a polling wait that re-issues it several times a second, multiplying the in-page term without adding proportionate round trips. In each of those the fix is usually to reshape the expression rather than to change syntax wholesale, because a badly shaped CSS selector such as a long descendant chain over a big table can cost more than a tightly anchored XPath. ## The habit to take away Two claims are worth being sceptical of in a Selenium performance discussion. The first is "XPath is slow", offered without a measurement; the cost model says the engine matters far less than the number of times you cross the process boundary. The second is "this loop is slow because the page is slow"; if the page renders fine for a human, and the loop issues six hundred commands, the loop is the story. The useful instinct is to count commands from the source before touching a profiler. It takes a minute, it is exact, and it points at the change that will actually move the number: fewer requests, not different selector syntax.

  • How would you confirm that command count, not matching, is the bottleneck?
    Time the loop, then collapse the 300 singular finds into one plural find and re-time with everything else unchanged. If wall-clock time falls roughly in proportion to the commands removed, and barely moves when you vary selector syntax on its own, the round trip is the cost.
  • When does swapping XPath for CSS actually pay for itself?
    When the expression is doing pathological work, such as a wildcard descendant scan with a subtree-text predicate over a very large DOM, or when the same find sits inside a polling wait that re-issues it several times a second. Then the in-page term stops being the small one.
  • Does running against a remote end over a network change the arithmetic?
    It pushes it further the same way. Every command still costs one request, and a request that leaves the machine costs more than a loopback one, so the round trip becomes an even larger share of the total. The remedy is unchanged: send fewer commands.

saying these in an interview costs you the question

  • Blames XPath for a loop that issues hundreds of separate commands
  • Rewrites every locator before measuring where the time actually goes
  • Counts the finds but forgets each getText is another round trip
  • Believes a shorter selector string makes the request itself faster
  • Assumes the browser, not the command count, is what got slower