skip to content

In Selenium XPath locators, why is ends-with() unavailable when starts-with() works fine?

level: middleimportance: nice to knowfreq 28%

answer

  1. Which engine evaluates the expression?
  2. The function library has a ceiling
  3. XPath 1.0 versus later versions
  4. starts-with shipped; the partner came later
  5. substring plus string-length emulates it

basics

~10 s

Browsers evaluate locators with document.evaluate, which implements XPath 1.0 only, and ends-with() arrived in XPath 2.0. Selenium 4 offers no way to raise that version, so you emulate it with substring and string-length.

solid answer

~40 s

Selenium 4 does not ship its own XPath engine: the remote end passes the string to the browser's `document.evaluate`, which implements DOM Level 3 XPath — that is, **XPath 1.0**. `starts-with()` is in the 1.0 core library; `ends-with()`, `matches()`, `lower-case()` and `replace()` are not, and no capability or flag raises the level. An unknown function is a compile failure in the evaluator, so `By.xpath` surfaces it as `InvalidSelectorException` rather than an empty result. The workaround is `substring()` plus `string-length()`: `//td[@class='slip'][substring(., string-length(.) - 4) = 'Dock1']` tests a five-character suffix, since `substring()` is 1-based. Case-insensitive matching uses `translate()` to fold letters one by one. The same ceiling explains `normalize-space()` and `contains()` being the workhorses of practical chart locators.

code

java · 25 lines
java
import java.util.List;
import org.openqa.selenium.By;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.WebElement;
import org.openqa.selenium.chrome.ChromeDriver;

public class SlipSuffixLookup {
  private static final String UPPER = "ABCDEFGHIJKLMNOPQRSTUVWXYZ";
  private static final String LOWER = "abcdefghijklmnopqrstuvwxyz";

  public static void main(String[] args) {
    WebDriver driver = new ChromeDriver();
    try {
      driver.get("https://harbour.example/moorings");
      List<WebElement> endsWithDock1 = driver.findElements(
          By.xpath("//td[@class='slip'][substring(., string-length(.) - 4) = 'Dock1']"));
      List<WebElement> reserved = driver.findElements(By.xpath(
          "//td[@class='status'][translate(normalize-space(), '" + UPPER + "', '" + LOWER
              + "') = 'reserved']"));
      System.out.println(endsWithDock1.size() + " " + reserved.size());
    } finally {
      driver.quit();
    }
  }
}

go deeper

for a junior

Know the handful of string functions you can actually use — contains, starts-with and normalize-space — and that copying an ends-with example from a reference will fail. That alone avoids most wasted debugging.

for a middle

Explain that the browser's document.evaluate implements XPath 1.0 and that the version is not configurable. Be able to write the substring plus string-length suffix test and the translate case-folding idiom from memory.

for a senior

Recognise an InvalidSelectorException as a compile failure in the evaluator rather than a missing element, and know that the result of a locator must be elements, so functions belong inside predicates only.

for a principal

Be ready to say when a locator has grown complicated enough that the answer is to change the markup or the data attributes rather than to keep emulating missing string functions across a suite.

## What the browser actually implements WebDriver's `xpath` location strategy is not a Selenium-side XPath engine. The remote end evaluates the string by calling the browser's own `document.evaluate`, which implements **DOM Level 3 XPath** — and DOM Level 3 XPath is **XPath 1.0**. There is no switch, no capability and no flag that raises it. Whatever XPath 1.0 lacks, a Selenium locator lacks, in every language binding. `ends-with()` is the sharpest example. It reads like a natural partner to `starts-with()`, it is in every reference you find by searching, and it does not exist in 1.0 — it arrived in XPath 2.0. Feed it to `By.xpath` and the expression is not merely unmatched: it fails to compile in the evaluator, and the Java binding surfaces that as `InvalidSelectorException`. ## The XPath 1.0 string library | Available in XPath 1.0 | Not available (2.0 or later) | |---|---| | `contains()`, `starts-with()` | `ends-with()` | | `substring()`, `substring-before()`, `substring-after()` | `matches()` — regular-expression matching | | `string-length()`, `normalize-space()` | `lower-case()`, `upper-case()` | | `concat()`, `translate()`, `string()` | `replace()`, `tokenize()` | | `last()`, `position()`, `count()`, `not()` | sequence functions of any kind | Everything in the left column is fair game inside a mooring-chart predicate. Everything in the right column has to be emulated or abandoned. Two entries deserve a note. `contains()` is the workhorse the missing functions push you toward, and it matches a substring **anywhere** in the value, so `contains(., 'A-12')` also matches a slip labelled `A-120`. And `count()` is available as a function but, as the last section explains, only inside a predicate — `//tr[count(td) = 4]` is a valid locator while `count(//tr)` is not. ## Emulating `ends-with()` The idiom takes the tail of the string and compares it. To ask whether a slip label ends with `Dock1`: 1. Measure the target: `Dock1` is five characters. 2. Take the last five characters of the value: `substring(., string-length(.) - 4)`. XPath's `substring()` is **1-based**, so the start offset for the last *n* characters is `string-length(.) - n + 1`. 3. Compare: `//td[@class='slip'][substring(., string-length(.) - 4) = 'Dock1']`. Written generically the shape is `substring($s, string-length($s) - string-length($t) + 1) = $t`. It is verbose, but it is the whole workaround and it is stable. ## Case-insensitive matching with `translate()` `translate()` maps characters one-for-one from a source set to a replacement set — it is the only case-folding tool 1.0 gives you: - `translate(., 'ABCDEFGHIJKLMNOPQRSTUVWXYZ', 'abcdefghijklmnopqrstuvwxyz')` lower-cases the context node's string value. - Wrap the comparison around it: `//td[@class='status'][translate(., 'ABCDEFGHIJKLMNOPQRSTUVWXYZ', 'abcdefghijklmnopqrstuvwxyz') = 'reserved']` matches `Reserved`, `RESERVED` and `reserved` alike. - It is character-mapped, not locale-aware, so any letter you want folded has to be listed in both strings explicitly. ## `normalize-space()` and the whitespace it removes `normalize-space()` takes a string, strips leading and trailing whitespace, and collapses every internal run of whitespace to a single space. Whitespace here means space, tab, carriage return and line feed. Called with no argument it operates on the **string value of the context node** — the concatenation of all its descendant text. That makes it the right default for matching rendered chart text, because markup like `<td class="boat">\n Sea Wren\n </td>` has a text node of `"\n Sea Wren\n "`. `//td[text()='Sea Wren']` fails against it; `//td[normalize-space()='Sea Wren']` succeeds. Two related distinctions: - `text()` selects the text-node **children** of an element, so `contains(text(), 'Wren')` converts a node-set to a string by taking only the **first** text node — it misses text that follows a nested `<span>`. - `.` inside a predicate is the context node itself, and converting it to a string concatenates **all** descendant text, so `contains(., 'Wren')` sees through nested markup. ## The other ceiling: the result must be elements XPath 1.0 can return four types — node-set, string, number and boolean — but the WebDriver location strategy accepts only elements. The remote end evaluates the expression as an ordered node snapshot and checks every node it got back; if any of them is not an element it answers with the `invalid selector` error rather than a result. So a function may appear **inside** a predicate but never as the whole expression: - `count(//tr)` returns a number — `InvalidSelectorException`, and the Java `InvalidSelectorException` javadoc names this exact case. - `//td[@class='boat']/text()` returns text nodes — rejected for the same reason. - `//td/@class` returns attribute nodes — also rejected. - `//td[@class='boat']` returns elements — accepted. To read text or an attribute, find the element and then call `WebElement.getText()` or `WebElement.getDomAttribute()`; the expression's job stops at selecting the element.

  • What happens when a Selenium XPath locator evaluates to a number instead of elements?
    The find fails with `InvalidSelectorException`. The remote end evaluates the expression as an ordered node snapshot and rejects any result that is not made of elements, so `count(//tr)`, `//td/text()` and `//td/@class` are all refused. Select the element, then read it with `WebElement.getText()` or `WebElement.getDomAttribute()`.
  • How do you match a status cell case-insensitively in XPath 1.0?
    Fold the case with `translate()`, which maps characters one-for-one: `//td[translate(., 'ABCDEFGHIJKLMNOPQRSTUVWXYZ', 'abcdefghijklmnopqrstuvwxyz') = 'reserved']`. There is no `lower-case()` in 1.0, and `translate()` is not locale-aware, so any letter you need folded must appear in both argument strings explicitly.
  • What exactly does normalize-space() do to a cell's text?
    It strips leading and trailing whitespace and collapses each internal run of whitespace — spaces, tabs, carriage returns and line feeds — into a single space. Called with no argument it works on the string value of the context node, which is why `//td[normalize-space()='Sea Wren']` matches markup that `//td[text()='Sea Wren']` misses.

saying these in an interview costs you the question

  • Assumes a Selenium capability or flag can enable XPath 2.0
  • Uses matches() or lower-case() and expects them to evaluate
  • Thinks an unknown XPath function silently returns no matches
  • Believes substring() in XPath is zero-based like Java's
  • Writes count(//tr) as a locator expecting a numeric result