skip to content

In Selenium, why can a CSS locator never match an element by its visible text?

level: middleimportance: must knowfreq 72%

answer

  1. Structure and attributes, not characters
  2. That pseudo-class came from jQuery
  3. It fails to parse, not to match
  4. Only anchors have a text strategy
  5. An attribute value is not rendered text

basics

~20 s

Because a CSS selector matches on tags, ids, classes, attributes, structural position and state, never on the characters inside an element. Selenium sends the string straight to the browser's querySelectorAll, so a locator gets exactly that power and no more.

solid answer

~40 s

`By.cssSelector` hands your string to the remote end under the `css selector` strategy, and the browser evaluates it with `querySelectorAll`. That engine matches on element type, id, class, attributes, structural position, state pseudo-classes, `:not()` and `:has()` — the text between the tags is not part of the matching surface. `:contains()` is remembered from a dropped CSS draft and from jQuery's own selector engine, so `button:contains('Resell')` is not merely unmatched, it fails to parse and returns the `invalid selector` error. The alternatives are `By.linkText` and `By.partialLinkText`, which match rendered text on `<a>` elements only, an XPath expression such as `//button[normalize-space()='Resell']` for any element, or a CSS search for candidates followed by filtering their text in your own code.

go deeper

for a junior

Be ready to say that a CSS locator matches tags, classes, ids and attributes, and that the visible text of an element is not among them. Naming By.linkText as the anchor-only exception is enough at this stage.

for a middle

Explain that the browser evaluates the string with querySelectorAll, so the locator has exactly that engine's matching surface, and that :contains() fails to parse rather than matching nothing.

for a senior

An interviewer expects you to choose deliberately between an attribute-based CSS locator, an XPath text predicate and client-side filtering of findElements results, and to explain what each costs in clarity and in round trips.

for a principal

Own the guidance about which facts the markup should expose as attributes so that most locators can remain plain CSS, and be explicit about where a suite accepts text matching as unavoidable.

## What a selector is allowed to test A CSS selector can test an element's tag name, its id, its classes, the presence or value of an attribute, its structural position among siblings (`:nth-of-type`, `:first-child`), its state (`:checked`, `:disabled`), a negation with `:not()` and, in engines that support it, a containment condition with `:has()`. That list is the whole surface. What is missing from it is the **character data between the tags** — the text a user actually reads. This is not a Selenium restriction. `By.cssSelector` sends your string to the remote end under the `css selector` strategy, the browser runs `querySelectorAll` with it, and the result is whatever the page's own engine can express. A locator gets exactly the power of `document.querySelectorAll` and not one feature more. So on a stadium ticket exchange, `button.resell` and `button[data-action="resell"]` are both fine, and "the button whose label reads Resell" is simply not a question CSS can be asked. ## Where `:contains()` came from An early working draft of the CSS Selectors module proposed a `:contains()` pseudo-class and it was dropped before the module became a recommendation. jQuery implemented it in its own selector engine, where it worked because jQuery was matching in JavaScript rather than asking the browser. A generation of page scripts used it, and the muscle memory outlived the API. The consequence in a test is concrete: `By.cssSelector("button:contains('Resell')")` reaches the browser, `querySelectorAll` cannot parse it, and the command comes back as the `invalid selector` error, surfacing in Java as `InvalidSelectorException`. It does not match nothing — it fails to run. The same is true of the text pseudo-classes borrowed from other tools' engines; if the browser did not ship it, the locator is not a locator. ## The three ways to reach text 1. **`By.linkText` and `By.partialLinkText`** are the only built-in strategies that match on rendered text, and the specification implements them by calling `querySelectorAll("a")`, reading each candidate's rendered text and comparing after trimming leading and trailing whitespace. That means they work on anchors and on nothing else — a `<button>` is out of reach whatever its label says. 2. **`By.xpath` with a text predicate** is the general fallback: `//button[normalize-space()='Resell']` matches on the element's string value, for any element, with whitespace folded. 3. **Find many in CSS, then filter in the client.** Locate the candidates with a CSS locator, read each one's text in your own code and keep the one you want. This is the honest choice when text is only one of several conditions, and it keeps the expensive part — the search — in the engine. ## Attribute values are not visible text This is where candidates slip. `button[aria-label="Resell seat 118-K"]` looks like text matching and is not: it compares an attribute value in the DOM, character for character, and that string may never be painted anywhere. The same applies to `option[value="118-K"]` versus the option's visible label, and to `[title]`, `[alt]` and `[placeholder]`. All of them are ordinary attribute selectors, and all of them are available to a CSS locator. | Locator | Matches on | Applies to | |---|---|---| | `By.cssSelector("button[aria-label='Resell']")` | an attribute value in the DOM | any element carrying the attribute | | `By.linkText("Resell")` | rendered text, exact after trimming | `<a>` elements only | | `By.xpath("//button[normalize-space()='Resell']")` | the element's rendered string value | any element | ## Where the boundary actually sits - The engine walks the element tree and reads attributes; the text nodes inside an element are content it never inspects when matching. - A pseudo-class that appears to match text — `:contains()`, or a text pseudo-class from another tool — is not CSS, and the browser's answer to it is a parse failure, not an empty result. - `By.linkText` compares the whole trimmed text; `By.partialLinkText` compares a substring. Neither takes a regular expression, and neither applies outside `<a>`. - Generated content produced by a stylesheet is not text either. It is not in the DOM at all, so no locator of any kind matches it. - If the markup already exposes the same fact as an attribute — a status, an action name, an accessible label — the locator can stay in CSS and never touch text. The compact way to say all of this in an interview: CSS matches the document's structure and its attributes, XPath can additionally match its rendered characters, and Selenium's only built-in text strategies are the two link-text ones. Everything else you have seen that appears to match text in a CSS string came from a different selector engine.

  • Does button[aria-label='Resell'] count as matching on text?
    No. That is an ordinary attribute selector comparing a value stored in the DOM, character for character, and the value may never be painted on screen. It happens to read like a label, but the engine is matching an attribute, not the element's rendered content.
  • How do By.linkText and By.partialLinkText actually work, and what are they limited to?
    The specification defines them as a `querySelectorAll("a")` call followed by a comparison of each candidate's rendered text, trimmed of leading and trailing whitespace — exact for `linkText`, substring for `partialLinkText`. They therefore work on anchor elements only; a `<button>` with the same label is out of reach.

The selector engine reads the building's wiring diagram, not its signage: it can tell you which box is the third seat cell in a sold-out row, but never what word is printed on it.

saying these in an interview costs you the question

  • Believes :contains() works in a CSS locator because jQuery had it
  • Says an unsupported text pseudo-class just matches nothing
  • Thinks By.linkText matches text on any element
  • Confuses an attribute value such as aria-label with visible text
  • Expects a CSS locator to match generated content from a stylesheet