skip to content

In Selenium 4, why can By.id("recall.due") find an element that By.cssSelector("#recall.due") cannot?

level: seniorimportance: should knowfreq 31%

answer

  1. One of them is generated, one is typed
  2. Something happens to the value first
  3. CSS punctuation has meaning the id does not
  4. A bare dot means a class, not text
  5. Leading digits get a code-point escape

basics

~10 s

Escaping. Selenium rewrites By.id into a CSS selector but escapes the value first, so the dot stays a literal character of the id. A hand-written selector reads that same dot as a class instead.

solid answer

~40 s

In Selenium 4 `By.id` has no wire strategy of its own, so the client builds a CSS fallback from the format `#%s`. Before formatting it runs the id through an escape pass that backslash-escapes whitespace, quotes, `#`, `.`, `:`, `-`, `[`, `(` and the rest of the CSS metacharacters, and rewrites a leading ASCII digit as a code-point escape, so `By.className("5foo")` serialises to `.\35 foo`. `By.id("recall.due")` therefore goes out as `#recall\.due`, asking for one element whose id is literally `recall.due`. Hand-writing `By.cssSelector("#recall.due")` skips that pass: CSS reads the bare dot as a class selector, so the query means id `recall` carrying class `due` and quietly matches nothing.

code

java · 13 lines
java
import org.openqa.selenium.By;
import org.openqa.selenium.json.Json;

public class RecallIdEscaping {
  public static void main(String[] args) {
    Json json = new Json();
    String[] ids = {"recall.due", "recall:2026", "2026-recall", "one two"};
    for (String id : ids) {
      System.out.println(id + " -> " + json.toJson(By.id(id)));
    }
    System.out.println(json.toJson(By.name("50%off")));
  }
}

go deeper

for a junior

Remember that By.id handles awkward id characters for you, so reach for it rather than hand-writing the hash form when an id contains dots, colons or digits at the front.

for a middle

Explain that By.id has no wire keyword and is built as a CSS fallback, and that the value is escaped before it is formatted into that selector. Naming a couple of escaped characters is enough.

for a senior

Interviewers want the diagnosis: a locator that silently matches nothing on ids emitted by a server-side template, traced to unescaped CSS punctuation rather than to timing or visibility.

for a principal

Own the guidance for tooling that generates locators: anything that builds selector strings from application data needs an escaping contract, and borrowing the client's is safer than inventing one.

## Why By.id turns into CSS in the first place `id` is not one of the five keywords the **W3C WebDriver** specification lists in its table of location strategies, so Selenium 4 cannot send it. `By.id` is instead built on the client's internal `PreW3CLocator` base class, which eagerly constructs a private `ByCssSelector` **fallback** from the format string `#%s`. `By.className` does the same with `.%s`. If the client simply pasted the raw id into that format string, every id containing a CSS metacharacter would turn into a *different query*. So before formatting, the value goes through an **escape pass**. ## What the escape pass covers The client backslash-escapes every occurrence of a character in this set: - whitespace, single quote, double quote and backslash; - the selector punctuation `#`, `.`, `:`, `;`, `,`, `!`, `?`, `+`, `<`, `>`, `=`, `~`, `*`, `^`, `$`, `|`, `%`, `&`, `@`, backtick, `{`, `}`, `-`, `/`, `[`, `]`, `(` and `)`. Escaping is unconditional, so even a harmless character is escaped: `By.id("recall-due")` serialises to `#recall\-due`, and `\-` is simply a valid CSS escape for a hyphen. The pass is not trying to be minimal, it is trying to be safe. ## The leading-digit rule A CSS identifier may not begin with a digit, so a leading **ASCII** digit is rewritten as a code-point escape rather than a backslash escape. `By.className("5foo")` serialises with the value `.\35 foo` — the escape is the hexadecimal code point of the character, and the trailing space terminates the escape sequence rather than acting as a descendant combinator. The rule is deliberately narrow. A non-ASCII digit such as U+0665, the Arabic-Indic five, is already a legal CSS identifier-start code point, so it is passed through untouched; escaping it would collide with the escape for a plain ASCII `5`. ## The same id, two very different selectors | You write | The client sends | A hand-written equivalent | |---|---|---| | `By.id("recall.due")` | `#recall\.due` | `By.cssSelector("#recall.due")` means id `recall` **and** class `due` | | `By.id("recall:2026")` | `#recall\:2026` | `By.cssSelector("#recall:2026")` is parsed as a pseudo-class | | `By.id("2026-recall")` | `#\32 026\-recall` | `By.cssSelector("#2026-recall")` is an invalid identifier | | `By.id("one two")` | `#one\ two` | `By.cssSelector("#one two")` means a `two` element inside id `one` | Every row on the right is a real query that either fails to parse or asks a different question, and none of them raises a complaint that names the id. That asymmetry is the whole point of the question: `By.id` is not merely a shorthand for the hash form, it is a *correct* one. For a dental-practice recall list whose ids come out of a server-side template — `patient.4821.recall`, `recall:overdue`, `2026-q1-recall` are all shapes real frameworks emit — this is not a curiosity. It is the difference between a working locator and a silent no-match. ## By.name escapes differently `By.name` builds `*[name='...']`, an attribute selector with a **quoted** value rather than an identifier, so it needs quoting rules, not identifier escaping: 1. A backslash in the value is doubled, so `By.name("a\b")` serialises with the value `*[name='a\\b']`. 2. A single quote is backslash-escaped so it cannot terminate the quoted string. 3. A percent sign is doubled internally, because the format string is fed to Java's `String.format`; `By.name("50%off")` serialises to `*[name='50%off']` with the percent intact and no format exception. The practical upshot is the same: hand-writing the attribute selector puts the quoting burden on you, while `By.name` handles it. ## What escaping does not do for you - **It does not validate the id.** An id that does not exist still simply matches nothing; escaping makes the selector *ask the right question*, not succeed. - **It does not apply to `By.cssSelector` or `By.xpath`.** Those two carry your expression verbatim to the remote end, because the expression is the whole point of the strategy. - **It does not change the keyword the locator declares.** `getRemoteParameters()` on `By.id("recall.due")` still reports `id` and the *unescaped* value; only the CSS fallback carries the escaped form. - **It does not help with duplicated ids.** Escaping is about parsing, not about how many elements come back. - **It is not a promise of stability.** The escape set is client implementation detail. Rely on the behaviour, which is that `By.id` addresses the literal id, not on the exact serialisation.

  • What happens to an id or class name that starts with a digit?
    The client rewrites the first character as a CSS code-point escape, because a CSS identifier may not begin with a digit. `By.className("5foo")` serialises to `.\35 foo`, where the trailing space terminates the escape rather than acting as a combinator. Non-ASCII digits such as U+0665 are already valid identifier-start code points and pass through untouched.
  • Does By.name escape its value the same way?
    Not quite. `By.name` builds `*[name='...']`, a quoted attribute selector, so it doubles backslashes and escapes single quotes to keep the quoted string intact, and doubles a percent sign so Java's `String.format` leaves it alone. `By.name("50%off")` serialises to `*[name='50%off']`. It is quoting rather than the identifier escaping `By.id` performs.
  • Does the escaped form show up in what the locator reports about itself?
    No. `getRemoteParameters()` on `By.id("recall.due")` still answers with the keyword `id` and the raw, unescaped value, and `toString()` prints `By.id: recall.due`. Only the CSS fallback used for the actual search carries the escaped selector, which is why a trace and a locator can appear to disagree.

saying these in an interview costs you the question

  • Says By.id and a hand-written hash selector are always identical
  • Thinks the browser, not the client, escapes the id value
  • Believes an id containing a dot cannot be located at all
  • Assumes escaping is applied to the value sent as using id
  • Claims a leading digit is simply stripped from the selector