In Selenium, why can WebElement.getText() return an empty string for an element whose markup has text?
answer
- Rendered, not raw
- The browser drew nothing to report
- Layout rules reach the string too
- Uppercase styling survives the read
- textContent answers the other question
basics
~20 sBecause it returns rendered text, not markup text. Anything the browser did not draw contributes nothing, so an element hidden by display none or visibility hidden yields an empty string even though its characters are in the page.
solid answer
~40 s`getText()` is specified to return the element's text as rendered, and the W3C WebDriver specification defines the result as exactly what Selenium's own visible-text function produces. Subtrees the browser did not draw are skipped, so a "LATE" badge on a route planner's stop row that is styled `display: none` returns `""` rather than its characters, and the same happens under `visibility: hidden` or inside a closed `<details>`. Rendering also changes the text you do get: whitespace is collapsed, block boundaries become newlines, the result is trimmed, and `text-transform` is applied, so a label whose markup reads `depot pickup` under `text-transform: uppercase` returns `DEPOT PICKUP`. When you want the characters as authored, Selenium 4's `getDomProperty("textContent")` reads the raw property instead.
code
java · 13 lines// <span class="stop-label"> depot pickup </span>
// .stop-label { text-transform: uppercase; }
WebElement label = driver.findElement(By.cssSelector("#stop-14 .stop-label"));
String rendered = label.getText(); // "DEPOT PICKUP"
String raw = label.getDomProperty("textContent"); // " depot pickup "
// <span class="late-badge">LATE</span>
// .late-badge { display: none; }
WebElement badge = driver.findElement(By.cssSelector("#stop-14 .late-badge"));
String hidden = badge.getText(); // ""
String hiddenRaw = badge.getDomProperty("textContent"); // "LATE"go deeper
Recall that the call returns the text a person would see, not everything written in the HTML, so a hidden element can be found successfully and still hand back an empty string.
Explain what rendering does to the string: hidden subtrees drop out, whitespace collapses, block boundaries become newlines, the result is trimmed, and text-transform is applied before you see it.
Show that you know when to reach past it. Be able to say why a comparison against rendered text is fragile under styling changes and what the raw textContent read gives you instead.
Own the convention: decide whether the suite reads what a user sees or what the template wrote, because mixing the two across page objects produces failures whose cause nobody can locate quickly.
## "Rendered text" is a narrower thing than the text in the DOM `WebElement.getText()` is specified to return an element's text **as rendered**. The W3C WebDriver specification defines the result as exactly what Selenium's own `bot.dom.getVisibleText` function produces — the specification says so directly, noting that Selenium's behaviour was in wide use before the standard and set the expectation. The command itself is `GET /session/{session id}/element/{element id}/text`. Rendered means what the browser laid out and painted. Text the browser did not render is not in the answer at all, which is why an element whose markup obviously contains characters can hand back an empty string. ## Why a hidden element returns an empty string The visible-text walk skips subtrees the browser did not draw. On a courier route planner's stop row, all of these produce `""`: - the element, or an ancestor, has computed `display: none`; - the element has `visibility: hidden` or `visibility: collapse`; - the text lives inside a closed `<details>` element that is not the `<summary>`; - the node is inside a collapsed panel the board has not expanded yet. The last one is the common production case: a "LATE" badge that exists in the DOM from first paint but is only shown once the route is recalculated reads as `""` until then, and a check that expects the badge's text gets an empty string rather than a missing-element failure. ## What the browser does to the text before you see it Rendered text is not the markup with the tags removed. The walk reproduces layout decisions: - **Whitespace is collapsed.** Runs of spaces, tabs and newlines become a single space under the usual `white-space` values; `pre` and `pre-wrap` are honoured instead of collapsed. - **Block boundaries become newlines.** A stop row made of two `<div>` children returns two lines, not one run-on string. - **`text-transform` is applied.** A label whose markup says `depot pickup` under `text-transform: uppercase` returns `DEPOT PICKUP`. This is the single most surprising line in the method for most candidates. - **Non-breaking spaces become ordinary spaces**, and zero-width characters and directionality marks are stripped. - **The result is trimmed** at the start and end, and each line is trimmed too. So a string comparison against `getText()` is a comparison against the *rendered* string, not against the template. ## `getText()` against the raw text When you want the characters as authored, read the DOM property instead. Selenium 4's `getDomProperty("textContent")` returns the node's `textContent`, which the browser does not filter. | read | hidden descendants | whitespace | `text-transform` | |---|---|---|---| | `getText()` | excluded | collapsed, lines trimmed | applied | | `getDomProperty("textContent")` | included | preserved as authored | not applied | | `getDomProperty("innerText")` | excluded | collapsed like rendering | applied | `innerText` is the browser's own rendered-text property and sits close to `getText()`, but it is the browser's definition, not the specification's, so the two can disagree on edge cases. ## It reads the whole subtree, not just the element The walk descends into descendants, so the value you get back is the visible text of the element **and everything under it**, with block-level pieces separated by newlines. - Reading a whole `<tr>` stop row returns every visible cell in that row as one multi-line string, which is almost never what the caller wanted. - Reading the one `<span>` that holds the stop name returns just that name. - An inline element's text joins its neighbours' text on the same line, because that is how the browser laid it out. Scoping the read to the smallest element that carries the value is what keeps the returned string predictable across markup changes. And like every other read on this interface, each call is one command and one round trip, so pulling four cells off a row is four commands. ## Reading it on a dispatch board 1. If a read comes back empty, check whether the element was drawn at all — `isDisplayed()` on the same element answers that in one more call. 2. If a read comes back with unexpected case or spacing, look at the CSS before the markup: `text-transform` and `white-space` are doing it. 3. If you genuinely need the characters the template wrote, including the ones the page is hiding, read `textContent` through `getDomProperty` and accept that you are no longer reading what a courier could see. ## The distinction worth stating out loud `getText()` answers "what does this element show a person right now". `textContent` answers "what characters are in this subtree". They are different questions, and the empty string is `getText()` answering its own question correctly — the element shows nothing, because the browser drew nothing.
- A label's markup says "depot pickup" but getText() returns "DEPOT PICKUP". What happened?CSS did it. Rendered text applies `text-transform`, so an uppercase transform on the label produces uppercase characters in the value Selenium returns, with nothing changed in the markup. The same class of surprise covers collapsed runs of whitespace, newlines introduced at block boundaries, and non-breaking spaces coming back as ordinary spaces. If you need the characters as authored, read `textContent` through `getDomProperty` instead.
- How do you tell an empty rendered text from an element that is not there at all?They fail differently. A missing element makes the lookup itself fail, so you never reach the read; an element present but undrawn is found successfully and hands back an empty string. When a read comes back empty, calling `isDisplayed()` on the same element separates the two cases: `false` means the element exists and was not drawn, which is a different defect from markup that never rendered the node.
- Does getText() include text from child elements?Yes. The walk descends the whole subtree and returns the visible text of the element and its descendants, joining block-level pieces with newlines. So reading a whole stop row returns every visible cell in it, which is usually more text than a caller expects - scoping the read to the specific cell is what keeps the value predictable.
saying these in an interview costs you the question
- Says getText strips the tags and returns the markup text
- Expects hidden text to appear in the returned string
- Assumes an empty string means the element was not found
- Believes CSS cannot change the characters that come back
- Treats getText and textContent as the same read