skip to content

In Selenium, why can catching StaleElementReferenceException and retrying the same WebElement never succeed?

level: seniorimportance: must knowfreq 56%

answer

  1. What the instance is actually bound to
  2. The interface documents all future calls
  3. Recovery needs a brand-new reference
  4. Ask whether the action already applied
  5. Repeating a write is not free

basics

~20 s

The instance is bound to one reference the browser has already declared dead, and Selenium documents that all future calls on it fail. Recovery needs a new lookup, and a repeated write needs proof the first one did not land.

solid answer

~40 s

A `WebElement` holds a fixed reference to one node; it is not a re-runnable locator. `WebElement` is documented to check freshness on every call and to fail all future calls once that check has failed, so retrying the identical instance is guaranteed to raise the same exception. Real recovery is a **new** `findElement` that mints a new reference. The senior half is what you do around that: a stale failure on a read is safe to repeat, but a stale failure on a write — clicking a cellar row's decrement control — may mean the click landed and the re-render it triggered is what killed the handle. Re-finding and clicking again then decrements twice. Establish whether the action applied before repeating it.

code

java · 9 lines
java
String readVintageAfterReRender(WebDriver driver, By rowBy) {
  try {
    WebElement row = driver.findElement(rowBy);
    return row.findElement(By.cssSelector(".vintage")).getText();
  } catch (StaleElementReferenceException detached) {
    WebElement fresh = driver.findElement(rowBy);
    return fresh.findElement(By.cssSelector(".vintage")).getText();
  }
}

go deeper

for a junior

Remember that the failing object cannot be revived. If you catch the exception, the next thing you do must be a fresh lookup, not another call on the handle you already have.

for a middle

Explain why: the instance is bound to one reference, the interface documents that all later calls fail, and only a new find mints a usable reference.

for a senior

Show the production judgment — classify the failing command as a read or a write, verify observed state before repeating a mutation, and recognise that a re-find after navigation can bind to the wrong node.

for a principal

Own where recovery is allowed to live and what it may repeat, so the suite never trades a visible failure for a silent double-mutation in the name of stability.

## The instance is not repairable `WebElement` in the Java binding is a thin object over one thing: a reference the browser minted for a single DOM node. It carries no locator it can re-run and no ability to look around for a replacement. The interface's own documentation is explicit — every method call performs a freshness check, and once that check fails, **all future calls to that instance will fail**. Selenium is not describing likely behaviour there; it is describing a guarantee. So the shape people reach for first is dead on arrival: - Catching the exception and calling the same method again on the same object raises the same exception. - Sleeping and then calling it again raises the same exception; staleness is a property of the node's attachment, not a timing state that decays. - Passing the handle into a helper that catches and repeats changes nothing about which reference is being sent. ## What a real recovery looks like Recovery always means **obtaining a new reference**, which means running a lookup again. The unit that must be retried is therefore not the failing command — it is the find plus the command. Keep the `By` rather than the `WebElement`, and let the recovery path perform its own `findElement`. Where that recovery loop lives — inline, in a helper, or in a polling wait that ignores the exception — is a separate design question with its own owner. The point that belongs to the failure itself is narrower and non-negotiable: **a retry that does not re-find is not a retry, it is the same failing command sent twice.** ## The question that separates seniors: did the action already apply? This is where blind re-finding does real damage. Consider a wine-cellar inventory grid with a decrement control on each row. The test clicks it and gets a stale reference. Two stories fit that evidence equally well: 1. The click never reached the node, because a background refresh had already replaced the row. Nothing happened; repeating is correct. 2. The click **did** land, the application decremented the bottle and re-rendered the grid, and the stale failure came from the *next* command against the now-detached row. A recovery that re-finds and clicks again is right in the first story and silently wrong in the second — the cellar loses two bottles instead of one, the assertion afterwards may still pass on some other check, and the defect surfaces later as data no one can explain. | | Stale on a read | Stale on a write | |---|---|---| | Example | `getText()` on a row's vintage cell | `click()` on the row's decrement control | | Repeating it is | idempotent and safe | potentially a second, unwanted mutation | | Safe recovery | re-find, read again | re-find, then verify state before acting again | | Worst case if wrong | one extra command | corrupted data and a green test | ## A checklist before you repeat the action 1. Classify the failing command. Reads and state queries are idempotent; anything that submits, clicks a control, or types into a field is not. 2. For a non-idempotent command, re-find and **assert the current state** — read the quantity the grid now shows — before deciding whether to act again. 3. Prefer a check the application already gives you. If the row exposes the resulting quantity, that is a cheaper and more honest oracle than reasoning about whether the click landed. 4. If nothing lets you tell the two stories apart, treat the ambiguity as the defect and change the test to act on a page state it has just observed, rather than one it is holding a handle to. ## When re-finding raises a different error instead A re-find is not guaranteed to succeed. If the cause was a navigation rather than a re-render, the element may not exist on the page you are now on, and the same locator raises a not-found failure instead. That is not the recovery misbehaving — it is the recovery telling you that your diagnosis of the cause was wrong. The converse is more dangerous. If the new page happens to contain markup the locator matches, the re-find **succeeds** and hands back a healthy handle to a completely different node in a different document. The test proceeds against the wrong element and the failure moves somewhere unrecognisable. A recovery path that re-finds should therefore be scoped to the case it was written for, not applied as a blanket wrapper around every command in the suite. ## What to take from it The mechanical fact is small: the handle is dead, only a new lookup helps. The judgment around it is what an interviewer is probing — whether you know that re-finding is safe for reads, requires a state check for writes, and can quietly bind to the wrong element after a navigation. A candidate who answers only the mechanical half will write a recovery that turns a loud failure into a silent one.

  • You cannot tell whether the failing click landed. What do you change in the test?
    Stop relying on the click's outcome being knowable after the fact. Re-find the row, read the quantity the grid now shows, and branch on that observed state. If the application exposes no such signal, that is worth raising as a testability gap rather than papering over it with a retry.
  • Why is a blanket wrapper that re-finds and repeats every failed command a bad default?
    Because it treats reads and writes alike, so it can repeat mutations, and because after a navigation the re-find may bind to a similar-looking node on a different page. It converts a loud, well-located failure into a silent wrong-element run that is far harder to diagnose.
  • If the recovery re-find raises a not-found error instead, what has that told you?
    That the cause was a navigation, not a re-render — the element does not exist on the page you are now on. The recovery has diagnosed your misdiagnosis. The fix is to work out why the page changed, not to widen the locator until something matches.

saying these in an interview costs you the question

  • Retries the same WebElement instead of finding it again
  • Re-runs a click without checking whether it already applied
  • Treats the exception as flake needing no diagnosis
  • Assumes a re-find after navigation returns the same element
  • Wraps every command in a catch that swallows the failure