skip to content

In a parallel Selenium suite, a case's first command lands on the previous case's session — what does that prove about the ThreadLocal driver holder?

level: seniorimportance: should knowfreq 47%

answer

  1. The thread outlives the case that used it
  2. Nobody cleared the previous case's entry
  3. Check what get returns before setup runs
  4. Log the thread name beside the session id
  5. Quit or not quit decides which symptom

basics

~20 s

It proves the previous case's driver was never removed from the holder. Runners hand cases to a pool of reused threads, so the entry survives the case that set it and the next case on that thread reads it.

solid answer

~50 s

It proves teardown never reached `holder.remove()`. The runner does not create a thread per case; it hands cases to a pool, so the entry the previous case set is still on that thread when the next one calls `get()`. Which symptom you see depends on whether the old driver was quit. If it was, the first command throws `NoSuchSessionException`, because in Selenium 4 any command issued after `WebDriver.quit()` is rejected by the remote end as an invalid session id. If it was not, the case silently drives the old browser, so the tracker is still showing the previous candidate's application board. Confirm it by logging the thread name beside `RemoteWebDriver.getSessionId()` at every set and get, and by asserting at the top of setup that the holder is empty. The usual gap is a teardown skipped by an exception, or one that quits without touching the holder.

go deeper

for a junior

Know that a browser session ends with quit and that a test which suddenly acts on another test's page is a harness bug rather than a product bug. Diagnosing the holder is not expected yet.

for a middle

Explain why a pooled runner thread keeps its thread-local map between cases, and name the two symptoms a leftover entry produces depending on whether the old driver was quit.

for a senior

Show the confirmation procedure: log the thread name next to the session id, assert the holder is empty at the start of setup, then find the teardown path that skipped the clear.

for a principal

Be ready to argue that this class of bug belongs in the harness design rather than in triage, and to say what evidence a run should produce to prove no session leaked.

## The one fact that makes this possible A thread-local entry belongs to the **thread**, not to the case. Test runners typically hand cases to a **pool of reused worker threads** rather than creating a thread per case, so the thread that ran *submit a new application to the tracker* is still alive, still pooled, and still carrying whatever its map held when that case finished. When the pool hands it *archive a rejected application*, the entry from the first case is right where it was left. That is the only way one case's command can land on another case's session without any driver ever being shared between two threads. The two cases ran **sequentially, on one thread**. Nothing raced; something simply was not cleared. This is also why the failure is so hard to place. It appears in the second case, but the defect is in the first case's teardown, and the two may sit in different classes written by different people. ## Two symptoms, one cause Which symptom you get depends entirely on whether the previous case's teardown quit the driver before it failed to remove the entry. | Previous teardown did | What the next `get()` returns | The first command then | How it reads in the report | |---|---|---|---| | `quit()`, no `remove()` | a driver whose session id is already cleared | throws `NoSuchSessionException` | looks like a remote-end or environment fault | | neither `quit()` nor `remove()` | a live driver on the previous browser | succeeds, on the wrong page | looks like a product bug or a stale assertion | The second row is the dangerous one. The tracker is showing the previous case's filtered board and the previous candidate's applications, so an assertion about *this* case's newly created application fails with a message about a missing row. Nothing in that message points at the holder. ## Confirming the holder is at fault The confirmation is cheap and does not need a debugger: 1. **Log the thread name beside the session id** at every `set()`, `get()` and `stop()`. `RemoteWebDriver.getSessionId()` gives you the identifier the remote end knows the session by. 2. **Assert in setup that the holder is empty** before it sets anything. A non-null value at the top of a case is proof, on its own, that the previous case on that thread leaked its entry. 3. **Correlate by thread, not by time.** Line up the failing case with the case that ran immediately before it *on the same thread name*, not the one before it in the report. 4. **Count `set()` calls against `remove()` calls** for the whole run. Any difference is a leaked entry, and the count tells you how many. 5. **Reproduce by shrinking the pool.** Fewer worker threads than cases forces reuse, so the failure stops being intermittent. If the log shows two different case names writing and reading the *same* session id on the *same* thread name, the diagnosis is finished. ## The teardown paths that actually skip remove() In practice the clear is missed in a small number of recurring ways: - An exception thrown in setup **after** `set()` but before the case body, where the runner then decides teardown does not apply. - A teardown that returns early when it sees a null local variable and never reaches the holder at all. - `quit()` and `remove()` written as two statements in sequence, with `quit()` throwing on an unreachable browser. - A failure listener that quits the driver to release the browser and never touches the holder that still points at it. - Teardown running on a **different thread** from setup, so `remove()` clears an entry on the wrong map and leaves the real one behind. - A per-class teardown paired with a per-method `set()`, so the clear runs once for many entries. ## Why this is not the shared-driver bug, and why ThreadGuard stays silent It is tempting to file this as two threads sharing one driver, and the fix is different, so the distinction matters. A genuinely shared driver fails everywhere and immediately: commands from two threads interleave on one session, and `ThreadGuard.protect(WebDriver)` raises a `WebDriverException` naming the constructing thread and the calling thread. A leaked entry produces none of that. The previous case created the driver **on the very thread that is now reusing it**, so `ThreadGuard` compares two identical thread ids and says nothing. A holder-based suite that has passed a `ThreadGuard` audit can still be leaking entries on every case, which is precisely why the audit you need here is the empty-holder assertion in setup rather than a thread-identity wrapper. ## Closing it The repair is the ordering and the ownership, not extra vigilance. Put `quit()` in a `try` and `remove()` in the matching `finally`, so a driver that cannot be quit still clears its entry. Route every case through one shared teardown rather than per-class copies. Make `get()` throw a message that names the current thread when the holder is empty, so a missing setup fails at the first helper instead of returning null into page code. Then keep the started-versus-stopped counter as a standing check, because the failure it catches is invisible until the day the pool gets smaller than the suite.

  • If the old driver was already quit, why does the stale entry still cost you anything?
    Because the case now fails for a reason unrelated to the feature under test. A `NoSuchSessionException` on the first command reads as an environment fault, so it gets re-run and re-triaged instead of fixed. The entry also pins a dead driver object on a live pool thread for as long as the pool lives.
  • Does calling quit() twice on the same driver instance throw?
    No. In Selenium 4 a `RemoteWebDriver` whose session id is already null returns immediately from `quit()`, so a second call is a no-op. That is what makes a defensively written teardown safe, but it also means a double quit is not the signal that tells you an entry leaked.
  • Why does the same suite pass on a runner that creates a fresh thread per case?
    A thread that dies takes its whole thread-local map with it, so a missed `remove()` never shows. The bug is latent rather than absent: the browser session is still leaked, and the day the runner switches to a pool, or the pool drops below the case count, the cross-case failures appear.

saying these in an interview costs you the question

  • Blames flaky product behaviour instead of a leftover holder entry
  • Assumes the thread and its map die when the test method returns
  • Thinks quitting the driver is enough to clear the holder
  • Re-runs the case instead of logging thread name and session id
  • Concludes two threads shared one driver when one thread reused an entry