skip to content

Why does a hung test on a remote browser grid keep consuming despite an idle timeout?

level: middleimportance: should knowfreq 49%

answer

  1. two clocks, not one
  2. the reaper watches the wire
  3. a poll is activity
  4. silence is what gets reaped
  5. the bound comes from your side

basics

~20 s

An inactivity timeout measures silence on the wire, not progress in the test. A hung case that polls keeps sending commands, so the timer keeps resetting and the session keeps its slot until a bound of your own stops it.

solid answer

~50 s

Because the reaper and the hang measure different things. A remote fleet's session timeout is an **inactivity** timeout: it fires when no command has arrived for a while. Selenium Grid documents its node `--session-timeout` as killing a session that has had no activity and releasing the slot for other tests; Selenoid, unmaintained by its own README, re-arms each session's idle timer on every request its proxy forwards. A test hung in a polling wait — waiting for a planning application's status to read *validated*, say — issues a command on every poll, and each one resets that timer, so the session is invisible to the inactivity reaper while making no progress. The other shape is different: a case blocked inside a long call sends nothing, the reaper does fire, and your next command then fails as an unknown session. Your own timeouts are what bound the first shape.

go deeper

for a junior

Know that a wait in a test is a loop of real commands, not a pause. The remote side sees that traffic and treats the session as busy rather than abandoned.

for a middle

Be able to separate the two hang shapes — a polling wait the reaper never sees and a blocked call it does — and say what each one looks like from the client's side.

for a senior

Expect to be asked how you bound the exposure. Layer the timeouts so each is smaller than the one above, and make sure something outside the case is the thing enforcing the case's own limit.

for a principal

You will be asked why runaway consumption keeps recurring. Argue that an unbounded wait anywhere in a suite is an open-ended commitment, and make bounded waits a reviewable rule rather than a matter of taste.

## Two different clocks A remote fleet's session timeout and a hung test measure different things, and confusing them is the mistake this question exists to catch. - **The fleet's clock measures silence.** It is an *inactivity* timeout: it fires when no command has arrived for that session for a while, and its job is to release a slot whose client has evidently gone away. - **Your clock measures progress.** A test is hung when it is getting nowhere, which has nothing to do with whether it is sending traffic. The open implementations state the first clock plainly. Selenium Grid's node flag `--session-timeout` is documented as automatically killing a session that has had no activity for that long, and explicitly as releasing the slot for other tests; `docker-selenium` surfaces that same setting to its images as `SE_NODE_SESSION_TIMEOUT`. Selenoid — unmaintained by its own README — re-arms each session's idle timer on **every request its proxy forwards**, and only a delete removes the session and releases its run slot. A hosted provider will have something of the same shape, but its window and its exact rule are facts to establish with the vendor rather than carried over from an open project. ## The two hang shapes | hang shape | what the wire shows | what the reaper does | what your client sees | |---|---|---|---| | a polling wait for a condition that never becomes true | a command on every poll | nothing — the timer keeps resetting | the case runs on until one of your own timeouts fires | | a block inside a long-running command with no timeout | silence | fires, deletes the session, releases the slot | the next command fails as an unknown session | Take the planning-application tracker suite. A case submits an application and waits for its status to read *validated*. The status never changes, because the background job that would have changed it did not run. The wait polls: it asks for the status element, sleeps briefly, asks again. Each ask is a command, each command re-arms the fleet's idle timer, and the session therefore looks perfectly healthy from the far side while making no progress whatever. It holds its slot and keeps consuming until something of yours stops it. The other shape is often misread. A case calls a navigation, or asks the browser to evaluate a long script, with no client-side timeout on the call. Nothing goes out for the duration, the reaper eventually fires, the session is deleted and the slot released — and then your next command fails with an unknown-session error long after the last step that worked. Read that as a bound firing, not as a flaky product. ## Where the bound has to come from Since the fleet's clock catches only the quiet shape, the bound for the noisy one has to be yours. Layer it, each level smaller than the one above: 1. **A timeout on the individual call**, so no command can block for an unbounded time. 2. **A timeout on the case**, enforced by something *outside* the case — the runner, or a watchdog that can interrupt or kill it. A case blocked inside a network call cannot check its own clock, and in-case guards run only between steps, which is exactly where a hung case is not. 3. **A budget on the job**, so a pathological run ends even if the levels beneath it fail. What a blocking pipeline invocation should promise inside its budget belongs to pipeline wiring; here it is only the outermost backstop. Two more belong in that stack and are easy to leave out: - **Cap retries.** A bounded wait inside an unbounded retry loop is unbounded again, and a retry that reopens a session multiplies the consumption rather than adding to it. - **Make the case timeout close the session too.** A case abandoned without a close hands the problem back to the reaper, which will take it only once the case has actually gone quiet. ## What this does to consumption - An unbounded wait is an open-ended commitment on a metered fleet, and it survives review because waiting for a condition reads as patience rather than as a blank cheque. - It is **not** the only construct with that property: an unbounded retry loop, and a test that re-navigates forever, do the same thing. - Some fleets impose a **maximum total session length** as well as an inactivity timeout. The two are different rules, and a total-length cap would catch the polling hang the inactivity timeout misses — but whether a given provider enforces one is a vendor fact to establish before you lean on it. ## Getting it wrong in the other direction The tempting fix on a fleet you operate yourself is to shorten the inactivity timeout. It does catch abandoned sessions sooner. It still does nothing for the polling hang, and it starts killing legitimate sessions whose tests have a genuinely quiet stretch — a long upload, a slow report generation, a deliberate wait on a background job. Shortening a reaper is a blunt instrument aimed at a different problem; the bound on a runaway case belongs in the suite that runs it.

  • A case fails with an unknown-session error long after its last successful step. What does that pattern suggest?
    That the session was reaped while the case sat blocked inside a long call. The client sent nothing for the length of the reaping window, the fleet deleted the session and released the slot, and the next command found nothing to talk to. Read it as a bound firing rather than as a flaky product, and go looking for the call that had no timeout of its own.
  • Where should a per-case time limit be enforced, and why not inside the case?
    Outside it — in the runner, or in a watchdog that can interrupt or kill the case. A case blocked inside a network call cannot check its own clock, so a guard written inside the case only runs between steps, which is precisely where a hung case is not. In-case guards catch slow cases; only an external one catches a stuck case.
  • Why isn't shortening the inactivity timeout on a grid you operate a general fix?
    Because it still only measures silence. It catches abandoned sessions sooner and does nothing at all for a polling hang, and it starts killing legitimate sessions whose tests have a genuinely quiet stretch — a long upload, a slow report generation, a deliberate wait on a background job. The bound on a runaway case belongs in the suite, not in the reaper.

saying these in an interview costs you the question

  • Believes an inactivity timeout caps total session length
  • Thinks a stuck test is silent on the wire
  • Relies on the far side to stop a runaway case
  • Reads an unknown-session error as a product flake
  • Sets a case timeout larger than the job's own budget
  • Assumes every provider reaps on the same rule