A Selenoid session dies mid-test though the client asked for a longer sessionTimeout. Why?
answer
- idle, not elapsed
- every command re-arms the clock
- waiting outside the browser is the trap
- the server's ceiling wins silently
- the same timer collects abandoned work
basics
~20 sSelenoid's sessionTimeout measures idle time between WebDriver commands, not elapsed test time, and the server silently clamps any value above its max-timeout flag. A test that waits without driving the browser idles out and the container is removed.
solid answer
~50 sTwo mechanisms produce that failure and both are server-side, in Selenoid — unmaintained by its own README, as all four Aerokube repositories are. First, its `sessionTimeout` is an **idle** timeout: Selenoid arms a timer per session and cancels and re-arms it on every proxied WebDriver command, so a test that blocks without touching the browser — waiting on a rota planner's overnight batch, polling a database in test code — idles out while it is busy. When the timer fires, Selenoid logs `[SESSION_TIMED_OUT]`, sends itself the session delete, and the container goes with it. Second, the value is **silently clamped**: Selenoid parses the duration and uses it only if it is at or under the server's `-max-timeout` flag, otherwise it substitutes `-max-timeout` and returns no error at all. So the operator's ceiling quietly beats the client's request, and only an unparseable value is rejected outright.
code
java · 8 linesMap<String, Object> selenoidOptions = new HashMap<>();
// Go duration string; Selenoid clamps it to the server's -max-timeout
selenoidOptions.put("sessionTimeout", idleTimeout);
ChromeOptions options = new ChromeOptions();
options.setCapability("selenoid:options", selenoidOptions);
WebDriver driver = new RemoteWebDriver(rotaGridUrl, options);go deeper
Know that the grid, not your test, decides when a session has been idle too long, and that the clock counts time since the last browser command. If your test waits without touching the browser, expect to lose the session.
Be able to explain that each proxied command re-arms the timer, and that the server holds a ceiling the client's request cannot exceed. Say what the client actually sees when a session is reaped.
Expect to be handed the symptom and asked to diagnose it. Go to the server log for the timed-out line, compare the requested value against the server's ceiling, then find the interval in the test with no browser traffic and close it there.
You set the ceiling for everyone. Balance recovering containers and slots from abandoned runs against killing legitimately slow work, and decide whether long-running suites get their own grid rather than a special-cased limit on a shared one.
This is one of the most reliably confusing failures on a self-operated container grid, because two independent server-side rules conspire and neither of them tells the client what it did. Selenoid is unmaintained by its own README, but the pattern is general: a hosted service also decides when your session has stopped being interesting, and it also does not ask. ## Rule one: it is an idle timeout, not a budget Selenoid attaches a timer to each session. Every WebDriver request proxied for that session cancels the pending timer and arms a fresh one. So the clock measures **time since the last command**, not elapsed session time. The consequences invert the usual intuition: - A test that drives the browser constantly for a very long time will never trip it. - A test that sits still for a stretch longer than the timeout trips it, even though the test is working hard. - The dangerous pattern is waiting *outside* the browser: sleeping in test code, polling a database directly, waiting for a message queue, waiting on a rota planner's nightly shift-generation job to finish before asserting on the roster page. - A debugger breakpoint held in a test does exactly the same thing, which is why sessions vanish during interactive debugging. When the timer does fire, Selenoid logs `[SESSION_TIMED_OUT]` with the session id and issues a `DELETE` for that session against itself. That runs the normal deletion path: the session leaves the map, the run-limit slot is released, and the container is force-removed. From the client's side the next command comes back as an unknown session, which is why the symptom is usually reported as "the driver lost the session" rather than as a timeout. ## Rule two: the value is silently clamped The per-session value arrives inside Selenoid's own namespaced capability map as `sessionTimeout`, written as a duration string. Selenoid parses it and then applies a ceiling: | What the client sends | What Selenoid uses | What the client is told | |---|---|---| | No value at all | The server's default idle timeout | Nothing | | A value at or under `-max-timeout` | Exactly that value | Nothing | | A value above `-max-timeout` | `-max-timeout` instead | **Nothing** | | A value it cannot parse | Nothing; the session is refused | `invalid argument`, "invalid sessionTimeout capability" | The third row is the trap. Asking for more than the operator allows is not an error and not a warning — the session is created and quietly runs on the server's ceiling. A malformed value, by contrast, is a hard refusal. So "it accepted my capability" tells you nothing about whether it honoured it. ## Diagnosing it on a real suite 1. Read the Selenoid log for the run and look for `[SESSION_TIMED_OUT]` against the session id your client reported losing. That log line separates a reaped session from a crashed browser. 2. Compare what your client asked for with the `-max-timeout` the server was actually started with. If yours is larger, yours never applied. 3. Find the gap. Instrument the test to log the time of each WebDriver call and look for the interval with none, then ask what the test was doing there. 4. Fix the gap rather than the timeout wherever you can: poll the page through the browser instead of waiting beside it, so the session keeps producing traffic while the wait happens. 5. Only then negotiate the ceiling with whoever operates the grid, because raising `-max-timeout` for you raises it for every abandoned session on that host too. ## The same reaper is a feature It is worth arguing the other side, because an interviewer often will. The timer is also the grid's garbage collector: - A runner that crashes, a pipeline cancelled mid-run, a developer who closes a laptop — none of them send the session delete. - Without the reaper each of those would strand a container **and** hold a run-limit slot forever, and a grid would slowly grind to a halt on abandoned work. - Because the reaper uses the ordinary deletion path, an abandoned session is cleaned up exactly like a healthy one — but only while Selenoid is alive to fire the timer, since it holds those timers in memory. - The lower the ceiling, the faster abandoned work is recovered, and the more likely a legitimately slow test is killed. That tension is the whole reason the setting is negotiable per session at all. ## What an interviewer is listening for - That you say **idle**, unprompted, and can explain what re-arms the timer. - That you know an over-large request is clamped without an error, so the accepted capability proves nothing. - That you would look in the server's log before changing anything in the test. - That you can defend the reaper's existence from the operator's side, not just complain about it from the suite's side.
- Your rota planner test must wait for an overnight batch to finish before asserting. How do you keep the session alive?Do the waiting through the browser rather than beside it. Poll the page for the state you need with the driver, so each poll is a WebDriver command that re-arms the idle timer. If the wait genuinely belongs outside the browser, drive a cheap command on a schedule while you wait, or end the session and open a fresh session for the assertion phase.
- How would you tell a reaped session apart from a browser that crashed?By the server's log. A reaped session is preceded by a timed-out entry naming the session id, and the deletion follows the normal path. A crashed browser shows up as the container dying or the driver failing to answer, without that line. From the client both look like commands failing against a session that no longer exists, so the log is what separates them.
- Is raising the server's max-timeout a reasonable fix for a slow suite?Rarely, and never unilaterally. The ceiling applies to every session on that host, including abandoned ones, so raising it means crashed runners hold containers and run-limit slots for longer. Fix the silent gap in the test first; if the work really is that slow, negotiate a higher ceiling knowing you have traded recovery time for it.
saying these in an interview costs you the question
- Reads sessionTimeout as a total wall-clock budget for the test
- Assumes a request above the server ceiling is rejected with an error
- Thinks a busy test can never be timed out
- Blames the browser for crashing without reading the server log
- Wants the ceiling raised without seeing that it applies to everyone
- Does not realise the same timer cleans up abandoned sessions