skip to content

Result Attribution

Joining an outcome recorded over there to the run you started here: telling the service how a case ended, and deciding whether a failure belongs to your code or to the machine it ran on.

on this pageshow

explore

questions

8

On a hosted browser fleet, what can end your ferry-booking session before the test does?

level: middleimportance: must knowfreq 63%

answer

  1. the far side owns a timer
  2. quiet on the wire is not idle
  3. the error names the symptom only
  4. unknown session on the next command
  5. compare last success to failure time

basics

~20 s

A remote end can end your session first: an inactivity timer that frees the slot, a ceiling on total session length, your own client hanging up, or the fleet reclaiming the machine. Later commands then fail on an unknown session.

solid answer

~50 s

The session lives on the far side, and several things there can end it before your test does. An inactivity timer reaps a session that has sent no commands for a while so the slot can serve someone else - Selenium Grid's node flag `--session-timeout` is documented in exactly those terms, and docker-selenium exposes the same setting as `SE_NODE_SESSION_TIMEOUT`. A limit on total session length can end a long run regardless of activity, your own side can hang up, and the fleet can reclaim an unhealthy machine. What you observe is not the cause but its shadow: the next command fails because the remote end no longer knows the session, which the W3C WebDriver specification names `invalid session id`. Separate the possibilities by comparing the time of the last successful command with the time of the failure.

go deeper

for a junior

Know that the session runs on the remote machine and can be ended there, and that the next command then fails saying the session is unknown. Be able to name inactivity as one reason that happens.

for a middle

Explain that inactivity is measured by commands arriving, not by whether your process is busy, and name what you would log to tell an idle reap from a limit on session length. Expect to be pushed on why the error message alone is not a diagnosis.

for a senior

Be ready to design the teardown and the session-start logging that make these failures attributable in minutes, and to say why keepalive traffic is usually the wrong answer. Describe a real case where a session ended under a test and how you proved which cause it was.

for a principal

Own the policy: what every harness must record about a session, what the estate assumes about cleanup when a runner dies, and when a limit you cannot change should reshape how long a single case is allowed to be.

## The session is not yours to keep A test that creates a browser session on a rented fleet gets back a handle to something running on somebody else's machine, under that machine's rules. Your harness can end the session deliberately. So can the far side, and it will do so for its own reasons - freeing capacity, enforcing a limit, recovering a host - none of which involve your test's progress. This is not unique to rented infrastructure. A grid you run yourself does it too, which is why the open implementations are the clearest place to see the shape: - **Selenium Grid** carries a node flag `--session-timeout`, documented as automatically killing a session that has had no activity for that period, explicitly in order to release the slot for other tests. - **docker-selenium** surfaces that same node setting as the environment variable `SE_NODE_SESSION_TIMEOUT`. - **Selenoid**, which declares itself unmaintained in its own README, has a server-side `-timeout` flag described as a session idle timeout, and its proxy cancels and re-arms the per-session timer on each request it forwards. What changes on a rented fleet is not the mechanism but the ownership: the value was chosen by somebody else, you may not be able to read it, and you may not be able to change it. ## The ways a session ends without you 1. **Inactivity.** No command has arrived for long enough, so the remote end reaps the session and frees the slot. 2. **Total length.** A limit on how long a session may live, applied regardless of how busy it has been. 3. **Your own side hanging up.** The harness crashes, the pipeline job is cancelled, the machine running the test disappears, or the connection drops. 4. **The fleet reclaiming the machine.** A host is drained, restarted or found unhealthy, and the sessions on it go with it. Do not assert which of these a provider does, when its timer fires, or what it logs internally. You cannot see any of it; you can only observe the symptom and establish the correlation. ## Idle means quiet on the wire, not idle in your process This is the part that catches people. An inactivity timer counts **requests received**, not effort expended. A test that is working very hard - building a large booking fixture, waiting on a slow external system, sitting on a breakpoint, sleeping in its own process - is, from the remote end's point of view, indistinguishable from a test that has stopped existing. Selenoid, unmaintained by its own README, shows the mechanism plainly: its timer is reset on each forwarded request, so anything that does not travel to the browser does not count as life. The inverse also holds, and it is why an idle timeout is not a safety net. A case stuck in a polling wait is sending commands continuously, so it looks perfectly alive while making no progress at all. ## What you see, and what you must establish The visible failure is almost never self-describing. The next command after the session ends fails because the remote end no longer recognises the identifier, which the W3C WebDriver specification names `invalid session id`. That error names the symptom - the session is gone - and says nothing about why, when, or on whose initiative. | observation | what it points at | |---|---| | a long quiet gap on the wire before the failure | an inactivity reap | | failure at a similar elapsed time from session start, whatever the step | a limit on session length | | your own process died or the job was cancelled first | your side, not theirs | | several concurrent sessions ending together | the machine or the route, not one test | So instrument for attribution rather than hoping the error explains itself: - Record the session identifier the remote end returned, and log it with every result. - Timestamp every command, and on failure report both the elapsed time since the last successful command and the elapsed time since session start. Those two values alone separate the first two causes. - Log whether your own process and connection were healthy at the moment of failure, so your side can be ruled in or out without argument. ## Cleanup is something you request, not something you assume A dropped connection is not a teardown. Selenoid, unmaintained by its own README, is explicit about this in its behaviour: its proxy error handler logs the client disconnect and returns a gateway error without removing the session, which then ends later on its own timer. Treat every rented session the same way: end it explicitly in a teardown block that runs even when the test throws, and never rely on the far side noticing you have gone. Two habits follow directly. - **Quit the session in a guaranteed-teardown path**, not at the end of the happy path. - **Keep the wire alive only when that is honest.** Do not invent keepalive commands to dodge a reaper while your test is genuinely stuck; that turns a session that would have ended into one that hangs, and hides the real defect. This is about a session that was created and then taken away mid-run. A request never granted in the first place, for want of free capacity, is a different situation with a different signature.

  • Your test was busy building a large booking fixture - why did the session still go idle?
    Idle is measured on the wire, not in your process. The remote end counts requests it has received, so a client working hard but issuing no commands looks exactly like one that stopped. Selenoid, unmaintained by its own README, re-arms its per-session timer on each request its proxy forwards, which is the same shape.
  • The pipeline job was cancelled. Does the session end the moment the client disconnects?
    Not necessarily. Selenoid, unmaintained by its own README, logs the client disconnect in its proxy error handler and returns a gateway error without removing the session, which then ends on its own timer. Assume nothing about the far side's reaction to a dropped connection, and make teardown something your harness requests.
  • How would you separate an inactivity reap from a limit on session length?
    Compare where the failure lands across runs. A reap follows a quiet gap on the wire and moves with whichever step goes quiet. A length limit lands at roughly the same elapsed time from session start whatever the test happens to be doing. Logging both elapsed values makes the answer immediate.
  • Would sending periodic no-op commands be a reasonable way to avoid being reaped?
    Only where the pause is legitimate and short, and even then it is a workaround. Keeping a stuck session alive converts a clean end into a hang and hides the defect that caused the pause. Prefer shortening the quiet stretch, or moving the slow work outside the session entirely.

saying these in an interview costs you the question

  • Assumes a session lives until the test deletes it
  • Thinks a busy test can never be judged inactive
  • Reads an unknown-session error as a product defect
  • Believes hanging up the client immediately frees the session
  • Claims to know when the provider's timer fires
open as a page

At a hosted browser provider, sessions from your driving-test booking suite carry no verdict — what are the causes?

level: seniorimportance: must knowfreq 57%

basics

~20 s

A missing verdict takes one of three shapes: the call was skipped before it ran, refused when it arrived, or lost when the process died. That is evidence about your reporting path, not about the test.

open as a page

What should a suite send to a hosted browser provider alongside each session's verdict, and why?

level: seniorimportance: must knowfreq 61%

basics

~20 s

Send the verdict, a case name that is stable across runs, something separating this attempt from the last, and a short failure reason. Keep secrets out of all of it: the payload lands on someone else's machine.

open as a page

Why is a ferry-timetable test slower on a hosted browser fleet than on your own machine?

level: juniorimportance: should knowfreq 70%

basics

~20 s

Every automation command is a separate request and reply, so a hosted fleet adds a network round trip to each one. A chatty test pays that cost per command, which is why it slows more than the browser itself does.

open as a page

Which ferry-booking assertions break on a remote browser whose window size, locale and clock you did not choose?

level: juniorimportance: should knowfreq 54%

basics

~20 s

Anything asserting rendered layout, formatted dates, currency or a date boundary can break, because the remote browser's window size, locale and clock were chosen by its operator. Read those settings back at session start and log them beside the result.

open as a page

Why is reporting a case's outcome to a hosted browser provider a separate call from those driving the browser?

level: middleimportance: should knowfreq 54%

basics

~20 s

Driving a rented browser and reporting the case's outcome are separate calls on separate channels, so either can fail while the other succeeds. Guard the reporting call so its own failure never turns a passing case red.

open as a page

Your ferry-booking suite fails only on the hosted browser fleet - how do you settle whether the test or the environment is wrong?

level: principalimportance: should knowfreq 48%

basics

~20 s

Treat a failure seen only on the hosted fleet as a claim needing evidence, not a verdict. Most remote-only causes exist locally in weaker form; what differs is that you neither own the setting nor observe the variance.

open as a page

Why does a hosted browser provider know your session ran but not whether your test passed?

level: juniorimportance: nice to knowfreq 50%

basics

~20 s

A hosted provider sees what crossed the automation connection: a session opening, its commands, and a close. Your assertion compares an expected value with an actual one inside your own process, so the verdict has to be sent separately.

open as a page