skip to content

Who records the video of a hosted browser session, and what must your test do to get it?

level: juniorimportance: must knowfreq 68%

answer

  1. the far side holds the camera
  2. no code in the test writes it
  3. one setting on the session request
  4. a capture process beside the browser
  5. a screen is grabbed, not a page

basics

~20 s

The service makes the recording on its own machine: a capture process beside the browser grabs the virtual screen for the whole session. A test asks with one setting in the session request and fetches the finished file afterwards.

solid answer

~50 s

The recording is made on the **far side**, by the service, not by your harness. A capture process sits beside the browser and records the virtual screen it paints on for the life of the session, so nothing in your museum membership-renewal test writes, encodes or uploads a frame. The run only asks: Selenoid — unmaintained by its own README, but a readable model of the mechanism — reads a boolean `enableVideo` from its `selenoid:options` map, while `docker-selenium` keeps a recorder container standing by beside the node and decides per session from the request. Because a screen is what gets grabbed, the file shows what a person sitting there would have seen: native dialogs, paint, animation, and the seconds after your assertion gave up. Retrieval is a separate step, and the finished file exists only once the session has closed.

code

json · 7 lines
json
{
  "browserName": "chrome",
  "selenoid:options": {
    "enableVideo": true,
    "enableVNC": true
  }
}

go deeper

for a junior

Be ready to say plainly that the service records the session on its own machine and the test only sets a flag on the session request. Interviewers ask this to check you know evidence can come from somewhere other than your own code.

for a middle

Expect to be pushed on what the file can and cannot contain. Practise the boundary out loud: painted pixels yes, request and response bodies no, and say where those actually live.

for a senior

You will be asked when recording should be on. Frame it as a default you choose at the account level and override per session, and be able to justify the override in terms of what somebody will actually open.

for a principal

Own the question of what recording everything obliges the organisation to do next: an artefact store that only grows, and a habit of switching capture on across teams with nobody owning what becomes of it.

## What a session video is A **session video** on a browser or device cloud is a recording made by the *service*, on its own machine, of the screen the browser was painting. It is not a capture your test takes, and not something your client library streams home. From the run's point of view the feature is one setting in the new-session request plus one file to fetch when the session is over. That inversion is the whole point. Your museum membership-renewal suite dials a remote endpoint, gets a session and drives it; meanwhile, on the far side, a capture process is watching the same virtual screen the browser draws on and writing an MP4 file. Nothing in the test knows, and nothing in the test has to. ## How a run asks for it Two open implementations show the shape end to end, in a way no closed product can be checked. - **Selenoid** — unmaintained by its own README, and read here only as a legible model — takes a boolean `enableVideo` inside its `selenoid:options` extension capability map. Its own documentation notes the feature works only when browsers are run in containers, and that it works with both VNC and non-VNC browser images. - **`docker-selenium`** deploys a recorder container beside the browser node. Its README says that recorder "is always up and stays in standby, watching for sessions"; whether a given session is actually recorded is decided per session from that session's request, falling back to the `SE_RECORD_VIDEO` environment default when the request says nothing. - On a **hosted provider** the same two shapes appear, described in ordinary words: recording is either on by default for the account, or requested per session through the provider's own namespaced settings map. You cannot verify a closed product's key names, so what you rely on is the shape — a setting on the session request, never code in the test. The consequence follows: **there is no failure hook to write.** You do not wrap the test, you do not catch the exception, you do not push bytes anywhere. A run that forgot to capture evidence still has a recording, provided the account or the request had recording on. ## Why the picture is a screen, not a page The recorder grabs a **display**, not a document. `docker-selenium`'s recorder runs `ffmpeg` with `-f x11grab` against an X display it addresses across the container network, so what lands in the file is literally the pixels the browser put on that screen. That decides what session video is good at and what it cannot do: - It shows things outside the page: a native file chooser over the membership form, a basic-auth prompt, an operating-system certificate warning, a browser crash tab. - It shows **time**: the spinner that never stopped, the layout settling half a second late, the order in which two things happened. - It shows the part of the run after your assertion failed and before teardown, which no assertion message contains. - It does **not** show anything that was never painted: a response body, a request header, a hidden DOM attribute, a console message. Those live in log channels, which is a different subject. ## Watching now against watching later These are two different mechanisms with two different lifetimes, and confusing them is the most common mistake newcomers make here. | | Live view | Session video | |---|---|---| | What it is | a stream of the screen while the session runs | a file written for the whole session | | When you can use it | only while the session exists | only once the session has closed | | Where it lives | a socket, gone when you disconnect | storage on the provider's side | | Selenoid's switch | `enableVNC` | `enableVideo` | Selenoid's live view is a raw bridge from the browser container's VNC port onto a WebSocket keyed by the session id; the moment that session is gone the bridge has nothing to attach to. The video is the opposite: it does not exist as a usable file *until* the session ends, and then it outlives it. A candidate who says "I watched the run, so I have the recording" has not understood either one. ## What the run still owns Three things stay yours even though the service does the recording: 1. **Turning it on** — either as an account-level default or per session on the request, and knowing which of the two you are relying on. 2. **Knowing which file is yours** — the far side names the artefact after its own session identifier unless you say otherwise, so something on your side has to carry that link. 3. **Ending the session** — the file is completed by the service's session-close path, so a run that never quits its session leaves an artefact that is not yet finished. There is a cost side too, and it is real without being a number: recording is trivial to switch on, and something has to decide later what becomes of what you switched on. Selenoid, for instance, intentionally ships no logic to remove old video files at all — how long a provider keeps footage is a separate subject in its own right, but the reason it matters starts here, with how easy it is to record everything.

  • If recording is on by default for the whole account, why would a suite still turn it off per session?
    Because every recorded session produces an artefact somebody must store, ship and eventually delete, and most sessions pass and are never watched. Turning it off for the runs you do not triage keeps the useful footage findable rather than buried, and keeps the encode off the host for sessions nobody will open. The decision is per session precisely so it can be made cheaply.
  • A colleague says session video makes browser console logs unnecessary. Why is that wrong?
    Video captures painted pixels. A console message, a failed request's status, a response body and a stack trace are never painted, so none of them appear in the file. Video tells you *when* and *what it looked like*; the log channels tell you *what was said*. They answer different questions and a triage that has only one of them is guessing about the other.
  • Why can a hosted provider record without the test asking, when a local run cannot?
    Because the browser is running on the provider's machine, inside the provider's container or virtual desktop, and the provider already controls that screen. Locally there is no second process watching your display unless you start one. The recording is a property of where the browser runs, not of the client library, which is why the same client code produces footage remotely and nothing locally.

saying these in an interview costs you the question

  • Thinks the client library records the video on the machine running the test
  • Expects the finished video to be downloadable while the session is still open
  • Believes recording requires a capture call inside every test
  • Confuses watching a session live with retrieving its recording afterwards
  • Expects network responses or console output to appear in the video