In Selenoid and docker-selenium, how does a session-video recorder capture a browser it never runs inside?
answer
- a third process, not the browser
- it grabs a display, not a document
- same host, so the encode competes
- one per session, or one standing by
- ffmpeg with x11grab on another container
basics
~20 sA separate container reads the browser's virtual display over the container network and encodes it, needing no hook in the browser and no driver command. The two projects differ in whether that recorder is per session or standing.
solid answer
~50 sThe recorder is a **separate process in its own container** that reads the browser's virtual display, so it needs no hook inside the browser and no command from your driver. `docker-selenium`'s recorder runs `ffmpeg` with `-f x11grab` against a display it addresses across the container network from `DISPLAY_CONTAINER_NAME`; it stands by beside the node and learns of each session either from the Grid event bus or by polling the node's status, selected with `SE_VIDEO_EVENT_DRIVEN`. Selenoid — unmaintained by its own README — inverts the deployment: it creates a video container **per session**, links it to that session's browser container, binds the video directory into it, and removes it when the session ends. Both encode on the same host as the browser, so the compression competes with the page for CPU; Selenoid exposes `videoCodec` for exactly that pressure.
code
yaml · 10 lines# docker-selenium: the recorder is its own service beside the browser node
chrome_video:
volumes:
- ./videos:/videos
depends_on:
- chrome
environment:
# the X display ffmpeg grabs, on the browser node's container
- DISPLAY_CONTAINER_NAME=chrome
- SE_NODE_GRID_URL=http://selenium-hub:4444go deeper
You are not expected to name the encoder flags. Do know that something separate from the browser is doing the recording, and that it is watching a screen rather than reading the page.
This is your tier. Be able to describe the recorder as an external process grabbing a virtual display, and to name at least one consequence of that: cost on the browser's host, or footage that starts before your first command.
Expect questions about the operational cost of recording every session on a busy fleet, and about how the recorder's lifetime interacts with node draining and session teardown.
Be able to argue when a fleet should record by default at all, weighing triage value against the encode load and the artefact volume it commits the organisation to handling.
## The shape of the thing A session recorder is not part of the browser and not part of WebDriver. It is a **third process**, started beside the session, that reads the virtual screen the browser is drawing on and writes an encoded file. Understanding that one sentence explains almost every behaviour that follows: what the footage can contain, why it costs CPU, why it starts before your first command, and why the file is not finished when your test is. The two open projects are worth reading precisely because they solve it differently while sharing the mechanism, and a closed provider's version cannot be read at all. ## How the recorder sees the screen `docker-selenium`'s recorder builds an `ffmpeg` command with `-f x11grab` and an input display it composes as `<container>:<number>` — the container name coming from `DISPLAY_CONTAINER_NAME` and pointing at the browser node's own X display, reached over the container network. It is grabbing an X display that belongs to another container, which is why: - the recorder needs no privileged access to the browser process, no extension and no driver command; - anything painted on that display is captured, including window furniture the page cannot see; - if the display is not up yet, there is nothing to grab, so the recorder waits for it; - the geometry of the file is the **screen's**, not the browser window's. Selenoid — unmaintained by its own README — attaches its recorder the same way in spirit: its docs say video recording "works by attaching a separate container with video capturing software to running browser container", and the code passes that browser container's address into the video container as an environment variable before starting it. ## Two deployments of one mechanism | | Selenoid | `docker-selenium` | |---|---|---| | Recorder lifetime | created per session, removed after | one container standing by beside the node | | How it learns of a session | Selenoid starts it deliberately | event bus or polling, per `SE_VIDEO_EVENT_DRIVEN` | | Per-session settings | capabilities in `selenoid:options` | capabilities read off the session, plus `SE_*` defaults | | Where the file goes | a host directory bound into the container | a volume the recorder writes into | Neither is obviously better and both are asked about. A per-session recorder starts clean and carries nothing over from the session before it, but it pays container creation on every single session and adds a step that can fail before the browser is even usable. A standing recorder amortises that startup and can react to session events for a whole node, but it must filter events to its own node and it is one more long-lived process to keep alive. ## What it costs, and where The encode happens **on the same host as the browser**. That is the fact behind almost every tuning knob a recorder exposes: 1. **CPU.** Compressing a moving screen is not free, and it competes with the browser rendering the membership-renewal form. Selenoid's own docs introduce `videoCodec` explicitly as the escape hatch when the default `libx264` "is consuming too much CPU", pointing at a cheaper encoder. 2. **Frame rate.** `docker-selenium` exposes `SE_FRAME_RATE` and Selenoid `videoFrameRate`; both default low, far below a real screen refresh, because triage footage does not need smoothness and every extra frame is more work and more bytes. 3. **Geometry.** Selenoid's `videoScreenSize` overrides the recorded width and height, and its docs warn that a size smaller than the actual screen trims the picture from the top-left corner rather than scaling it — so a badly chosen value silently crops the part of the form you needed. 4. **Disk and network.** The file is written next to the session and then has to go somewhere, which is why both projects grow an offsite-shipping story around the recorder. ## Why the picture starts early and stops late Selenoid starts the video container **after** creating the browser container but **before** waiting for the browser's WebDriver port to answer. The recording therefore begins during browser startup, not at your first command — which is exactly what you want when the complaint is "the session never became usable", and which explains why the footage is longer than the test. At the other end, the recorder is stopped deliberately: Selenoid signals its video container to terminate and waits for it before tearing down the browser, and `docker-selenium` asks `ffmpeg` to quit and only escalates if it does not. Neither can be replaced by killing the container, because an encoder that is killed has not written the trailer that makes the file playable — `docker-selenium` hedges against exactly that by writing fragmented MP4 (`-movflags frag_keyframe+empty_moov+...`) so that a truncated file still plays. ## What to say in an interview Say the mechanism first: an external process grabbing a virtual display. Then name the two consequences that matter operationally — the encode is a tax on the browser's own host, and the recorder's lifetime is not the test's lifetime. Then, only if asked, reach for the identifiers. On a hosted provider none of these knobs are yours and none of them are checkable, but the shape is the same, and reasoning from the shape is what lets you predict that a recorded session will be slightly slower and slightly longer than an unrecorded one.
- Why does a recorded session usually finish slightly later than the same session unrecorded?Two reasons, both structural. The encode runs on the browser's own host and takes CPU the page would otherwise have, so the run itself is marginally slower. And the session-close path now has extra work in it: the recorder must be told to stop, must flush and exit, and the file must be finalised before teardown completes. Neither is dramatic, but both are real and both scale with how many sessions a host carries.
- What happens to the recording if the recorder is started but the browser never becomes usable?You still get footage, and it is often the most useful footage you will get. Selenoid starts its video container before waiting for the browser's WebDriver port to answer, so a browser that crashes or hangs during startup is captured on screen. The session request fails and the client sees an error, but the recorder saw the desktop the whole time.
- Why must the recorder be filtered to its own node in a distributed grid?Because session lifecycle events are broadcast on the grid's event bus to every subscriber, not addressed to one recorder. A standing recorder beside one node would otherwise react to sessions created on other nodes, whose displays it cannot see, and would start encodes that capture nothing. `docker-selenium` resolves the node's identity from its status endpoint on startup and matches events against it.
saying these in an interview costs you the question
- Thinks the browser or the driver produces the video itself
- Assumes the recorder captures the DOM rather than a rendered screen
- Believes recording is free to the session because the provider runs the encoder
- Expects the recording to start at the first WebDriver command
- Thinks killing the container is a fine way to stop a recording