skip to content

Log Channels

The text a rented session writes down and the trip it makes to reach you: produced on a host you do not own, listed or pushed elsewhere, and removed on somebody else's schedule.

on this pageshow

explore

questions

4

Should a hosted fleet's session logs be pulled from the provider or pushed to storage you own?

level: principalimportance: must knowfreq 66%

answer

  1. two channels, not one decision
  2. where the artefact rests between runs
  3. the producing host is the least durable part
  4. a push needs somewhere to push to
  5. check whether the local file survives the trip

basics

~20 s

Pull leaves each artefact on the machine that made it and fetches only what a failure needs; push sends every artefact to storage you control as it is produced. Decide by how often you will actually read them.

solid answer

~50 s

Treat it as two channels with different failure modes. **Pull** leaves the artefact on the host that produced it and you fetch it by name later. Selenoid, unmaintained by its own README, models this exactly: started with `-log-output-dir` it serves that directory under `/logs/`, where a client can list it, download one `.log` file, or `DELETE` one. **Push** means the server ships each finished artefact outward as it is created; Selenoid's S3 uploader fires on a file-created event and builds the object key from `-s3-key-pattern`. The trap is that push is not automatically an extra copy: unless Selenoid is started with `-s3-keep-files`, a successful upload deletes the local file, so the `/logs/` surface stops answering for it. Most suites end up pushing, because the host that produced the artefact is among the least durable parts of the estate.

code

bash · 14 lines
bash
# Selenoid is UNMAINTAINED per its own README; read as a model of the mechanism.
# The -s3-* flags exist only in a build made with the s3 build tag, absent by default.
# Both channels on at once: write locally AND ship offsite, keeping the local copy.
./selenoid \
  -log-output-dir /opt/selenoid/logs \
  -save-all-logs \
  -s3-endpoint "$OBJECT_STORE_ENDPOINT" \
  -s3-bucket-name "$OBJECT_STORE_BUCKET" \
  -s3-key-pattern '$quota/$date/$fileType/$sessionId$fileExtension' \
  -s3-include-files '*.log' \
  -s3-keep-files

# Without -s3-keep-files a successful upload removes the local file,
# and GET /logs/ stops listing it.

go deeper

for a junior

Know that a session's log is written on the remote machine, not yours, and that getting it back is a separate step your harness has to take deliberately.

for a middle

Be ready to describe both channels and what each costs to run. Knowing that a fetch path and a delivery path are different mechanisms with different failure modes is the expected depth here.

for a senior

Be ready to say how you verified which shape you are on, and to name the check for whether delivery copies or relocates the artefact. Interviewers look for the verification step, not the preference.

for a principal

Own the decision and its blast radius: what the estate runs, who holds the storage credential, and how a change of channel is staged so no window of evidence is lost.

A rented browser session writes text on a machine you do not own, and that text has to make a trip before anyone on your team reads it. The trip has two shapes, and choosing between them is a design decision with operating consequences rather than a preference. ## Pull: the artefact waits where it was made In a pull design the log is written to disk on the host that ran the browser and stays there until somebody asks for it by name. Aerokube's Selenoid, which declares itself unmaintained in its own README and is read here only as a legible open model of the mechanism, implements this directly. Started with the `-log-output-dir` flag it serves that directory, and writes a finished session's log into it whenever that session asked or the operator switched capture on for everyone: - `GET /logs/` serves a listing of the directory. - `GET /logs/?json` returns the same directory as a JSON array of file names. - `GET /logs/<name>.log` downloads one saved file. - `DELETE /logs/<name>.log` removes one, answering not-found if it has already gone. Pull is cheap to operate. There is no second system to run, no storage credential to hold, no key layout to design, and nothing to reconcile. Its weakness is structural: the only copy sits on one of the most disposable things in the estate. A host that is rebuilt, drained, scaled down or simply reclaimed takes the evidence with it, and nothing in the pull channel notices. ## Push: the artefact is sent onward as it is produced In a push design the server that made the artefact delivers it to storage you name. Selenoid models this too, in a build made with its `s3` build tag, which its own documentation says is absent by default. When a session ends and its file is finalised, a file-created event reaches an uploader, which computes an object key from the `-s3-key-pattern` flag and writes the file into the bucket named by `-s3-bucket-name` at `-s3-endpoint`. The glob flags `-s3-include-files` and `-s3-exclude-files` decide which artefacts make the trip, matched against the file's base name, so a suite can ship logs and leave heavier artefacts behind. Push costs you the things pull did not: storage to own, a credential to hold, and a key layout to design, because the key layout is what the bucket is later navigated by. ## The trap: a push is not automatically a second copy This is where confident designs go wrong. In Selenoid the uploader removes the local file after a successful upload unless it was started with `-s3-keep-files`. The push therefore *moves* the artefact by default rather than copying it, and the moment it succeeds the `/logs/` listing stops answering for that file. A team that wires push on top of pull without that flag and assumes both surfaces still work has one surface, not two, and discovers it during an incident. The same reasoning applies to any hosted service: establish whether a delivery mechanism duplicates the artefact or relocates it, because "we have it in both places" is an assumption, not an observation. ## Choosing between them | what you are weighing | argues for pull | argues for push | |---|---|---| | how often a log is actually read | rarely, on a real failure | routinely, by tooling | | how long the producing host lives | long-lived, stable hosts | short-lived, reclaimed hosts | | who must be able to read it | one engineer, on demand | a pipeline, a search index, an auditor | | what you are willing to operate | nothing extra | storage, a credential, a key layout | For most suites the answer is push, because pulling ties evidence to the lifetime of infrastructure you do not control, and that lifetime is the one variable a tenant cannot set. The honest counter-argument is that push obliges you to run and own the destination, so a suite that reads a log only when a person is already looking at a failing build can legitimately stay on pull. ## Working it through on a real suite Take a port-logistics container-tracking dashboard whose end-to-end suite runs on rented browsers each night. Its failures cluster in the shipment-status view, and triage happens the following morning. 1. Decide who reads the artefact. If the answer is "a person, the next morning", pull is survivable only while the producing host is still there the next morning, which on elastic capacity it may not be. 2. Decide what the artefact is worth once it is stale. If nobody reads a passing run's log, push everything but expire aggressively at the destination, which is a decision you can make only once the destination is yours. 3. Decide the key layout before the first upload, not after. Renaming a bucket's worth of objects later is work; getting the prefix right on the first day is a line of configuration. ## What to say in an interview Name the two shapes, say what each obliges you to run, and then name the trap: verify whether the delivery mechanism leaves a local copy behind. Candidates who stop at "we save the logs somewhere" have not made a decision; candidates who say "we push, we keep the local copy during the migration, and we verified the key resolves before we trusted it" have.

  • What decides whether a pushed artefact can be found again long after the run?
    The key layout, because the object key is the only index the store has. A pattern that encodes the account, the date and the session identifier lets a person narrow by prefix; a pattern that emits only the bare file name produces a flat namespace that can be listed but not navigated. Selenoid, unmaintained by its own README, exposes this as `-s3-key-pattern` server-side, overridable per session.
  • How would you move a suite from pulling to pushing without losing a week of evidence?
    Run both channels during the overlap. Configure the uploader to keep local files so the existing fetch path keeps answering, then verify that keys resolve in the destination for a full cycle of the suite, including a failing run. Only after that do you retire the fetch path. The order matters because the default behaviour deletes the local copy.
  • Which half of this channel does a tenant actually control on a hosted provider?
    Usually only the asking half. The tenant can request capture for a session and can fetch what the service offers, but the directory, the delivery mechanism and the moment of deletion belong to the operator. That is why the design question is phrased as an investigation: find out which shape you are on before you build triage around it.

Pull is leaving parcels at the depot that packed them and collecting the few you need. Push is a forwarding service that moves your post rather than copying it: everything reaches you, but nothing is waiting at the old address any more.

saying these in an interview costs you the question

  • Assumes pushing offsite always leaves a second local copy
  • Treats a listing endpoint as a durable archive
  • Plans to fetch a log from the producing host long after the run
  • Turns capture on everywhere without deciding where it lands
  • Thinks the key layout is cosmetic rather than the only index
open as a page

In Selenoid (unmaintained), what must be true before a session's log file is saved?

level: middleimportance: should knowfreq 57%

basics

~20 s

Two conditions, not one. Selenoid, unmaintained by its own README, must have been started with a log output directory, and the session must have asked with its log capability, unless the operator switched capture on for every session.

open as a page

In Selenoid (unmaintained), what does the s3KeyPattern capability decide about a pushed artefact?

level: middleimportance: should knowfreq 49%

basics

~20 s

It decides the object key the artefact is stored under, meaning both its prefix path and its name in the bucket. Set per session, it overrides the server-wide pattern for that session's files only, and it is expanded from placeholders.

open as a page

In Selenoid (unmaintained), how does following a session under /logs/ differ from downloading its saved .log file?

level: seniorimportance: nice to knowfreq 41%

basics

~20 s

One path serves two unrelated channels. Where an output directory is configured, a remainder that is empty or ends in .log is answered as a file from it; anything else opens a websocket that follows the running browser container's output.

open as a page