Your ferry-booking suite fails only on the hosted browser fleet - how do you settle whether the test or the environment is wrong?
answer
- only there is a claim
- list the differences before judging
- close the axes you own
- ownership and visibility, not impossibility
- record the environment with the result
basics
~20 sTreat a failure seen only on the hosted fleet as a claim needing evidence, not a verdict. Most remote-only causes exist locally in weaker form; what differs is that you neither own the setting nor observe the variance.
solid answer
~50 sA failure that reproduces only on a rented fleet has two possible authors: the environment, or a test assumption your own machine happened to satisfy. Start by writing down what actually differs - a network hop on every command, hardware you share and did not size, a session lifetime somebody else's timer controls, and a window size, locale and clock somebody else chose. Then close the axes you control: set the window size explicitly rather than maximising, assert on data rather than formatted text, and replace fixed sleeps with waits on observable state. Whatever still fails after that is the honest candidate for a genuinely remote cause. The distinction that matters is not *impossible locally* - a busy laptop contends too - but **unowned and unobserved**: you cannot read the provider's timer, and you cannot correlate its contention with anything you did.
go deeper
Be ready to say what you would check first when a test passes on your machine and fails on a rented browser fleet. Naming one concrete difference and how you would look at it beats any general theory.
Know the families of difference a remote engine introduces and be able to give the cheap experiment that separates two of them. Expect to be pushed on why 'it works locally' is not evidence that the test is correct.
You are expected to run the attribution to a named cause and to design the session-start logging that makes the next one cheap. Be ready to describe a case where the fleet was blamed and the test turned out to be wrong.
Own the standing rule. Decide what every run records, what evidence a failure filed against the provider must carry, and how much local-remote parity is worth buying. Be ready to defend leaving a difference unclosed because closing it costs more than the failures do.
## "Only there" is a claim, not a diagnosis A ferry-crossing timetable and booking suite that is green on a developer machine and red on a rented browser fleet has produced one piece of information: the runs differ somewhere. It has not said which side is wrong. The reflex to file it against the provider is strong, because the provider is what changed, and treating that reflex as a default costs teams weeks. The disciplined move is to convert the observation into a list of differences and then close them one at a time, starting with the ones you own. ## What actually differs Four families of difference come with running the engine on somebody else's machine. - **A network hop on every command.** The automation protocol is request and response, so a case's wall-clock cost grows with how many commands it issues, not with how much work the browser does. - **Hardware you share and did not size.** The machine running the browser may be serving other tenants, and the processing available to your session varies between runs for reasons you cannot see. - **A session lifetime somebody else's timer controls.** The far side may end a session that has gone quiet, or one that has simply run long, on a schedule you cannot read. - **A window size, locale and clock somebody else chose.** Those defaults were set by whoever operates it, not by you. None of these is strictly impossible on your own machine, and pretending otherwise loses the argument. A laptop compiling in the background contends. A grid you run yourself reaps idle sessions: Selenium Grid's node flag `--session-timeout` is documented as killing a session with no recent activity in order to release the slot, and docker-selenium exposes the same setting as the environment variable `SE_NODE_SESSION_TIMEOUT`. A colleague working in another country has a different clock. **What makes the remote versions different is ownership and visibility, not possibility.** | property | your own machine | the rented fleet | |---|---|---| | who set the value | you | whoever operates it | | can you read it | yes, directly | only what is published or returned | | can you change it | yes | only what is offered to you | | can you correlate a change in behaviour | yes, against your own load | no | ## Close the axes you own first 1. **Remove the assumptions your machine happened to satisfy.** Fixed sleeps tuned to local speed, waits that assume an element has already rendered, and clicks on controls a narrower viewport pushed below the fold are test defects that a fast, familiar machine was hiding. 2. **Make the environment legible.** At session start, read back and log the window size, the reported locale and the reported time-zone offset. A moment spent at session start beats any later argument about what it probably had. 3. **Assert on what you actually mean.** If the case is about a ferry crossing's departure data, assert on the data. If it is about how that departure time is rendered, state the locale and time zone the case requires and fail loudly when the browser reports another. 4. **Re-run the failing case locally with the remote's observed settings applied.** If it now fails on your own machine, the test was always wrong and the fleet merely stopped flattering it. ## What survives, and what to do with it Whatever still fails after that is the honest candidate for a genuinely remote cause. Foreign defaults are closed by the steps above, so the remaining hypotheses are the hop, the hardware and the timer, and each has a cheap distinguishing observation. - A failure landing at roughly the same elapsed time from session start, whichever step it reaches, points at a limit on session length. - A failure landing after a long quiet gap on the wire points at an inactivity reap. - A failure whose timing wanders run to run on the same steps points at contention. - A failure that moves when nothing changes but the route points at the path rather than the engine. None of those is a conclusion on its own; each is a hypothesis with a cheap next experiment attached, and the teams that resolve these are the ones that actually run it. ## Keeping the verdict honest at team scale Two habits keep this from decaying into folklore. - **Record the environment with the result.** The session identifier the remote end returned, the browser build it reported, the window size, locale and offset read back at start, and per-command timings. Without them, the next occurrence restarts the argument from nothing. - **Do not let the verdict be free.** Blaming the fleet costs nothing in the moment and produces no change, which is exactly why it accumulates. Requiring a named difference and its evidence is what stops a suite's failures becoming background noise. This is narrower than it sounds, deliberately. What a flaky test is, what causes intermittency in general, and how to tell a flake from a genuine intermittent product defect are a separate subject, as are quarantine and where a retry belongs in a suite. A here-against-there split is a different shape: it may be perfectly reproducible on both sides, and the question is which side's assumptions are wrong, not how often the result wobbles.
- Which single piece of evidence most often collapses an 'only there' failure into a test defect?The environment read back at session start. Once the window size, reported locale and time-zone offset the remote session actually had are recorded, most such failures reproduce on your own machine by applying those same settings - at which point the test, not the fleet, is what was wrong.
- What should a run record so that 'only there' can be checked later rather than remembered?The route or region the endpoint address selected, the session identifier the remote end returned, the browser build the session reported, the window size, locale and time-zone offset read back at start, and per-command timings. Without those, the next occurrence restarts the argument from nothing.
- How do you keep this from turning into a standing grudge against the provider?Make a named difference the price of the verdict. A failure filed against the fleet should carry which of the four families it belongs to and the observation that supports it. Unattributed blame is free, so it accumulates until nobody investigates anything.
saying these in an interview costs you the question
- Says a failure seen only remotely must be the provider's fault
- Treats 'it passes locally' as proof the test is correct
- Assumes remote causes are impossible on a developer machine
- Reaches straight for a re-run instead of naming a difference
- Blames the network without measuring anything on the wire