skip to content

Why is a ferry-timetable test slower on a hosted browser fleet than on your own machine?

level: juniorimportance: should knowfreq 70%

answer

  1. the protocol is request and reply
  2. count the conversation, not the work
  3. each command waits for its own reply
  4. cost scales with command count
  5. one scripted call beats many reads

basics

~20 s

Every automation command is a separate request and reply, so a hosted fleet adds a network round trip to each one. A chatty test pays that cost per command, which is why it slows more than the browser itself does.

solid answer

~50 s

The automation protocol is request and response: each find, click, text read and navigation is its own message to the remote end and its own reply back. On your own machine that exchange stays inside the box; on a hosted fleet it crosses a network, usually through an intermediary that routes to the machine actually running the browser; a private site under test adds a further hop of its own, on the browser's traffic rather than on your commands. The extra wall-clock is therefore roughly the round trip multiplied by how many commands the case issues - a property of your test's chattiness, not of the browser's speed. A case that walks a ferry timetable cell by cell suffers badly; one that reads the whole table in a single scripted call barely notices. Sleeps and timeouts tuned against local latency are the first things to break.

go deeper

for a junior

Be able to say that each automation command is its own request and reply, and that a hosted fleet adds a network trip to every one of them. Giving one example of a chatty test and its terser rewrite is enough here.

for a middle

Expect to be asked where the time actually goes and how you would measure it. Know that command count, not page weight, is the thing you control, and be able to explain why a polling wait costs more remotely.

for a senior

Be ready to size timeouts from measured remote behaviour and to argue for a harness change that cuts command count across a whole suite. You should also be able to say which failures this explains and which it does not.

for a principal

Decide how much latency the estate should design around rather than fight: where runners sit, whether chattiness is a review concern, and when a suite's shape should change instead of its configuration values.

## The conversation, not the work Browser automation is a conversation over a protocol that is strictly request and response. Finding an element is one message and one reply. Clicking it is another. Reading its text is another. Navigating is another. Nothing in that design changes when the browser moves to a rented fleet - but the distance each message travels does, and every message now waits for its own reply before the next one is sent. That single fact explains most of the slowdown people attribute to "the fleet being slow": - The browser is not necessarily rendering pages more slowly; the exchange with it is longer. - The added time is paid **per command**, not once per case or once per session. - A case's remote wall-clock therefore tracks how many commands it issues far more closely than how heavy its pages are. - Two cases doing the same user-visible work can differ enormously if one of them walks the page element by element. For a ferry-crossing timetable, this is easy to feel. A case that reads every cell of a crossings table one command at a time issues a message per cell. The same assertion written as a single scripted evaluation that returns the whole table issues one. Locally the difference is barely noticeable. On a hosted fleet it is the difference between a case that finishes and one that times out. ## What the extra distance is made of | hop | present locally | present on a hosted fleet | |---|---|---| | your process to the automation endpoint | yes, inside the machine | yes, across a network | | endpoint to the machine running the browser | usually none | usually one, inside the provider | | browser out to the site under test | direct | direct, or through a tunnel when the site is private | The tunnel hop is worth naming because it is the one piece of this that has no local counterpart at all: it exists only because the engine is somewhere your private site is not. How such a tunnel is opened and where its relay process runs is a subject of its own; for attribution purposes the point is simply that it adds a hop to traffic between the browser and the site, and that it is not the same hop as the one your commands travel. ## What breaks first 1. **Fixed sleeps.** A pause chosen because it "was enough locally" was really sized against a machine with almost no latency. Remotely, the same pause is either far too short or wasteful - and because it is a constant, it is wrong in a different direction on every different route. 2. **Timeouts.** A per-command or per-wait timeout sized on local feel expires on a remote engine even though nothing is wrong, which produces failures that look like product defects and are not. 3. **Aggressive polling waits.** Each poll is itself a command and therefore its own round trip. A wait that polls hard spends most of its budget on the wire rather than on the condition it is waiting for. 4. **Per-element assertions over collections.** Any loop that issues a command per row, per cell or per option multiplies the round trip by the collection's size. ## Reducing it on your own side You cannot shorten the network much, and you certainly cannot change how the provider routes internally, but the command count is entirely yours. - Read what you need in one call rather than walking the page. One scripted evaluation returning a whole crossings table replaces a command per cell. - Locate by a single precise selector rather than navigating a chain of parents and children. - Wait on a condition, not on the clock, and choose a polling interval deliberately rather than accepting the fastest one available. - Size timeouts from measured remote behaviour. Record how long commands actually take against the fleet and set the value from that, rather than from how a local run felt. - Reuse a single session across assertions that legitimately belong together, so session setup is not paid repeatedly for work that could share it. ## Measuring rather than guessing The cheap instrumentation is to time each command inside the harness and record, per case, both the total time and the number of commands issued. Two patterns then become obvious at a glance. A case whose duration tracks its command count is paying for the conversation, and the fix is to hold fewer conversations. A case whose duration is dominated by a small number of long commands is waiting on the browser or the site, and shortening the conversation will not help it. It is worth being precise about what is remote-only here. A local run also pays a hop - it is simply a very short and very steady one. What the rented fleet changes is the size of that hop, its variability, and the fact that neither is yours to control or even to watch. That is the honest version of the claim, and it is the version that survives someone pointing out that a local grid has a network in it too.

  • Does moving the test runner closer to the provider's region remove the problem?
    It shrinks the round trip; it does not change that one is paid per command. A chatty case still pays the smaller cost many times over, so cutting command count usually wins more than cutting distance - and the distance you can remove is bounded by where the provider actually runs the browser.
  • Why does a polling wait get more expensive on a hosted fleet than locally?
    Each poll is a command and therefore a round trip. A wait that polls hard spends most of its budget on the wire instead of on the condition. Poll less often, or express the condition as a single scripted evaluation the browser can settle without another exchange.
  • How would you measure how much of a run is spent on the wire?
    Time every command in the harness and record, per case, the total duration and the number of commands issued. A case whose duration tracks its command count is paying for the conversation; one dominated by a few long commands is waiting on the browser or the site instead.

Driving a browser on your own machine is like talking to someone in the same room; driving it on a hosted fleet is like exchanging letters. The conversation is no longer, but every sentence now waits for the post.

saying these in an interview costs you the question

  • Thinks the remote browser itself renders pages more slowly
  • Assumes a faster remote machine removes the per-command cost
  • Believes latency is paid once per case rather than per command
  • Widens every sleep instead of reducing the command count
  • Confuses time spent in the browser with time spent on the wire