skip to content

How do you decide whether a suite on a busy hosted browser fleet waits for a slot or fails fast?

level: principalimportance: nice to knowfreq 41%

answer

  1. not a timeout, a routing choice
  2. is there anywhere cheaper to go
  3. fail fast while alternatives remain
  4. let the last attempt wait
  5. an intermediary can overrule the client

basics

~20 s

Waiting is right only when there is nowhere else to go and giving up costs more than the wall-clock. With an alternative host, region or later run available, refusing at once and moving on is the better default.

solid answer

~50 s

Frame it as a routing decision, not a timeout setting: the governing variable is whether the request has anywhere else to go. Ggr — Aerokube's unmaintained load balancer — encodes exactly that policy. On every attempt it deletes any `X-Selenoid-No-Wait` header the client sent and re-adds it only while more than a single candidate host remains, so a busy host is abandoned instantly while alternatives exist and the last attempt is allowed to queue. Generalise it: fail fast while you hold a cheaper option — another region, another lane, a later run — and wait only on the final attempt, inside a deadline your own job budget can afford. Two consequences follow. The client's preference is advisory, because an intermediary may discard it; and a fleet-wide fail-fast policy without a final waiting attempt converts a busy period into a wave of hard failures.

code

go · 16 lines
go
// aerokube/ggr (unmaintained per its own README), proxy.go - the routing loop.
// Per attempt: discard the client's own preference, then decide for itself.
for h, i := choose(hosts); ; h, i = choose(hosts) {
    count++
    r.Header.Del("X-Selenoid-No-Wait")
    if len(hosts) != 1 {
        // alternatives remain: tell the downstream Selenoid not to queue us
        r.Header.Add("X-Selenoid-No-Wait", "")
    }
    if h == nil {
        break loop
    }
    // on the final candidate the header is absent, so this attempt may wait
    resp, status := session(r.Context(), h, r.Header, c)
    // ... browserStarted / browserFailed / seleniumError handling
}

go deeper

for a junior

Understand that failing immediately and waiting are both legitimate responses to a busy fleet, and that your suite behaves very differently under each. Ask which one your setup does before tuning anything.

for a middle

Be able to explain why the choice depends on having an alternative to route to, and to describe what happens to a suite when every lane fails fast during a busy hour.

for a senior

Expect a scenario where one lane gates merges and another runs overnight, and be ready to give them different policies with reasons rather than one number for both.

for a principal

Own the argument that this is a routing and admission-control decision, and be ready to defend staggering or deferring lane starts over any tuning of wait behaviour.

## The question behind the question "Should we wait or fail fast?" is usually asked as a timeout-tuning question and is really a routing question. A wait is only ever worth its wall-clock when the request has **nowhere cheaper to go**. If another host, another region, another lane or a later run could satisfy the same request sooner than the queue will, waiting is strictly worse than moving. If this is the last option, refusing converts a delay into a failure and buys nothing. So the decision variable is not "how long can we wait" but "what is the alternative, and what does it cost". ## A working policy, read from an open router Ggr, Aerokube's unmaintained load balancer, encodes exactly this and is the clearest available model. Its routing loop picks a candidate host, and on **every attempt** it: - deletes any `X-Selenoid-No-Wait` header the client itself supplied, and - re-adds it only while more than a single candidate host is still in the list. The effect is a policy with two regimes. While alternatives remain, Ggr tells the downstream Selenoid — also unmaintained — "do not queue me" so that a busy host is abandoned instantly and the next one is tried. On the final candidate the header is absent, so that attempt is allowed to sit in the queue and wait. Fail fast while there is somewhere to go; wait when there is not. Two consequences of that code are worth carrying into your own design: - **The client's preference is advisory.** Ggr deletes the header unconditionally before deciding for itself. A suite that sets a fail-fast hint and then routes through an intermediary has expressed a preference, not a guarantee. - **A fleet-wide fail-fast policy with no waiting attempt is a bad policy.** It converts every busy period into a wave of hard failures. The last attempt is what makes the whole thing safe. ## Designing your own Three inputs decide it, and none of them is a timeout number: 1. **Do you have an alternative?** A second region, a self-operated grid for the cheap cases, or simply "run this lane after the other one finishes". If yes, prefer refusal and route. 2. **What does the delay cost?** A merge-gating smoke suite for a ticketing site is on a human's critical path; a nightly regression sweep is not. The same wait is expensive in the first and free in the second. 3. **What does a hard failure cost?** If a red run means a person investigates, a failure that was really "we were busy" is expensive noise, and waiting is cheaper. The typical shape that falls out is a hybrid: the merge-gating lane fails fast and is retried by routing it elsewhere or deferring it, while the long overnight lane waits, because nothing is behind it and a failure there costs an engineer's morning. ## Where the policy is actually enforced It is worth being precise about who can enforce what, because this is where designs go wrong: - **The service** decides its default. Some hold, some refuse, and the same product may do either depending on how it was started. - **An intermediary** can override the client, as Ggr does. If your traffic goes through a router you did not configure, its policy wins. - **The client** can always bound its own patience, because it owns the socket. That bound is the one guarantee available to you regardless of what the far side does. - **The scheduler around the suite** decides how many requests are made at all, and admission control there is usually cheaper than any of the above. That last point is the one most often skipped. If the ceiling is the binding constraint, the cheapest fix is not a better wait policy but asking for less: start fewer lanes at once, or move a lane to an hour when nothing else is asking. How that ceiling is structured, and which consumers of it contend, is a separate question from this one. ## Judging it in an interview A strong answer names the alternative as the decision variable and refuses to give a single number. It distinguishes the merge-gating path from the batch path and gives each a different answer. It notes that the last attempt should wait even when the policy is fail-fast, and it acknowledges that an intermediary may silently override the client's preference. A weak answer proposes one global timeout, or argues for pure fail-fast because "retries are cheap", which is true right up to the moment every lane fails simultaneously during a busy hour and nobody can merge. Finally, hedge the claim you make about the service. Whether a given hosted product holds or refuses, and whether it lets you influence that, is something to establish by running a deliberately over-concurrent probe against it — not something to assume from another product's behaviour.

  • Why should the final attempt wait even when the policy is fail-fast?
    Because a fail-fast policy with no waiting attempt converts every busy period into a wave of hard failures. While alternatives remain, refusing is free — you simply try the next one. Once there is no alternative left, refusing buys nothing and costs the run, so the last attempt should be allowed to queue inside whatever deadline your job budget affords.
  • Your suite sets a fail-fast hint and it appears to be ignored. What are the likely causes?
    Either the far side does not implement that hint, or something between you and it rewrote the request. Ggr, an unmaintained load balancer by its own README's admission, unconditionally deletes the `X-Selenoid-No-Wait` header a client supplied and re-adds it only on its own terms. Treat a client-side hint as a preference; confirm the actual behaviour with an over-concurrent probe rather than assuming the header travelled.
  • When is the right answer neither waiting nor failing fast?
    When the ceiling is the binding constraint. Then the cheapest fix is asking for less: start fewer lanes at once, stagger their starts, or defer the lane nobody is waiting on. Admission control on your side is usually cheaper than any wait policy, because it removes the contention instead of scheduling around it — how the ceiling itself is structured is a different conversation.

saying these in an interview costs you the question

  • Answering with a single global timeout for every lane
  • Arguing for pure fail-fast because retries are assumed cheap
  • Assuming a client-side fail-fast hint always reaches the service
  • Treating the decision as tuning rather than as routing
  • Giving a merge-gating suite and an overnight sweep the same policy