skip to content

One browser-provider account serves your nightly suite and every pull-request run, and they starve each other. What do you do?

level: seniorimportance: should knowfreq 44%

answer

  1. a shared resource with no owner
  2. sized in isolation, summed by accident
  3. measure the ceiling, do not recall it
  4. declared shares beat inherited defaults
  5. the third pipeline is what breaks it

basics

~20 s

Treat the account ceiling as a shared budget with an owner: establish it by observation, enumerate every claimant on that account, give each a declared share of it, and guard the sum so a later pipeline cannot silently exceed it.

solid answer

~50 s

The fault is not either pipeline; it is that a shared resource has no declared division. Each was sized in isolation, so the sum of independent decisions exceeds what the account holds, and nothing in either configuration shows it. I would **establish** the ceiling by instrumenting concurrently-open sessions during a deliberately wide run and writing the result down, then **enumerate** every claimant on that account — including ad-hoc and developer runs — because the sum of claims is what must fit. Each claimant then gets a declared share of the established ceiling rather than whatever default its runner chose, with priority decided deliberately: the run a human waits on usually deserves the wider share. Finally I would guard the sum in CI, since the usual regression is a third pipeline added months later and sized sensibly in isolation.

code

python · 11 lines
python
# Grocery-delivery slot picker: each pipeline derives its width from the shared
# account ceiling it is one claimant on. Our side only - no provider surface here.
ceiling = config["observed_account_ceiling"]   # measured, not remembered
share   = config["pipeline_share"]            # agreed with the other pipelines

workers = max(int(ceiling * share), config["min_workers"])
run_suite(parallel_workers=workers)

# Guard that runs in CI, not inside the suite: the shares must still add up.
claimed = sum(p["pipeline_share"] for p in config["pipelines"])
assert claimed <= config["whole_of_the_ceiling"], "pipelines over-claim the account"

go deeper

for a junior

Be ready to spot that two pipelines on one account draw from the same pool, so once that shared ceiling is the binding limit, widening one of them takes width from the other.

for a middle

Be ready to explain why moving a pipeline to another runner or CI service changes nothing, and why sharding it wider raises simultaneous demand rather than relieving the contention.

for a senior

Be ready to give the sequence: establish the ceiling by observation, enumerate claimants, assign declared shares, decide priority on purpose, and guard the sum against a pipeline added later.

for a principal

Be ready to say who owns the account's capacity as a budget, how that ownership is exercised when a team adds work, and on what evidence you would take the commercial conversation about widening it.

## Name the real problem first The problem is not the nightly grocery-delivery slot-picker regression, and it is not the pull-request runs. It is that a single shared resource — the account's session ceiling — has no owner and no declared division. Each pipeline was sized on its own, probably by whoever set it up, probably against no measurement at all, and the sum of those independent decisions has quietly exceeded what the account can hold. That framing matters because it rules out the fixes that feel natural and do nothing: - Moving a pipeline to a different runner or a different CI service changes where requests come from, not how many are admitted. - Sharding the suite wider raises the number of simultaneous requests, which deepens the contention rather than relieving it. - Retrying harder converts a capacity problem into a longer capacity problem. - Raising either pipeline's worker count takes width from the other, since the ceiling they share is already the binding limit here. ## The sequence I would actually follow 1. **Establish the ceiling by observation.** Instrument the harness to record concurrently-open sessions, not requested ones, and run one deliberately wide suite at a quiet time. Write the result down somewhere durable. A ceiling remembered from a conversation is not an established ceiling, and this is the step teams skip. 2. **Enumerate every claimant.** Not just the two pipelines — the ad-hoc runs, the developer debugging from a laptop, the scheduled smoke check, anything else presenting the same account. The sum of claims is the quantity that has to fit. 3. **Give each claimant a declared share.** Replace the default worker count each runner picked for itself with a width derived from an agreed fraction of the established ceiling. The value of this is not the arithmetic; it is that the division becomes a thing written down and reviewable rather than an emergent property of several unrelated configuration files. 4. **Decide priority deliberately.** A run a human is waiting on usually deserves the wider share; a scheduled run can take the remainder, or take its full width only in a window when nothing else is scheduled. Either is defensible. Leaving it to whichever run happens to arrive first is not. 5. **Guard the sum in CI.** Add a check that the declared shares still add up. The failure mode this catches is the common one: someone adds a pipeline months later, sizes it sensibly in isolation, and pushes the total over without anyone touching the existing two. 6. **Alert on observed drift, not on configuration.** Compare observed concurrency against the established ceiling over time. Configuration says what you intended; observation says what happened. ## Two details that change the arithmetic **Arrival shape.** A nightly regression typically offers all its requests at once, at a fixed time. A pull-request run offers its requests whenever someone pushes. When the two overlap, the nightly is usually already holding the width and the interactive run finds none — which is why this fails asymmetrically and why teams describe it as the nightly "stealing" capacity even though both pipelines are behaving exactly as configured. **Occupancy, not just admission.** A session your side opened and did not close still occupies the ceiling until the far side reclaims it. How and when it is reclaimed is the provider's business and not something to assert, but the consequence for you is direct: a harness that leaks sessions on an abnormal exit makes a pipeline's real share larger than its declared one, and the symptom appears in the *other* pipeline. Reliable teardown on every exit path is therefore a capacity control, not only a hygiene measure. ## What a strong answer includes that a weak one does not A weak answer reaches immediately for more capacity, or for the service's behaviour when it is over the limit, and never establishes what the ceiling is. A strong answer treats the ceiling as a budget: measured, divided on purpose, guarded against drift, and owned by someone. It also declines to over-claim. You cannot see the provider's internal accounting, so the honest position is that you sized against an observed figure with a margin, and that you will re-observe it when the account's arrangement changes. And the reserved case is worth stating explicitly, because it is where senior candidates slip: width that has been paid for in advance is still a hard admission ceiling. Prepaying removes the marginal cost of using the width; it does not remove the top of it, and a suite sized above a reserved width hits the wall at exactly the busiest moment.

  • Why does this failure look asymmetric, with the nightly seeming to steal capacity?
    Arrival shape. A scheduled regression offers all its requests at once at a fixed time; a pull-request run offers its requests whenever someone pushes. When they overlap, the nightly is usually already holding the width and the interactive run finds none. Both pipelines are behaving exactly as configured, which is why nobody's configuration looks wrong.
  • How does a harness that leaks sessions distort the division you agreed?
    A session your side opened and did not close still occupies the ceiling until the far side reclaims it, so a leaking pipeline's real share is larger than its declared one. The symptom shows up in the other pipeline, not the leaking one. Reliable teardown on every exit path is a capacity control, not only hygiene.
  • Why is buying more capacity the wrong first move here?
    Because nothing has been established yet. Without a measured ceiling and a list of claimants, more width is divided by the same absent rule and the contention returns as soon as another pipeline appears. Establish, divide and guard first; then, if the divided shares are genuinely too small for the work, the commercial conversation is an informed one.

saying these in an interview costs you the question

  • Reaches for more capacity before establishing what the ceiling is
  • Shards the suite wider, which raises simultaneous demand instead
  • Relies on retries, turning a capacity problem into a longer one
  • Sizes each pipeline in isolation and never checks the sum
  • Assumes a leaked session releases the ceiling as soon as the run exits
  • Blames flakiness because the failure only appears when runs overlap