One long-lived deployed target serves every case and cannot be duplicated per test worker — how do you set the suite's isolation model?
answer
- The target is the constraint, not the runner
- Partition inside it, or lane around it
- Count the cases touching instance-global settings
- Cap the width at what it sustains
- Make each case declare what it needs
basics
~20 sChoose between two models: hand each worker its own slice of a scope the product already models, or run a wide pool of self-contained cases beside a small exclusive lane holding named locks. What the product scopes decides which.
solid answer
~50 sWhen the target cannot be duplicated, the runner's isolation stops mattering — every worker converges on one instance — so the model is built from what the cases ask that instance to do. - **Partition inside it.** If the product models a real scope — an organisation, a workspace, an account family — give each worker its own and let everything fan out. This is the durable option: isolation becomes a property of the data. - **Classify and lane.** If no such scope exists, sort cases into a wide pool that touches only its own data and a small exclusive lane, run one at a time, for cases that change something global to the instance. What picks between them: whether that scope exists, what fraction of cases touch instance-global settings, and the width the instance sustains. Cap the pool at that width, and make each case declare what it needs.
code
pseudocode · 10 lines# Model A: partition inside the one instance
worker.slice = organisations[worker_index] # every case fans out
# Model B: classify, then lane
case.declares(needs = ["instance_wide_toggle"]) # exclusive lane, one at a time
case.declares(needs = []) # wide pool, fans out
schedule:
pool_width = min(available_workers, width_target_sustains)
exclusive_lane = 1go deeper
Be ready to say why one shared deployed instance limits parallel runs at all: every worker is talking to the same place, so anything that place holds once is held for all of them.
Explain the two shapes: slice the instance by a scope the product already has, or classify cases and give the few that change instance-global settings a lane of their own. Say what each shape requires to work.
Show that you cap the pool at what the instance sustains and can tell contention failures from product defects. Explain how each case would declare what it needs so the model is enforced rather than remembered.
Own the choice and its conditions: what fraction of cases touch instance-global settings, whether the product models a scope you can hand a worker, and what the exclusive lane costs. Say when you would stop refining the model and fund a target that can be created per run.
## The constraint is the target, not the runner When one long-lived deployed instance serves every case and cannot be duplicated per worker, the runner's isolation stops mattering very quickly. Threads, processes, separate machines — all of them converge on the same instance, the same data behind it and the same settings that are global to it. The isolation model therefore has to be built at the level of *what the cases ask that instance to do*, not at the level of how the runner dispatches them. There are two defensible models. Most real suites end up with a blend, but the choice of which one is primary is the decision worth making deliberately. ## Model A — partition inside the target Give every worker its own slice of a scope the **product already models**: an organisation, a workspace, an account family, a region, a project. Each worker signs in inside its slice and every case fans out. This is the better model when it is available, because it needs no classification and no scheduling cleverness — the isolation is a property of the data, so it holds no matter how the suite grows. It requires three things to be true: 1. The product genuinely has a scoping concept, and it is a real boundary rather than a display filter. 2. Nothing the cases assert crosses the boundary — no shared counter, no global list, no cross-slice report. 3. There is capacity to keep several slices populated, and creating one is cheap enough to do per run. ## Model B — classify and lane Sort the cases into two groups: a wide pool of cases that touch only their own data, and a small **exclusive lane** for cases that touch something global to the instance — a toggle, a system-wide configuration value, a job that runs once, a resource the product models as singular. The pool fans out to the width the target sustains; the lane runs its cases one at a time holding a named lock. This is the model when the product has no scope to hand a worker, or when the scope exists but a minority of cases deliberately reach outside it. Its whole risk is the honesty of the classification. A case in the wide pool that quietly touches something global produces failures that land on other cases, which is the most expensive failure shape there is. ## Choosing between them | | Partition inside the target | Classify and lane | | --- | --- | --- | | **Needs from the product** | a real scoping boundary | nothing | | **Needs from the suite** | slice-aware setup | an honest, maintained classification | | **How it ages** | holds as the suite grows | degrades as people forget to declare | | **Failure shape when wrong** | assertions cross slices and fail directly | one case breaks a different case | | **Ceiling** | slice capacity on the instance | the run time owned by the exclusive lane | The conditions that actually decide it: - **What fraction of cases touch instance-global settings.** A tenth is a lane. Half is not a lane, it is a serial suite with a decoration. - **Whether the product models a scope you can hand a worker.** If it does, use it; nothing you build will be as durable. - **What the instance sustains.** Beyond some width, latency rises, unrelated cases time out, and you are measuring your own load rather than the product's behaviour. Cap the pool at that width and treat the cap as part of the model, not as a temporary workaround. - **The cost of cross-talk debugging.** When one case's failure is caused by another's, the investigation crosses cases, and that cost is paid by whoever is on the failure, repeatedly. ## Making the model enforceable Whichever you pick, the model has to live in the code rather than in people's memory. Have each case **declare what it needs** — a slice, or a named exclusive resource — and have the scheduler refuse to place a case that declares an exclusive need into the wide pool. Then audit the declarations by periodically running the wide pool at a higher width than usual: anything that starts failing was misclassified, and the failure names the case rather than the run. Also publish two numbers, because they are what the next decision turns on: the share of run time owned by the exclusive lane, and the width at which the instance stops behaving. ## Knowing when to stop There is a point where every isolation model is a workaround for the single instance. If the exclusive lane owns most of the run, or the sustainable width is low enough that fanning out buys nothing, the honest recommendation is not a cleverer scheduler — it is funding a target that can be created per run, and saying plainly what the current arrangement costs in wall-clock and in debugging time. Owning that recommendation, with the two numbers behind it, is the judgement the question is really asking for.
- How do you keep the classification of which cases need exclusive access honest as the suite grows?Make it declared and enforced rather than remembered: a case states what it needs, and the scheduler refuses to place a case with an exclusive need into the wide pool. Audit it by running the pool at higher width on a schedule — anything that starts failing was misclassified, and the failure names the case rather than the run.
- What tells you the concurrency against the shared target has gone too wide?Failures that move rather than repeat: timeouts spreading across unrelated cases, latency climbing with worker count, and the same case failing at high width and passing at low. At that point you are measuring your own load, not the product's behaviour, so cap the pool at a sustainable width and treat the cap as part of the model.
- When is the honest answer to stop refining the model and fund a duplicable target instead?When the exclusive lane owns most of the run time, when the sustainable width is low enough that fanning out buys nothing, or when cross-talk debugging costs more than the wall-clock it saves. Bring the two numbers — the lane's share and the sustainable width — and make the recommendation with them rather than around them.
saying these in an interview costs you the question
- Widens worker count without asking what the target sustains
- Assumes the runner's isolation extends to the deployed target
- Leaves which cases need exclusivity as tribal knowledge
- Picks a model without naming the condition that picked it
- Blames intermittent cross-talk on the cases rather than the model
- Never measures how much run time the exclusive lane owns