Should a whole test matrix share one tunnel to a browser provider, or should every job open its own?
answer
- one path inward or many
- who shares a failure, who shares reach
- the routed set becomes a union
- setup paid once or paid per cell
- measure setup against the shortest job
basics
~20 sSharing trades isolation for setup: one path serves everything, so its failure is everyone's and its routed set is everybody's union. Per-job tunnels isolate blast radius and reach, at the price of setup in every cell and more to leak.
solid answer
~50 sTreat it as **shared fate against per-cell overhead**, and decide it from facts you can measure rather than from taste. A shared tunnel is set up once, watched as one thing, and adds nothing to any job's wall-clock; in exchange it is a single point of failure for the whole matrix, its routed set must be the union of what every consumer needs rather than what any one job needs, and it gives you no way to attribute traffic to a job. Per-job tunnels make reach, failure and attribution follow the job, which is what you want when cells differ in what they may touch — but setup is paid in every cell rather than once, and every cell is another chance to leave one running. The deciding facts are setup time against the shortest job, how much cells differ in reach, and whether teardown is reliable.
go deeper
You are unlikely to own this decision, but know that the two shapes exist: one path everything attaches to, or one path per job. Find out which your pipeline uses before you change anything about it.
Be able to name one concrete cost on each side — a shared path failing for everyone at once, and per-job setup being paid in every cell — rather than declaring one of the shapes correct in general.
Expect to be pushed on the routed set. Explain that sharing forces it to the union of every consumer's needs, so a job gets reach it never asked for, and say what you would measure before changing the shape.
You own this trade. Frame it as shared fate against per-cell overhead, name the measurements that settle it, and be ready to defend a hybrid — one path per environment rather than per cell — when neither pure shape fits.
## What the choice is actually between Take the matrix as given — its width, how many jobs run at once and how the suite is sharded are other people's decisions, and this one sits downstream of all of them. What is left is a shape question: does every job attach to one long-lived path into your network, or does each open and close its own? The honest framing is **shared fate against per-cell overhead**; everything else here is a detail of one of those two. A shared path is set up once, monitored as one thing, and adds nothing to any job's wall-clock. A per-job path makes reach, failure and attribution follow the job that owns them, and charges setup in every cell instead of once. | | One shared tunnel | One tunnel per job | |---|---|---| | Setup cost | paid once for the matrix | paid in every cell | | A failure takes down | every attached job at once | one job | | Routed set | the union of everyone's needs | exactly this job's needs | | Attribution of traffic | not from the path itself | trivial; one path, one job | | Leftovers | one path to watch | as many chances to leak as cells | | Naming | one reusable name | one name per job, derived per job | ## What a shared path costs everybody on it - **Shared fate.** When it goes, every in-flight job goes with it, and every queued one starts failing at the same moment. A restart to fix anything — a configuration change, an upgrade on your own side — is an outage for the whole matrix rather than one cell. - **A routed set that is the union.** One path serves every consumer, so the names it carries must cover what all of them need. A job that only ever touches one branch's payroll staging host can now reach every host in the union. That is a least-privilege argument, and unlike the reliability one it survives a path that never fails. - **No attribution.** With many jobs on one path you cannot say from the path which traffic belonged to which job. That matters the morning after, when the question is which cell caused something. - **A reusable name.** Sharing one path means sharing one name, and a reusable name is precisely the condition under which a path that outlived its job becomes a correctness problem for the next run rather than merely an untidy one. ## What per-job costs you - **Setup in every cell.** The weight of that cost is not its absolute size but its size *relative to the job*. Setup that is a small fraction of a long job is noise; the same setup in front of a short job dominates it, and it is paid in every cell rather than once. - **More moving parts.** Every cell now derives a name, opens a path and is responsible for closing it. That is more code in the pipeline and more places for it to be wrong. - **More chances to leak.** A matrix that opens many paths has many opportunities to leave one behind, so per-job shapes raise the value of reliable teardown and of a reconciliation somebody actually runs. - **Isolation you may not be using.** If every cell needs identical reach and cells never interfere, per-job buys separation from a problem you did not have, and charges setup for it in every cell. ## Deciding it from measurements Argue this from numbers you can take rather than from taste. In rough order of weight: 1. **Setup time against the shortest job in the matrix.** That ratio is the per-cell overhead, stated honestly. Measure it; do not estimate it. 2. **How much cells differ in reach.** If different legs need different hosts, per-job (or per-leg) keeps the routed set honest. If every cell needs the same hosts, that argument is not available. 3. **How often the shared path has actually failed, and how many jobs each failure took.** Shared fate is only expensive if it happens. 4. **Whether teardown is reliable on the cancelled and killed paths**, not just the successful one. If it is not, multiplying paths multiplies leftovers. 5. **Whether anyone needs to attribute traffic to a job.** If nobody has ever asked, do not pay for it. ## The shape between the two The pure alternatives are rarely the best answer. A path **per environment, per team or per matrix leg** keeps the routed set close to what its consumers need, limits a failure to one group rather than the whole matrix, and pays setup once per group rather than per cell. The right unit is whatever shares both a reach requirement and a tolerance for failing together — those two properties are what the decision is really about. Two overstatements to avoid, because an interviewer is listening for the hedge. **"Per-job is always safer"** is false: it is safer against shared fate and worse against leftovers, and it buys isolation you may have no use for. **"Sharing is cheaper"** is true only of setup. Whether it is cheaper overall depends on how often the shared path fails and what a failure costs, which is a measurement about your estate rather than a property of the shape.
- What would you measure before moving a matrix from a shared tunnel to one per job?Setup time for a single tunnel against the length of the shortest job in the matrix, because that ratio is what per-cell overhead actually costs. Then how often the shared path has failed and how many jobs each failure took with it, and whether cells genuinely differ in what they need to reach. Without that third fact, per-job buys isolation nobody is using.
- Why is the routed set an argument against sharing even when the shared path is reliable?Because a shared set has to be the union of everything any consumer needs. A job that should only reach one branch's staging host can then reach every host in the union, so reach stops matching intent. That is not a reliability argument at all, it is a least-privilege one, and it survives a path that never fails.
- Is there a shape between one tunnel for everything and one per job?Yes, and it is often the right answer: one per environment, per team or per matrix leg. It keeps the routed set close to what its consumers actually need and limits a failure to one group, while paying setup once per group rather than once per cell. The unit should be whatever shares both a reach requirement and a tolerance for failing together.
saying these in an interview costs you the question
- Says per-job tunnels are always the safer choice
- Ignores that a shared routed set is everybody's union
- Forgets a shared path fails every attached job at once
- Assumes setup cost is negligible in every cell
- Cannot attribute traffic on a shared path to a job
- Decides from taste rather than from measured setup time