Several parallel jobs each open a tunnel to a hosted browser provider — how does a session attach to the right one?
answer
- several paths inward at once
- something must say which one
- the two ends have to agree
- a name computed twice can differ
- derive once from pipeline variables
basics
~20 sAn agreed identifier matches a session to a tunnel: once more than one is up, the request must name the tunnel it wants. Derive that name once per job from a pipeline variable, and read it back on both sides.
solid answer
~50 sWhile only one tunnel is ever up, nothing needs deciding. Once a matrix runs jobs side by side, the far side needs something in the session request saying which tunnel this session's private traffic belongs to — in practice a name the tunnel was started with and the request repeats. The failure that actually bites is **drift**: two steps of one job each computing a name rather than reading one, so the tunnel is up under a name no session ever asks for. Derive it once, from variables the pipeline already holds — the run identifier, the branch, the matrix cell — publish it into the job's environment, and let the tunnel startup and the session request both read that same variable back. Make it unique per job, so a name left by a dead run can never be attached to by accident.
code
python · 16 lines# PSEUDOCODE - entirely our own pipeline. No provider surface named.
# ONE place derives the name, from variables the pipeline already has.
def tunnel_name(run_id, branch, cell):
return slug("payroll-" + branch + "-" + run_id + "-" + cell)
# At job start: derive once, then publish it for the rest of the job.
name = tunnel_name(PIPELINE_RUN_ID, BRANCH_NAME, MATRIX_CELL)
set_job_variable("TUNNEL_NAME", name)
# Both ends read the SAME stored value back - they cannot drift apart.
open_tunnel(read_job_variable("TUNNEL_NAME"))
build_session_request(read_job_variable("TUNNEL_NAME"))
# Teardown is registered for every exit path, not only for success.
register_on_any_exit(lambda: close_tunnel(read_job_variable("TUNNEL_NAME")))go deeper
Know that more than one tunnel can be open at once, so a session has to say which one it wants. If a job cannot reach an internal address, check the name it asked for before suspecting the site itself.
Be ready to explain where that name comes from. Derive it once per job from the pipeline's run identifier and matrix cell, publish it, and have the tunnel startup and the session request both read the one stored value back.
Expect the drift scenario: a tunnel up under a name no session asks for. Show how you would prove attachment from the run's own recorded output rather than inferring it from a green result.
You will be asked what happens when two jobs claim one name. Argue that the collision must be made impossible by construction rather than handled, because the far side's behaviour on a collision is not yours to depend on.
## Why a tunnel has a name at all A tunnel exists so that a browser running on somebody else's machine can reach a name that resolves only inside your own network — here, a payroll portal whose staging host is built fresh for every branch. While exactly one tunnel is ever up, there is nothing to decide: one path leads inward, and traffic for the private name takes it. A matrix breaks that. Several jobs are live at the same time, each with its own reason to reach inside, and the far side has to send each session's private traffic down one path rather than another. What tells it which is information carried in the session request, matched against information the tunnel was started with. Nothing else in the exchange reliably differs between two of your own jobs: they commonly share an account, a network and often a runner. In ordinary words: **the tunnel has a name, and the request asks for that name.** Everything else about attaching to the right tunnel follows from that one fact — though not the separate question of what that tunnel can reach, which is a different setting altogether. ## Deriving the name so it cannot drift The mechanism is trivial. Getting the two ends to agree on it is where teams lose afternoons, and a few habits make it reliable: - **Derive the name in exactly one place.** A value computed twice is a value that can differ twice. Compute it once, at the start of the job. - **Build it from variables the pipeline already holds** — the run identifier, the branch, the matrix cell. Together they are unique to this job, and each is already available to every step, so nothing has to be invented or coordinated between them. - **Never derive it from anything that moves.** A clock reading, a random value, a retry counter: each of those answers differently in the second step than it did in the first. - **Publish it, then read it back.** The tunnel startup and the session assembly should read one stored value rather than both calling one recipe. A shared value cannot drift; a shared recipe can. - **Make it unique per job, and make an old one look old.** Embedding the run identifier means a name belonging to a finished run is recognisable on sight, which is what makes an abandoned tunnel findable later. Note that the unit of uniqueness is the **job**, not the run. A name unique per pipeline run is still shared by every cell of that run's matrix, which is precisely the case the naming was supposed to separate. ## What goes wrong, and what you actually see - **The request names a tunnel that is not up.** The symptom is usually that the private name is unreachable from the session — and often not at session creation but at the first navigation, which is confusing, because the session itself looked healthy. Whether a provider refuses such a request outright is that provider's own behaviour: establish it, do not assume it. - **Two jobs claim the same name.** From outside, those two tunnels are indistinguishable. What happens next — one refused, one replaced, traffic spread across both — is the provider's internal practice, undocumented in your repository and not stable over time. Do not build a matrix whose correctness depends on the answer. Make the collision impossible instead. - **A name is reused across runs.** This is the same defect displaced in time: a tunnel that outlived its job still claims the name, and the next run asking for it can be routed wherever that dead job pointed. - **The name is right and the reach is still wrong.** Attaching to the correct tunnel says nothing about which of your hostnames that tunnel carries. Those are two independent settings and they fail independently, so diagnose them separately. ## Proving the attachment instead of assuming it A green run is weak evidence here, because the worst of those failures are the quiet ones — the ones that route you somewhere real and wrong. Two cheap controls are worth more than a lot of debugging: 1. **Make the first thing the session does prove where it is.** Open the branch staging host and assert something only that host could produce — the branch name rendered on the page, a build identifier in a footer. "It loaded" is not the same statement as "it loaded mine." 2. **Record the name alongside the run.** Write the tunnel name into the job's own output next to the session identifiers. When somebody asks the next morning which path a failing job used, that should be readable rather than reconstructed from memory. ## Where this question stops The shape of the thing making the connection — what runs it, where it sits, and what an extra hop costs on every command — is a separate subject, as is deciding how wide the matrix should be in the first place. This question takes the matrix as given and asks only the routing problem it creates: with several live paths inward, which one is this session on, and how do you make that answer impossible to get wrong? The answer is an agreed identifier, derived once from what the pipeline already knows, read back by both ends, and unique to the job that owns it.
- What would you put in a per-job tunnel name, and why that rather than a random value?The pipeline's run identifier, the branch and the matrix cell. Each is already unique and already available to both steps, so no coordination is needed. A random value is unique too, but it is unique per computation rather than per job, so the two ends can disagree — and it carries no information, which makes a leftover tunnel impossible to trace back to the run that abandoned it.
- A job's tunnel is up and its session is created, but an internal page will not load. Where do you look first?At the name, on both sides, compared as literal strings. A trailing separator or a differently sanitised branch name is enough to break the match. If the two agree, the next question is not attachment but reach: whether that hostname is inside the set of names the tunnel carries at all. Those two settings fail independently and are worth ruling out in that order.
- Is one tunnel name per pipeline run enough?Only if one job in the run ever needs it. A matrix runs its cells side by side, so a name unique per run is still shared by every cell of that run. The unit of uniqueness has to match the unit that starts a tunnel, and that unit is the job rather than the run.
saying these in an interview costs you the question
- Assumes only one tunnel can be up at a time
- Computes the tunnel name separately in each step
- Builds the name from a clock reading or random value
- Assumes the provider refuses a duplicate tunnel name
- Treats a created session as proof the right path was used
- Confuses attaching to the right tunnel with reaching the right hosts