skip to content

Attacker-Controlled Inputs

Everything a job reads before it runs — its own definition, the scripts it invokes, the event payload, the steps it pulls — is reachable by someone. Most CI compromises begin as an input problem.

on this pageshow

explore

questions

14

In a CI pipeline, which parts of the triggering event are attacker-controlled, and which are not?

level: juniorimportance: must knowfreq 68%

answer

  1. who typed it, a person or the platform?
  2. free text is the dangerous half
  3. hashes and event names are computed
  4. branch names allow backticks and $( )
  5. untrusted only matters at a sink

basics

~20 s

Anything a person typed is attacker-controlled: branch and tag names, commit messages, git author names, issue and comment bodies, pull request descriptions. Commit hashes, the repository name and the event type are produced by the platform, not by the submitter.

solid answer

~50 s

Split the event payload into fields a human chose and fields the platform computed. Free text a contributor supplies is attacker-controlled: the commit message, the git author and committer name and email, the branch or tag name, an issue or comment body, a pull request description, a label someone can add. Fields the forge derives are not: the commit hash, the event name, the repository and organisation identifiers, the run number. The distinction only matters at a sink. Reading `$COMMIT_MESSAGE` into a build log is harmless; splicing it into a command, a JSON body or a generated file makes it program text. Note that platform validation is not a safety boundary: a git ref cannot contain a space, but it can contain quotes, semicolons, backticks and `$( )` — plenty to close one command and start another.

go deeper

for a junior

Be ready to name three attacker-controlled fields and two platform-derived ones without hesitating, and to say that the git author name is simply whatever the contributor configured locally.

for a middle

Explain why platform validation is not a boundary, using the ref-name grammar as the concrete case, and state the rule that a value is harmful only once something parses it as syntax.

for a senior

Show how you turn this into review practice: a named list of never-interpolate fields, and a habit of asking who can set each value on a public repository versus a private one.

for a principal

Own the framing that this is a taint problem, not a character-filtering problem, and resist standards that enumerate bad characters — they age badly and give teams false confidence.

## The question behind the question Every CI job starts by reading an event: someone pushed, someone opened a pull request, someone commented, someone pushed a tag. The payload of that event is a mixture of two very different kinds of data, and the whole of expression injection begins with failing to separate them. **Attacker-controlled (free text a person supplied):** - the commit message and the commit body - the git author name, committer name, and both email addresses (these are whatever the contributor set locally; nothing verifies them) - the branch name and the tag name - a pull request title, description, and the name of the source branch - an issue body, an issue comment, a review comment - file paths and file contents in the diff **Platform-derived (not chosen by the submitter):** - the commit hash, and any tree or blob hash - the event name and the action that fired it - the base repository and organisation identifiers, the run identifier and attempt number - the timestamp the forge recorded A useful mental test: could I make this field say `hello"; whoami; #` by typing it? If yes, it is untrusted. ## Who the attacker is, and it is rarely the obvious person The positions differ per field, and that changes how much you should care. - **An anonymous internet user** can usually set a comment body, an issue body and a pull request description on a public repository. No account relationship with you is required. - **A low-privilege contributor** chooses the branch name of the pull request they open, and the commit message and author name of every commit in it. - **A colleague with write access** chooses tag names. "Only maintainers can push tags" narrows the attacker set; it does not empty it, and it does nothing against a compromised account or a bot token. - **A compromised automation account** — a bot that opens dependency pull requests, for example — supplies branch names and commit messages that no human ever reads. ## Validation is not a boundary The most common junior mistake is deciding a field is safe because the platform restricts it. Git ref names are the classic case: the ref grammar rejects spaces, control characters and a handful of symbols, so a branch cannot be called `my branch`. It can perfectly well be called `x$(id)` or ``x`id` `` or `x';whoami;'`. The characters that matter for command construction are almost all legal. The same reasoning applies to length limits, to "it must start with a letter", and to "our contributors are all employees". None of these were designed as security controls, and none of them survive an attacker who reads the grammar. ## Merging does not sanitize Engineers often assume that once a pull request is reviewed and merged, its metadata is trusted, so a job that runs on the default branch is safe. It is not. The merged commit carries the contributor's message and author name verbatim into the default branch's history. Reviewers read the diff; almost nobody reads the author field. A job triggered by the push to the default branch reads exactly the same untrusted strings, now with production-grade credentials in scope. ## Untrusted only matters at a sink Marking a field untrusted is not the same as banning it. The payload is genuinely useful — you want the branch name in a container tag, the author in a notification, the message in a changelog. The rule is about where the value is allowed to go: - **Safe:** passed to a program as an argument or through the process environment, written to a file with a serializer, printed as data. - **Dangerous:** substituted into a command line, a script body, a JSON or YAML document, a query, an HTML report, or any other text that something later parses as syntax. So the practical output of this triage is a short list — "these fields never get concatenated into anything" — that a reviewer can hold in their head. Everything else in this area is the consequence of that list being missing. ## In an interview Answer with the split, give two or three fields on each side, and volunteer the ref-name detail — it demonstrates that you understand why platform validation reassures people wrongly. Then close on the sink: the field is not harmful until something treats it as code.

  • A branch name cannot contain a space. Does that make it safe to interpolate into a command?
    No. The ref grammar rejects spaces and a few symbols, but permits quotes, semicolons, backticks and `$( )` — everything you need to terminate one command and start another. Platform validation was designed to keep refs parseable, not to keep them harmless, so treating it as a security boundary means relying on a rule that can change and was never yours.
  • The job only runs on the default branch after review and merge. Are the event fields still untrusted?
    Yes. A merge copies the contributor's commit message, author name and email into the default branch unchanged. Reviewers read the diff, not the author field, so nothing in the review path inspects those strings. The job now reads the same untrusted text with more privilege than the pull request build had.
  • Is it enough to reject payload values that contain shell metacharacters?
    Denylisting characters is the weakest option: you have to enumerate every metacharacter of every interpreter the value will reach, and you break legitimate input such as an apostrophe in a name. Prefer keeping the value out of program text entirely; where the shape is genuinely known — a semantic version, a hex hash — validate against an allowlist pattern instead.

saying these in an interview costs you the question

  • Believes only pull request titles are attacker-controlled
  • Assumes review and merge sanitize commit metadata
  • Calls branch names safe because they cannot contain spaces
  • Confuses the commit hash with the commit message
  • Trusts a field because only collaborators can set it

context

open as a page

When a fork's pull request triggers CI, what decides whether that run can read the base repository's secrets?

level: juniorimportance: must knowfreq 70%

basics

~20 s

The trigger event, and whose context it runs in. Runs in the fork's context get no repository secrets and a read-only token; a second family of events runs in the base repository's context with full secrets and a write token.

open as a page

What is poisoned pipeline execution, and how does direct PPE differ from indirect PPE?

level: juniorimportance: must knowfreq 60%

basics

~20 s

Poisoned pipeline execution means getting CI to run attacker-supplied code by changing something the build reads. Direct PPE changes the pipeline definition itself; indirect PPE changes a file the pipeline invokes, such as a build or test script.

open as a page

A release job builds a shell command from the pushed git tag; what does passing that tag through an environment variable fix, and what does it not?

level: middleimportance: must knowfreq 58%

basics

~20 s

An environment variable keeps the tag out of the script text, so the shell reads it as data instead of parsing it as code. It fixes nothing else: eval, or re-embedding the value in another interpreter, reopens the hole.

open as a page

Your fork-PR pipeline holds the chart publish credential and checks out the fork's Helm chart to lint it. What has that handed an attacker?

level: middleimportance: must knowfreq 58%

basics

~20 s

Use of the publish credential. Once the fork's files are on the runner, almost any tool run over them becomes attacker-controlled execution inside a job that already holds the credential, so the identity that publishes your charts becomes theirs.

open as a page

A third-party build step pinned to a commit digest downloads a toolchain tarball at run time. What does the pin still guarantee?

level: middleimportance: must knowfreq 62%

basics

~20 s

A commit digest fixes only the step's own source at that revision. Whatever the step downloads while running - a toolchain tarball, an installer, packages - is unpinned, so the pinned code is a loader for unpinned code.

open as a page

In indirect poisoned pipeline execution, which repo files are executable inputs that reviewers read as config?

level: middleimportance: should knowfreq 45%

basics

~20 s

Anything the job invokes rather than merely reads: build scripts, task-runner files, lint and test configuration that loads plugins, code-generation steps, and properties files that inject work into dependency restore. Reviewers see settings; the runner sees a program.

open as a page

A CI job passes an untrusted commit author name through an environment variable, then a later step builds a JSON webhook payload from it — what is still wrong?

level: seniorimportance: should knowfreq 34%

basics

~20 s

The environment variable only protects the shell. Concatenating the name into a JSON string makes it syntax for the JSON parser: a crafted author name closes the string and adds fields, forging the deploy record that the notification leaves behind.

open as a page

You gate privileged fork-PR runs behind maintainer approval or a safe-to-test label. How can a contributor still get newer code run?

level: seniorimportance: should knowfreq 44%

basics

~20 s

By pushing after the human looks. Approval and labels attach to the pull request, not to a commit, so the job resolves the branch head when it starts, and a gate that survives later pushes blesses code nobody reviewed.

open as a page

Branch protection covers main, but CI runs a credentialed job on every branch push. How do you map and close the PPE paths?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Enumerate every trigger — branch pushes, tag pushes, path-filtered runs, schedules, manual dispatch — and record for each who can cause it and what code it executes. Then remove credentials from any job an unreviewed ref can reach.

open as a page

What criteria admit a third-party build step into an infrastructure pipeline that holds cloud admin credentials?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Start from what the step can reach, not how popular it is: does it need a secret at all, who publishes it, is the pinned source readable, what does it do at run time, is a first-party alternative cheaper.

open as a page

In a Jenkins multibranch pipeline, why is building a Groovy string from the branch name dangerous?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

A double-quoted Groovy string is interpolated by the pipeline engine before the command text reaches the agent, so a branch name a contributor chose becomes part of the command. The fix: keep the value out of the program text.

open as a page

Merge requests from a contracted partner's fork run on your internal-network CI runners. What do you change, and what do you accept?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

A contract is not a technical control. Run partner merge requests like any untrusted fork: no credentials, disposable runners, no internal network position. If they genuinely need internal access, bring them inside the repository where it can be scoped and audited.

open as a page

When is vendoring a fork of a third-party build step into an internal repository worth its upgrade debt?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Vendor when you need what a pin cannot give: the ability to modify the step, durability if upstream disappears, and one choke point to patch. Otherwise pin, because a stale fork nobody upgrades is worse.

open as a page