skip to content

In a CI pipeline, which parts of the triggering event are attacker-controlled, and which are not?

level: juniorimportance: must knowfreq 68%

answer

  1. who typed it, a person or the platform?
  2. free text is the dangerous half
  3. hashes and event names are computed
  4. branch names allow backticks and $( )
  5. untrusted only matters at a sink

basics

~20 s

Anything a person typed is attacker-controlled: branch and tag names, commit messages, git author names, issue and comment bodies, pull request descriptions. Commit hashes, the repository name and the event type are produced by the platform, not by the submitter.

solid answer

~50 s

Split the event payload into fields a human chose and fields the platform computed. Free text a contributor supplies is attacker-controlled: the commit message, the git author and committer name and email, the branch or tag name, an issue or comment body, a pull request description, a label someone can add. Fields the forge derives are not: the commit hash, the event name, the repository and organisation identifiers, the run number. The distinction only matters at a sink. Reading `$COMMIT_MESSAGE` into a build log is harmless; splicing it into a command, a JSON body or a generated file makes it program text. Note that platform validation is not a safety boundary: a git ref cannot contain a space, but it can contain quotes, semicolons, backticks and `$( )` — plenty to close one command and start another.

go deeper

for a junior

Be ready to name three attacker-controlled fields and two platform-derived ones without hesitating, and to say that the git author name is simply whatever the contributor configured locally.

for a middle

Explain why platform validation is not a boundary, using the ref-name grammar as the concrete case, and state the rule that a value is harmful only once something parses it as syntax.

for a senior

Show how you turn this into review practice: a named list of never-interpolate fields, and a habit of asking who can set each value on a public repository versus a private one.

for a principal

Own the framing that this is a taint problem, not a character-filtering problem, and resist standards that enumerate bad characters — they age badly and give teams false confidence.

## The question behind the question Every CI job starts by reading an event: someone pushed, someone opened a pull request, someone commented, someone pushed a tag. The payload of that event is a mixture of two very different kinds of data, and the whole of expression injection begins with failing to separate them. **Attacker-controlled (free text a person supplied):** - the commit message and the commit body - the git author name, committer name, and both email addresses (these are whatever the contributor set locally; nothing verifies them) - the branch name and the tag name - a pull request title, description, and the name of the source branch - an issue body, an issue comment, a review comment - file paths and file contents in the diff **Platform-derived (not chosen by the submitter):** - the commit hash, and any tree or blob hash - the event name and the action that fired it - the base repository and organisation identifiers, the run identifier and attempt number - the timestamp the forge recorded A useful mental test: could I make this field say `hello"; whoami; #` by typing it? If yes, it is untrusted. ## Who the attacker is, and it is rarely the obvious person The positions differ per field, and that changes how much you should care. - **An anonymous internet user** can usually set a comment body, an issue body and a pull request description on a public repository. No account relationship with you is required. - **A low-privilege contributor** chooses the branch name of the pull request they open, and the commit message and author name of every commit in it. - **A colleague with write access** chooses tag names. "Only maintainers can push tags" narrows the attacker set; it does not empty it, and it does nothing against a compromised account or a bot token. - **A compromised automation account** — a bot that opens dependency pull requests, for example — supplies branch names and commit messages that no human ever reads. ## Validation is not a boundary The most common junior mistake is deciding a field is safe because the platform restricts it. Git ref names are the classic case: the ref grammar rejects spaces, control characters and a handful of symbols, so a branch cannot be called `my branch`. It can perfectly well be called `x$(id)` or ``x`id` `` or `x';whoami;'`. The characters that matter for command construction are almost all legal. The same reasoning applies to length limits, to "it must start with a letter", and to "our contributors are all employees". None of these were designed as security controls, and none of them survive an attacker who reads the grammar. ## Merging does not sanitize Engineers often assume that once a pull request is reviewed and merged, its metadata is trusted, so a job that runs on the default branch is safe. It is not. The merged commit carries the contributor's message and author name verbatim into the default branch's history. Reviewers read the diff; almost nobody reads the author field. A job triggered by the push to the default branch reads exactly the same untrusted strings, now with production-grade credentials in scope. ## Untrusted only matters at a sink Marking a field untrusted is not the same as banning it. The payload is genuinely useful — you want the branch name in a container tag, the author in a notification, the message in a changelog. The rule is about where the value is allowed to go: - **Safe:** passed to a program as an argument or through the process environment, written to a file with a serializer, printed as data. - **Dangerous:** substituted into a command line, a script body, a JSON or YAML document, a query, an HTML report, or any other text that something later parses as syntax. So the practical output of this triage is a short list — "these fields never get concatenated into anything" — that a reviewer can hold in their head. Everything else in this area is the consequence of that list being missing. ## In an interview Answer with the split, give two or three fields on each side, and volunteer the ref-name detail — it demonstrates that you understand why platform validation reassures people wrongly. Then close on the sink: the field is not harmful until something treats it as code.

  • A branch name cannot contain a space. Does that make it safe to interpolate into a command?
    No. The ref grammar rejects spaces and a few symbols, but permits quotes, semicolons, backticks and `$( )` — everything you need to terminate one command and start another. Platform validation was designed to keep refs parseable, not to keep them harmless, so treating it as a security boundary means relying on a rule that can change and was never yours.
  • The job only runs on the default branch after review and merge. Are the event fields still untrusted?
    Yes. A merge copies the contributor's commit message, author name and email into the default branch unchanged. Reviewers read the diff, not the author field, so nothing in the review path inspects those strings. The job now reads the same untrusted text with more privilege than the pull request build had.
  • Is it enough to reject payload values that contain shell metacharacters?
    Denylisting characters is the weakest option: you have to enumerate every metacharacter of every interpreter the value will reach, and you break legitimate input such as an apostrophe in a name. Prefer keeping the value out of program text entirely; where the shape is genuinely known — a semantic version, a hex hash — validate against an allowlist pattern instead.

saying these in an interview costs you the question

  • Believes only pull request titles are attacker-controlled
  • Assumes review and merge sanitize commit metadata
  • Calls branch names safe because they cannot contain spaces
  • Confuses the commit hash with the commit message
  • Trusts a field because only collaborators can set it

context