skip to content

Your CI must post a coverage comment on pull requests opened from forks, so a team proposes running the fork's build in a job that already holds the repository write token. Why is that dangerous, and what is the safe pattern?

level: seniorimportance: should knowfreq 44%

answer

  1. fork code is untrusted code
  2. the whole build executes, not just tests
  3. trusted pipeline plus untrusted checkout
  4. two stages, artifact in between
  5. the artifact is data, not instructions

basics

~20 s

A fork pull request is untrusted code, and any credential present in the job's environment is available to it — build scripts, test hooks and dependency install hooks all execute before anyone reviews the diff. Split it: run the untrusted build with no credentials, hand its output to a separate trusted job that never checks out the fork's code.

solid answer

~60 s

The proposal inverts the trust model. A pull request from a fork contains arbitrary code from someone with no write access, and CI runs it — not just the tests, but the build scripts, the test configuration, dependency lifecycle hooks, and anything the pipeline file on that branch says. Whatever credential is in that job's environment is therefore the contributor's credential: it can be printed in an obfuscated form, sent to an external host, or used directly to push to your repository. Log masking is a display backstop, not a boundary. The safe pattern is a two-stage handoff. Stage one executes the untrusted code with no secrets, a read-only token, no route to internal networks, and on a disposable runner — never on a self-hosted machine inside your network — and writes its result to an artifact. Stage two is triggered by that run finishing, uses the pipeline definition from your default branch, never checks out the fork's code, downloads the artifact, treats its contents as untrusted data, and posts the comment with the write token.

go deeper

for a junior

Understand that CI executes the contributor's code, not just their tests, so a pull request from a stranger's fork must not run with access to any secret.

for a middle

Explain what is actually executable in a pull request — build scripts, dependency install hooks, test configuration — and why a job holding a write token while building that code hands the token to its author.

for a senior

Design the split: an unprivileged stage that produces an artifact and a trusted stage triggered by its completion that never checks out the fork's code, plus runner-pool separation and approval gating for first-time contributors.

for a principal

Set the policy across repositories: which pools may accept outside contributions, what the maximum token scope on a pull-request path is, and when a convenience feature is simply not worth the trust boundary it would require.

## The rule underneath the question There is one rule and everything else follows from it: **whatever a job can reach, the author of the code that job runs can reach.** A pull request from a fork is code written by someone you have granted no access to. If the job that executes it holds a write token, you have granted them write access — you have simply routed it through your CI system. The subtlety candidates miss is *how much* of the pull request is executable. It is not only the test files. It is the build script, the task runner configuration, the dependency manifest with its install hooks, the linter plugins, the compiler plugins, and — depending on the platform — the pipeline definition itself. A malicious contribution does not need to look malicious; a single line in a lifecycle hook of a dependency, or a one-character change to a build script, executes before a human ever reads the diff. ## Why the tempting shortcut exists The feature people want is legitimate: post a coverage or benchmark comment, apply a label, update a status. All of those need a credential with write permission, and the job that produced the numbers is the obvious place to use it. So teams reach for the mechanism their platform offers for running a pipeline in a privileged context on pull requests — the pipeline definition is taken from the *target* branch, which is trusted, and the elevated token is available. That part is sound. The mistake is the next line: adding an explicit checkout of the pull request's head so the job can actually build the contributor's code. Now trusted pipeline definition plus untrusted source code plus a write token are in one execution context, and the trust of the definition is irrelevant, because the definition tells the job to run the contributor's build. GitHub Actions names such a trigger `pull_request_target`, and the same shape exists on any platform that lets you take the pipeline from one ref and the code from another. The signature of the bug is always identical: *the pipeline is trusted, the checkout is not, and they run together with a credential.* ## The two-stage pattern Split the work along the trust line. **Stage one — untrusted execution.** Triggered by the pull request itself. It checks out the contributor's code and runs the build and tests with: - no repository secrets and no cloud credentials of any kind; - a token scoped to read the public repository, or no token at all; - a disposable runner from a pool that never touches anything internal — a provider-hosted instance is the natural choice, and self-hosted machines inside your network must be excluded outright; - no route to internal registries, databases or metadata services; - and a required maintainer approval before it runs at all, for contributors who have not landed a change before. It writes its result — the coverage number, a benchmark JSON, a report file — as an artifact, along with the pull request number. **Stage two — trusted publication.** Triggered by stage one *completing*, not by the pull request. It runs the definition from your default branch, so a contributor cannot change what it does. It never checks out the fork's code. It downloads the artifact and posts the comment using the write token. ## The part that is still sharp The artifact crossing that line is attacker-controlled data. Stage two must handle it as data: read the file, parse it, validate the shape and bounds, and never interpolate its contents into a shell command, a template that gets evaluated, or a comment body that could carry markup with side effects. Also validate the pull-request number itself — if stage two takes a number from the artifact and posts to it, a contributor can direct the comment anywhere. Give stage two the narrowest token you can: permission to write comments on pull requests, nothing else. If it can only comment, the worst outcome of a mistake is a rude comment rather than a push to your default branch. ## Simpler answers that are often better Before building the handoff, ask whether the comment is worth it. Two alternatives are frequently sufficient: publish the result as a check or status from a trusted context that reads the numbers from your own storage, or accept that fork pull requests get a link to the run instead of a rendered comment. A large fraction of the incidents in this category exist because a convenience feature was allowed to determine the trust boundary.

  • Why is masking the token in logs not a defence here?
    Masking is a display filter applied to the log stream. The value is still a real credential inside the process, so the code can use it directly, or encode it — base64, reversed, split across lines — so the filter no longer matches what is printed. It protects against accidental echoing by trusted code, not against code that wants the value.
  • What should stage two verify about the artifact it receives?
    That it is the shape expected: parse it strictly, reject unexpected fields or sizes, bound the numbers, and never pass its contents to a shell or an evaluated template. Also verify the pull-request identifier independently rather than trusting the one in the artifact, so a contributor cannot redirect the comment to another pull request.
  • Which runners should fork pull requests be allowed to use?
    Disposable ones with no privileged access: a provider-hosted instance, or a self-hosted pool that is single-use, network-isolated from internal systems and holds no other pipeline's credentials. Never a persistent self-hosted machine on your internal network — it turns an untrusted contribution into code execution behind your perimeter with whatever the last job left on disk.

saying these in an interview costs you the question

  • Only the test files run, so a diff review is enough
  • The token is masked in the log, so it cannot leak
  • Fork pull requests are safe because we require an approval to merge
  • A read-only checkout makes the untrusted code harmless
  • Just run fork builds on our self-hosted runners like everything else

context