skip to content

Why is a build runner that keeps its state between jobs treated as a supply-chain risk, not just a flakiness risk?

level: juniorimportance: must knowfreq 60%

answer

  1. no boundary in time
  2. the previous job writes what you read
  3. PATH, caches, home directory, key stores
  4. different repo, different trust level
  5. contamination runs both directions

basics

~20 s

A reused runner lets one job write what the next job reads. Whatever an earlier job left on disk - tools on PATH, caches, unlocked keys - silently shapes the next artifact and exposes that job's secrets.

solid answer

~50 s

A persistent runner has no boundary in time. A job can write to the tool cache, the dependency cache, the shell profile, anything on `PATH`, and the home directory of the account jobs run as - and the next job, which may belong to a different repository at a different trust level, reads all of it. That makes the build partly controlled by earlier, less-trusted code: it can substitute a compiler wrapper, drop a package into a shared cache, or leave signing material unlocked. The flow runs both ways in time. The earlier job can tamper with the later job's artifact, and the later job can read what the earlier one left behind - checked-out source, a token written to disk, a credential that has not expired. Flaky builds are the benign version of the same mechanism; the security version is that you shipped an artifact partly built by code nobody reviewed.

go deeper

for a junior

Be ready to say plainly that a reused runner has no boundary between jobs, and to name three things that survive: the tool cache, the dependency cache and the home directory of the build account.

for a middle

Explain the two directions - an earlier job tampering with a later job's artifact, and a later job reading the earlier one's source and tokens - and why a workspace clean-up step touches neither.

for a senior

Show the production judgment: separate pools by trust tier, keep long-lived secrets off the host, make the toolchain immutable, and destroy rather than clean a host you suspect.

for a principal

Own the framing that reuse is a risk-acceptance decision, not an implementation detail: know which fleets cannot be made single-use, what compensating controls you require in exchange, and who signs off on the exception.

## What persistent means here A build runner is the machine or container that executes CI jobs. A **persistent** (or long-lived, or reused) runner takes job after job on the same operating-system installation: it registers once, picks up work, finishes, and waits for the next job with its disk intact. An **ephemeral** runner takes exactly one job and is then destroyed and replaced from a known image. The security property that separates them is simple: a persistent runner has **no boundary in time**. Everything a job writes is still there for whoever runs next. ## What a job can actually write More than people expect, because build tooling is designed to persist things for speed: - The **workspace** where the source is checked out. - The **tool cache** and language-version shims - the directories a build reads its compiler, SDK or interpreter from. - **Dependency caches** for the package ecosystems in use (npm, PyPI, Maven, and so on). These exist precisely so they survive between jobs. - Anything on `PATH`, plus shell profiles and environment files that later shells source. - The **home directory** of the account jobs run as: tokens, package-manager configuration, SSH keys and known-hosts, credential helpers. - Host-level agents and unlocked key stores - an agent holding a decrypted key, or a signing keychain a previous job unlocked and never re-locked. - The local image store, and background processes a job started that nobody reaped. ## Two directions of contamination **Earlier job attacks later job (integrity).** This is the one people miss. Consider a fleet of hardware-bound macOS hosts building mobile apps, where the tooling prefix is writable by the build account. One job replaces a build tool on `PATH` with a wrapper that calls the real tool and adds one line to the output. Every later job on that host builds an artifact the reviewer never approved. The source is clean, the lockfile is clean, the review was honest - and the binary is not. **Later job reads earlier job's leftovers (confidentiality).** The other direction is theft. A job that lands on that same host after a release build finds the checked-out source of a repository it has no access to, a token a step wrote to a file, a package-manager configuration with an upload credential in it, or a key store left unlocked and usable without a passphrase. ## Why the trust level is the crux Contamination only matters because the jobs on a shared pool are **not equally trusted**. A pool that takes work from many repositories mixes a release build that holds signing material with a job triggered by a low-privilege contributor's pull request. The pull-request job is arbitrary code - the contributor can change the build script, add a test, or add a dependency whose install hook runs. Once that code runs on the host, the trust boundary you thought you had between the two repositories does not exist, because the operating system underneath them is shared and writable. ## Why the flakiness framing is not enough Teams meet this first as flakiness: a build passes on the shared machine and fails on a fresh one, so someone adds a clean-up step and moves on. That reasoning fails in two ways for security. First, a clean-up step that deletes the workspace does not touch tool caches, the home directory, package prefixes, or a running agent. Second, a clean-up step runs **after** the untrusted job - the attacker's code has already executed with write access to the host and may have disabled or subverted the clean-up itself. A control that the attacker runs before is not a control. ## What actually helps In rough order of strength: 1. **One job per host.** Destroy the machine after the job and rebuild it from a known image. This is preventive and removes the class rather than the instance. 2. **Do not mix trust levels on a pool.** If some hosts must persist, keep them single-tenant - one repository, or one trust tier - so contamination stays inside a blast radius you already accepted. 3. **Keep long-lived secrets off the host.** Issue short-lived, job-scoped credentials at job start; have signing done by a service the host calls rather than by key material sitting on the host. 4. **Make the toolchain immutable** - delivered in the build image and not writable by the job account. 5. **Record which job ran on which host.** That is detective, not preventive, but without it you cannot bound the damage later. ## Saying it in an interview Name both directions - the earlier job tampering with the later job's artifact, and the later job harvesting the earlier job's secrets - and say that the fix is a fresh machine, not a cleaner script.

  • If the pipeline deletes the workspace at the end of every job, is the problem solved?
    No. The workspace is one of several writable locations; tool caches, dependency caches, the home directory, package prefixes, the local image store and any process a job left running all survive it. Worse, a clean-up step runs after the untrusted job has already executed, so attacker code could disable it. Clean-up is hygiene, not a boundary.
  • Every job on the runner executes as the same OS user. Why does that matter?
    Because file permissions then give you nothing between jobs. One shared account means every job can read and write every other job's files, agents and configuration by design. Separate accounts help a little, but they share the kernel, the tool cache and any root-owned agent, so they narrow the problem rather than removing it.
  • The pool only serves private repositories. Does that make reuse safe?
    No. Private means the code is not public, not that everyone who can trigger a build is trusted with everything the pool touches. A contributor with write access to one low-value repository, a compromised developer account, or a malicious dependency's install hook all put code on the host. Private repositories change who the candidate attackers are, not whether contamination works.

It is a shared kitchen with no cleaning between shifts. The cook before you may have swapped the salt for something else, and you may find their wallet in the drawer.

saying these in an interview costs you the question

  • Says leftover state only causes flaky builds, never attacks
  • Assumes wiping the workspace makes the host clean
  • Thinks a private-repo-only pool means every job is trusted
  • Believes secrets disappear when the job ends
  • Treats a clean-up step run after untrusted code as a control

context