A build passes on your team's long-lived shared CI machines but fails on a freshly provisioned one. What kinds of leftover state on a reused runner cause that, and what removes the whole class of problem?
answer
- undeclared dependency on leftovers
- the home directory, not just the workspace
- cleanup steps are skipped when it matters
- one job, one machine, then destroy
- warm state is also leaked state
basics
~20 sA reused runner carries state forward: leftover workspace files, globally installed tools, package-manager config and credentials in the agent user's home directory, stray processes, and cached images. The build silently depends on it. Single-use runners — one job per instance, then destroyed — remove the class.
solid answer
~50 sThe build has an undeclared dependency on something a previous job left behind. The usual suspects, in rough order of frequency: files in the reused workspace that are not in version control (generated code, stale build output, an untracked config file); tools installed globally by an earlier job rather than declared in the manifest; configuration and credentials in the agent user's home directory, such as a package-registry token or a container-registry login; background processes and bound ports; and leftover containers, volumes and images. The same mechanism is a security problem, not just a flakiness problem — whatever the last job left readable, the next job can read, and the next job may belong to another team. The structural fix is a single-use runner: an instance registered for one job and destroyed after it, so nothing can outlive a job. Cleanup steps are a weaker substitute because they are skipped exactly when a job crashes or is cancelled.
code
bash · 9 lines# Job 1 (earlier, on a reused runner)
npm install -g some-internal-cli
npm config set //registry.internal/:_authToken "$REGISTRY_TOKEN" # writes ~/.npmrc
nohup ./test-fixtures/mock-api --port 8081 & # never stopped
# Job 2 (later, same machine, different pipeline)
which some-internal-cli # found: left by Job 1, not declared anywhere
cat ~/.npmrc # Job 1's registry token, still readable
curl localhost:8081/health # tests pass against Job 1's stray processgo deeper
Recognise the symptom: if a build works only on one particular machine, suspect something installed or left there rather than something in the repository. Know that a clean machine per job avoids it.
Enumerate the state that survives a job — workspace, global tools, home-directory config and credentials, stray processes, cached images — and explain why a trailing cleanup step does not cover it.
Treat workspace reuse as a credential-disclosure path, not just flakiness. Argue for single-use instances, pool separation by trust domain, and an explicit cache strategy that replaces the accidental warm state you gave up.
Own the fleet policy and its cost: single-use runners raise per-job provisioning time and cache-service spend, so decide where that is mandatory, how runner images are built and rotated, and how teams are stopped from re-introducing shared mutable machines for speed.
## The two failures of a reused runner A persistent runner executes many jobs on the same machine, in the same working directory, as the same user. That produces two distinct failures with one cause. The first is the *dirty workspace*: a build succeeds only because of something an earlier job left behind, so it fails anywhere else — on a colleague's laptop, on a new machine, or the day the runner is replaced. The build has an undeclared dependency, and the shared runner has been hiding it. The second is *cross-job disclosure*: whatever one job leaves readable, the next job can read. If those two jobs belong to different teams, one team's pipeline has just read the other's credentials with no exploit involved. ## Where the state actually hides Candidates usually name the workspace and stop. The full list is longer, and the entries below the workspace are the ones that cause the confusing failures. - **The workspace directory.** Untracked files, generated sources, stale compiled output, an old dependency tree that satisfies an import nobody declared. A clean checkout into a dirty directory is not a clean state. - **Globally installed tools.** A job that ran `npm install -g` or `pip install --user` or added a system package leaves that tool on the PATH for every later job. - **Version-manager state.** A tool-version manager switched to a version and left it as the default, so later jobs get a runtime nobody selected. - **Home-directory configuration and credentials.** This is the dangerous one: a package-registry token, a container-registry login, a cloud CLI's cached session, a git credential helper store, an ssh-agent socket. These are files written into the agent user's home by ordinary, correct steps, and they simply stay there. - **Processes and ports.** A database or mock server started for tests and never stopped will make the next job's tests pass — or fail with a port collision. - **Container and image state.** Cached images that no longer match their tag, dangling volumes, and eventually a full disk. - **The machine itself.** OS packages, DNS or hosts-file edits, kernel parameters, and a clock or locale nobody set deliberately. ```bash # Job 1, last week, on a reused runner: npm install -g some-internal-cli # global, not in package.json npm config set //registry.internal/:_authToken "$TOKEN" # writes ~/.npmrc # Job 2, today, same machine, different team: some-internal-cli build # "works" only because Job 1 installed it cat ~/.npmrc # and Job 1's token is still readable here ``` ## What actually removes the class The structural answer is a **single-use runner**: the instance is created (or registered) for exactly one job and destroyed when that job ends, whether it passed, failed, was cancelled, or the machine was killed mid-step. Nothing survives, because the thing that would hold the state no longer exists. Platforms implement this as an ephemeral or just-in-time registration mode, usually driven by an autoscaler that provisions a fresh VM or container per queued job. Weaker substitutes, in descending order of strength: 1. **Run each job inside a fresh container** on a persistent host. The workspace, PATH, home directory and processes are new; the host disk and any mounted caches are still shared, and a privileged job still escapes. 2. **Wipe the workspace at job start** rather than at the end. Better than a trailing cleanup, but it touches only the first item on the list above. 3. **A cleanup step at the end of the job.** The weakest, because it is skipped precisely when it matters — a cancelled job, a crashed step, a runner that lost its connection. Never rely on it for anything security-relevant. If you must keep persistent runners, at least separate pools by trust domain, so a job that leaves a credential behind can only leak it to jobs that were already entitled to it. ## The cost, and how to pay it honestly Single-use runners give up the warm state that made builds fast: the dependency cache, the container image cache, sometimes the source checkout itself. That is a real cost, and the answer is to make the fast path explicit rather than accidental — restore dependencies from the platform's cache service at the start of the job, pull base images through a local registry mirror, and bake slow-moving toolchain into the runner image. The point of the exercise is not that caching is bad; it is that a cache you declare is reproducible, while a cache that is merely left over is not.
- Why is a cleanup step at the end of a job a weak defence?Because it only runs when the job reaches the end. A cancelled run, a step that kills the shell, an out-of-memory crash, or a lost connection to the coordinator all leave the machine exactly as the job left it — and those are the cases most likely to have secrets on disk. Cleanup belongs to the lifecycle that destroys the instance, not to the job.
- If single-use runners lose the warm cache, how do you keep builds fast?Make the fast path explicit. Bake slow-moving toolchain into the runner image so provisioning does not install it, restore dependencies from the platform's cache service or a shared remote cache at job start, and pull base images through a local registry mirror. Each is declared and reproducible, unlike state that merely happened to survive.
- You cannot make the runners single-use yet. What is the highest-value interim control?Run every job inside a fresh container on the host, so the workspace, PATH, home directory and process table are new each time, and separate runner pools by trust domain so a leaked credential can only reach jobs already entitled to it. Also drop privileged mode and host socket mounts, which would defeat the container boundary anyway.
A persistent runner is a shared kitchen where nobody washes up: your recipe works because someone else's ingredients are still on the counter, and your notes are still readable by whoever cooks next.
saying these in an interview costs you the question
- Deleting the workspace at the end is enough cleanup
- Only the checkout directory carries state between jobs
- Reused runners are just a speed optimisation, not a security issue
- A fresh git clone guarantees a clean build environment
- Ephemeral runners are impossible because caching would break