Someone plants a compiler shim in a shared runner's tool cache. How do you detect that builds were tampered with?
answer
- the source stays honest
- what is inside versus how it was built
- rebuild clean and diff
- records must live off the host
- destroy the machine, do not clean it
basics
~20 sRarely from the artifact alone. Detection comes from rebuilding on known-clean infrastructure and comparing, from build records kept off the host that say which host and toolchain produced each artifact, and from integrity monitoring of the tool-cache path.
solid answer
~50 sStart by admitting the hard part: the source is honest, the lockfile is honest, and a list of declared components will not show a tool that was substituted on the host. Three things do catch it. A **comparison rebuild** on a clean single-use machine, diffed against the suspect artifact, as far as the build is deterministic enough to compare. **Build records stored off the host** - which host, which build image, which toolchain versions - so a build that ran somewhere it should not have stands out. And **file-integrity monitoring** of the tool cache and everything on `PATH`, shipped off-host as it happens, because whoever had write access to that machine also had write access to its local logs. Then bound the blast radius from job-to-host assignment records, destroy the host rather than cleaning it, and rotate everything it touched.
go deeper
Understand the basic shape: a build tool on the host can be replaced so the artifact differs from the reviewed source, and nothing in the repository would show it.
Explain why a list of what is inside an artifact cannot reveal a substituted tool, and why a record of how the artifact was built, stored off the host, can.
Show the incident judgment: comparison rebuild on clean infrastructure, off-host build records and integrity monitoring, blast radius from last-known-clean, destroy the host and rotate its credentials.
Own the tradeoff between paying for detection on shared hosts and paying to remove the class with single-use machines and immutable toolchains, and decide what evidence you must retain to answer a customer asking whether their release was affected.
## Why this class is nasty A shim is a small program placed where a build expects to find a real tool - a compiler, a linker, a language-version launcher. It calls the real tool and adjusts something on the way past. On a shared persistent runner, the tool cache is often writable by the build account, so one repository's job can plant a shim that every later job on that host silently uses. What makes this hard to see is that all the usual evidence stays clean. The source is what was reviewed. The lockfile pins what was declared. An inventory of what is inside the artifact describes the components that were **declared and packaged**, so a tool that altered code generation on the way through typically does not appear in it. This is precisely the gap that a record of **how the artifact came to be** exists to close, as opposed to a record of **what is inside it**. ## Detection, in order of usefulness ### 1. Rebuild somewhere clean and compare Build the same commit on a fresh single-use machine from a known image, and compare the output to the suspect artifact. Any difference that is not explained by non-determinism in the build is your signal. This works only as far as your builds are comparable at all - if timestamps, file ordering and embedded paths vary run to run, you may need to compare disassembly or specific sections rather than whole files. Even a partial comparison is powerful, because the attacker's change is usually in a place a diff will reach. ### 2. Build records that live off the host For every artifact, record and store externally: the job identity, the commit, the host or runner identity, the build image identity, and the toolchain versions actually used. Two anomalies then become queryable rather than invisible - an artifact built on a host that should not have taken that job, and a build whose reported toolchain does not match the image it claimed to run in. This only works if the record is produced by something the job cannot rewrite and is stored where the host cannot reach back and edit it. ### 3. Host integrity monitoring Baseline the tool cache, any language shim directory and everything on the default `PATH`, and alert on writes. On a multi-tenant build host, a write to a shared toolchain path by a job is almost never legitimate. Ship those events off the machine as they happen. The reason is the same one that makes this incident hard: whoever planted the shim had write access to the host, so anything stored only on the host - including its own logs and its own baseline database - is evidence you cannot fully trust. ## Bounding the damage The first question after detection is always "what else did this touch", and it is answerable only if you kept **job-to-host assignment records**. Without them, you cannot say which repositories built on that machine and are reduced to treating every artifact from the whole pool in the window as suspect. With them, the blast radius is every job that ran on that host from the earliest possible plant time - which you should treat as the last time you know the host was clean, not the time you first noticed - plus everything those artifacts were promoted into or bundled inside downstream. ## Response - **Destroy the host, do not clean it.** You do not know everything that was installed, and a cleanup performed with the attacker's tooling still on `PATH` proves nothing. Rebuild from a known image. - **Rotate every credential that host could reach**, including any long-lived tokens on disk, machine identities and cached registry credentials. - **Rebuild and republish** affected artifacts from clean infrastructure, and compare the new outputs against what you shipped. - **Preserve a disk image first** if you want forensics, but do it before rebuilding, not after. ## Prevention beats all of it The reason this question is asked at senior level is that the good answer ends by conceding that detection here is expensive and partial. The cheaper controls are structural: deliver the toolchain inside the build image rather than caching it on the host, make toolchain paths not writable by the job account, keep mutually untrusted repositories off the same host, and use single-use machines so a plant has no next job to catch. ## Saying it in an interview Open with "how would I even know" - naming that the source, the lockfile and a component inventory are all blind to a substituted tool - then give the three detections, then bound the blast radius from assignment records, then destroy rather than clean. A candidate who jumps straight to "we would see it in the scan results" has missed what happened.
- How far back does the blast radius go, and what lets you answer that?Back to the last point you can prove the host was clean, not to when you noticed. Every job that ran on it since then is suspect, plus every artifact those builds were promoted into or bundled inside. You can only scope it if you retained job-to-host assignment records off the machine; without them you must treat the whole pool's output for that window as suspect.
- Why not just clean the tool cache and put the runner back into service?Because you do not know the full set of changes. The plant may include a persistence mechanism in a shell profile, a scheduled task, or a modified agent, and any cleanup runs with the attacker's tooling potentially still on PATH. Rebuilding the machine from a known image gives you a state you can actually reason about; cleaning gives you a hope.
- Why are the runner's own logs weak evidence here?Because the attacker had write access to that host. Local logs, local audit trails and a locally stored integrity baseline are all in reach of the same account that planted the shim, so absence of evidence there proves nothing. Evidence has to be shipped off the machine as it is produced, to somewhere the host can write to but not rewrite.
saying these in an interview costs you the question
- Expects a component inventory to reveal a substituted build tool
- Trusts the compromised host's own logs as evidence
- Cleans the tool cache and returns the runner to service
- Dates the blast radius from discovery rather than last-known-clean
- Assumes a passing build means an untampered artifact