In a CI pipeline, what is cache poisoning and why does the build not notice it?
answer
- a cache is a build input
- restore is not rebuild
- trust level equals whoever can write
- no signature on a cache blob
- green build, poisoned bytes
basics
~20 sCache poisoning means getting tampered content into a cache entry a later build restores. The build treats restored files as if it had produced them: nothing is signed, nothing is compared, and the log shows only a cache hit.
solid answer
~50 sA CI cache is a key-to-blob store. A job asks for a key, and if an entry exists the stored directory is unpacked into the workspace before the build runs. That means the cache is an **input** to the build, and its trust level is whoever can run a job that writes that key. Poisoning is writing attacker-chosen bytes under such a key: a patched dependency tree, a swapped compiled object, a tampered generated file. The build does not notice because a restore performs no verification beyond matching the key. There is no signature on a cache blob, no producer identity, and typically no re-check of extracted or compiled contents against any manifest. The log line says `cache hit`, not what came back. The failure is silent by design: the build stays green and the poisoned bytes are linked into whatever it publishes.
go deeper
Be ready to say what a cache restore actually does: match a key, unpack a directory, run the build. Then say the one consequence that matters: nothing about the contents was checked.
An interviewer expects you to explain where verification stops - checksum files cover downloaded archives, not extracted trees or compiled outputs - and why a cache hit line in a log proves nothing about content.
Show that you treat the cache as an input with a trust level, and that you know detection is weak: prevention is scoping and write permissions, with cold-build comparison as the only cheap after-the-fact check.
Own the rule that decides policy: a cache may hold only what could be recomputed, written only by something trusted to compute it. That is what lets you keep caching at all instead of banning it for speed's sake.
## What a cache is, mechanically Every CI system offers roughly the same thing: a job computes a **key**, asks a shared store whether an entry exists for it, and if one does, the stored directory is unpacked into the workspace before the real work starts. At the end of the job the directory may be packed back up and stored under the key. The purpose is speed — skip a dependency download, skip recompiling code that has not changed. The security-relevant fact is what that means for the build's inputs. A cached dependency tree, a cached compiled-object directory, a cached generated-code folder — all of these are consumed exactly as if the job had produced them itself. They get compiled, linked, packaged and shipped. So the cache is not an optimisation sitting beside the build; it is a **source of build inputs**, and it deserves the same question you would ask of any input: who could have put this here? ## What poisoning means Cache poisoning is arranging for a cache entry that a later, more privileged build will restore to contain content of your choosing. Concretely that can be: - a **patched dependency**: the extracted tree of a library with one function altered; - a **swapped compiled artifact**: an object file or class file in a compile cache, which no dependency manifest describes at all; - a **generated or downloaded auxiliary file**: a code-generation output, a schema, a data file the build embeds; - a **tool** placed on the path from a cached tool directory. The attacker does not need to break the cache store. In most setups they only need to be able to *run a job that writes that key* — which may mean pushing to a throwaway branch, opening a pull request, or being any of several hundred engineers with commit rights somewhere in the org. ## Why the build does not notice Three properties combine: 1. **A restore verifies nothing but the key.** The entry is not signed, and the restoring job has no record of who wrote it or what it should contain. 2. **Checksum coverage stops early.** A lockfile-style checksum file records hashes of the archives a resolver downloads. A cache usually holds the *extracted* tree and, worse, *compiled outputs* — and nothing in the project records what those should hash to. A cache key computed from the checksum file makes the entry look integrity-protected when the bytes inside it are not covered by it at all. 3. **The log is unhelpful.** A hit prints one line. Nobody diffs a restored directory against a fresh download, because the entire point of the cache was not to do that work. The result is that a successful poisoning looks like a completely normal green build. The wrong mental model — "if the cache were corrupt the build would fail" — is what makes the attack worth mounting. A corrupted cache that breaks the build is the *benign* outcome; the dangerous one is a build that succeeds and ships. ## Cache versus inter-job artifact It helps to separate two things that look similar. A **cache** is best-effort and persists across runs: a miss must be survivable, and its contents are supposed to be reconstructible from declared inputs. An **inter-job artifact** is a required handoff inside one run: the deploy job cannot proceed without it. Both are unverified stores handed to a later consumer, but their blast radii differ. A poisoned cache entry can affect every future build that computes the same key; a substituted artifact affects the run it lives in. ## The containment principle The rule that makes the rest follow: **a cache may only hold things you could recompute, and it may only be written by something you would trust to compute them.** Two consequences: - Scope entries so the trust level of the writer matches the trust level of the reader. If a job on any branch can write an entry that a release build reads, the release build has inherited that branch's trust level. - Keep the cache droppable. If deleting the whole cache changes the output of a build, you are not caching, you are storing an unreproducible input. Detection is weak here, so prevention carries the weight. What you can usefully record is *which run and which ref wrote each key*, so that after an incident you can ask which builds restored it, and a periodic cold build (cache disabled) whose output you compare against the warm build tells you whether the cache is changing results at all. ## What to say in an interview Name the cache as an input, name the trust level as "whoever can write the key", and say plainly that a restore checks nothing but the key. Then note that the attack is invisible: the build is green. That framing gets you to the design question — scoping and writability — which is where the conversation actually goes.
- If the cache only holds third-party dependencies that the resolver already checksum-verified, can it still be poisoned usefully?Yes. Checksum files cover the archives a resolver downloads, not the extracted trees or the compiled outputs a cache typically stores, and most builds do not re-hash extracted content on every run. A cache that holds compiled objects is worse still: nothing in the project records what those bytes should be, so there is no comparison for a restore to fail.
- How would you tell after the fact whether a given build restored a poisoned entry?You need two records that most setups do not keep by default: which key each build restored and hit, and which run and ref wrote that key. With those you can enumerate every build that consumed a suspect entry. A cold build with caching disabled, compared against the warm build's output, tells you whether the cache is currently changing results.
- Is deleting the cache after an incident enough of a response?It stops further spread but does not tell you what already shipped. You still have to identify the builds that restored the entry, re-run them cold, and compare the artifacts they now produce against what was released. And unless you also fix who can write the key, the next job on an untrusted ref simply reseeds it.
A cache restore is like accepting a package left on your doorstep because the label matches your name. You never asked who put it there, and nothing about the label proves what is inside.
saying these in an interview costs you the question
- Assuming a tampered cache would make the build fail
- Believing cache entries are signed or verified on restore
- Saying checksum files already cover everything a cache holds
- Treating caching purely as a performance concern
- Thinking only the cache store's operator can poison an entry