A Go build cache keyed only on the go.sum hash is shared by every branch's CI jobs - what is the risk?
answer
- the key describes inputs, not authorship
- same key on every branch
- compiled outputs are covered by nothing
- prefix fallback widens the family
- namespace by trust class, enforce on write
basics
~20 sAny engineer who can push a branch computes the same key, writes that entry, and the release build restores it. A key describes content inputs, never who produced them, so one namespace across all branches erases the trust boundary.
solid answer
~50 sThe key answers "what were the inputs?" and nothing else. A hash of `go.sum` is identical on a throwaway branch and on the release branch, so a low-privilege insider with write access to any branch can run a job that populates the entry the release build later restores. Two things make it worse in a Go project: `go.sum` covers the module archives the toolchain downloads, not the **compiled outputs** the build cache stores, so the key looks integrity-bearing while its contents are covered by nothing; and prefix-fallback restore keys let a trusted build accept the newest entry merely *starting* with a prefix, widening the set of entries it will take. The fix is to make the namespace, not just the key string, carry trust: separate namespaces for protected and unprotected refs, with untrusted refs read-only against the trusted one, and no cross-class fallback.
go deeper
Know that a cache key is computed from build inputs and that identical inputs on any branch produce the same key. That alone explains how one branch's job can seed another build's cache.
Be ready to explain what the key encodes and what it omits, where checksum coverage stops, and how prefix-fallback restore keys let a build accept entries it never intended to.
Demonstrate that you separate namespaces by trust class and enforce it through write permissions rather than key strings, and that you keep records of which ref wrote and which build restored each key.
Own the tradeoff with the platform team: which cache classes may be shared org-wide for speed, which must be per-trust-class, and how you say no to a shared compile cache without killing build performance.
## What a cache key is and is not A cache key is a string a job computes, usually from the things that determine the cached content: an OS identifier, a toolchain version, and a hash of a dependency manifest or checksum file. It answers exactly one question — *what inputs was this content derived from?* It answers none of these: - who produced the entry; - from which ref, run or identity; - whether the producer was allowed to influence a release. In a system where the key alone decides what you get, **the trust level of every restore equals the trust level of the least-trusted principal who can write that key**. Share one namespace org-wide and that is "anyone who can push a branch". ## Walking the concrete attack A Go service caches its module cache and build cache directories under a key like `go-<os>-<hash of go.sum>`, stored in a project-wide namespace. An engineer with commit rights — no admin, no release privileges — pushes a branch that does not change dependencies, so `go.sum` is untouched and the key is byte-identical to the release build's key. Their job runs, does whatever they wrote, and the cache-save step packs the directory back up under the shared key. The release build starts, computes the same key, hits, and unpacks a directory containing at least one compiled package archive that did not come from the source it just checked out. The resulting binary is signed, published and deployed, and every check the pipeline runs passes. Note the attacker position: not an anonymous outsider, but an authenticated low-privilege insider using an entirely legitimate ability — to push a branch and have CI run. Nothing about that is anomalous in the logs. ## Why the checksum file is a false comfort It is tempting to argue that a Go build cannot be poisoned this way because `go.sum` pins module content. Be precise about coverage: | What is cached | What records its expected bytes | | --- | --- | | Downloaded module archives | the project's checksum file | | Extracted module trees | nothing, once extraction happened | | Compiled package archives (the build cache) | nothing at all | Compiled outputs are the whole point of a build cache and are the part no manifest describes. So keying on the checksum file makes the entry *look* integrity-bound while the most valuable content in it is unverifiable. The same holds in any ecosystem whose cache holds compiled or generated output rather than only downloads. ## Restore-key fallbacks widen the target Most cache implementations support a fallback: if the exact key misses, accept the newest entry whose key begins with a given prefix, so a dependency bump still gets a warm start. Security-wise this means the trusted build no longer restores *one* entry it can reason about; it restores whichever of a whole prefix-matching family happens to be newest. An attacker no longer has to match the release build's key exactly — they only have to write something under the prefix, recently. Every widening of the fallback widens the blast radius. ## The fix, in the order it actually helps 1. **Namespace by trust class.** Entries written by jobs on protected refs live in one namespace; everything else lives in another. Trusted builds read only the trusted namespace. Untrusted refs may *read* it for speed but never write it. 2. **Enforce it with permissions, not string discipline.** Putting a branch name in the key is not a boundary: keys are attacker-chosen strings, and someone can compute whatever string they like. What separates trust classes is who is allowed to *write* into a namespace, decided by the CI system's credentials, not by convention. 3. **Constrain fallback across the boundary.** Prefix fallback within a trust class is fine; falling back from the trusted namespace into anything an untrusted job could have written is not. 4. **Cache only what is reconstructible, and keep it droppable.** If wiping the cache changes a build's output, the cache is holding an input, not a derived result — and that must be fixed before any scoping helps. 5. **Consider what belongs in the cache at all.** Downloads are cheap to re-verify; compiled outputs are not. Some teams cache dependency downloads across trust classes and keep compile caches strictly per-trust-class, which is a defensible split. ## Recording, for the day you need it Even with scoping, log which run and which ref wrote each key and which builds restored it. Without that pair of records, the question "which releases could have consumed this entry?" has no answer, and the only honest response to an incident is to rebuild everything cold. ## The one-line version A cache key describes content; a cache namespace decides trust. Design the key for correctness and the namespace for security, and never let one namespace span two trust levels.
- Does adding the branch name to the cache key fix this?Only partly, and not as a security control. It stops accidental sharing, but a key is just a string the job computes, so anyone who can run a job can compute any string. And if prefix fallback is enabled, the release build will still accept another branch's entry. The boundary has to be enforced on write permissions to a namespace, not on the text of a key.
- Your CI system already scopes caches per ref automatically. What still needs checking?Three things: what the fallback rules do across those scopes; which jobs can run on the protected ref itself, since anything running there writes into the trusted scope; and whether a shared remote build cache sits behind the per-ref scoping and quietly reunites everything. Automatic scoping is a default, not a guarantee you have verified.
- Why is caching compiled output riskier than caching downloaded archives?A downloaded archive has a recorded expected hash somewhere in the project, so a consumer could in principle re-verify it. A compiled package archive has no such record anywhere - it is a derived artifact whose only definition is "whatever the compiler produced last time" - so a substituted one is indistinguishable from a legitimate one and goes straight into the linked binary.
- The team says scoping will halve the cache hit rate. How do you argue it?Measure it rather than concede it: most of the hit rate usually comes from dependency downloads, which can stay shared, while compile caches are the part that must be per-trust-class. Then price the alternative honestly - the shared namespace means any branch author can influence a release binary, which is a privilege escalation, not a performance setting.
saying these in an interview costs you the question
- Treating the branch name in a key as a security boundary
- Claiming the checksum file protects everything in the cache
- Assuming only maintainers can write a cache entry
- Ignoring prefix-fallback restore behaviour entirely
- Designing the key for hit rate and calling scoping done