How would you design the cache key for a CI job's dependency cache, and what are the two ways a cache key can be wrong?
answer
- hash of everything that determines the content
- lockfile plus toolchain plus platform
- two symmetric failures, one silent
- never hits versus hits when stale
- keep a version prefix you can bump
basics
~20 sDerive the key from a digest of every input that determines the cached content: the lockfile, the toolchain version, the operating system and architecture. Too-specific keys never hit and waste effort; too-loose keys hit while stale and silently corrupt builds.
solid answer
~50 sI treat a cache key as a content-addressing problem: the key must be a function of everything that determines what is inside the entry. For a dependency cache that is normally a digest of the lockfile, plus the runtime or compiler version, plus the OS and CPU architecture, plus a short prefix naming what the entry is and a version number I can bump by hand. Then there are exactly two failure modes and they are symmetrical. A key that is too specific — one that folds in the commit SHA, a timestamp or the run number — never matches, so every run pays full cost and the store fills with dead entries. A key that is too loose — the lockfile digest alone, say — hits when the content no longer matches reality, and the build quietly uses the wrong dependencies. The second failure is far worse, because the first is visible in the run time and the second is visible nowhere.
code
bash · 4 lines# Purpose prefix + manual version + platform + toolchain + lockfile digest.
KEY="deps-v3-$(uname -s)-$(uname -m)-node$(node -v)-$(sha256sum pnpm-lock.yaml | cut -d' ' -f1)"
echo "cache key: $KEY"
# Bumping v3 -> v4 abandons every stale entry in one edit.go deeper
Know that the key decides whether a restore happens, and that it should be built from the lockfile rather than from anything that changes every run, such as the commit SHA.
Be able to enumerate the components — purpose prefix, manual version, platform, toolchain, lockfile digest — and explain the two symmetrical failures and why the silent one matters more.
Show how you detect both failures in practice: hit rate as a tracked metric, scheduled cache-disabled runs to catch inputs escaping the key, and a rotation lever built into the key before you need it.
Own the standard across repositories — a shared key convention, agreement on which data classes may be cached at all, and the budget argument that a wasted cache is money while a wrong cache is an incident.
## A cache key is a content-addressing problem A cache entry is a promise: *if the key matches, the stored bytes are the bytes this job would have produced.* The whole discipline follows from taking that promise literally. The key must be a function of every input that determines the content, and of nothing else. Anything determinant left out of the key produces wrong hits; anything non-determinant folded into the key produces missed hits. So the design question is not "what string do I put here" but "what actually determines what ends up in this directory?" ## What belongs in the key For a dependency cache, the usual answer is four things: 1. **A purpose prefix and a manual version.** Something like `deps-v3-`. The version digit is your escape hatch: when you discover an entry is bad, you bump it and the whole namespace rotates. 2. **A digest of the dependency manifest that pins versions** — `package-lock.json`, `pnpm-lock.yaml`, `poetry.lock`, `go.sum`, `Gemfile.lock`, `Cargo.lock`. Hash the *lock* file, not the loose manifest, because the loose one can say `^1.2` and resolve differently on two days. 3. **The toolchain identity.** The Node, Python, Go, JDK or compiler version. Wheels, native modules and compiled extensions differ between them, and this is the single most commonly forgotten component. 4. **The platform.** OS and CPU architecture, because a cache written on x86 Linux is not valid on arm64 or on a different libc. ```bash # key derived from every input that determines the content KEY="deps-v3-$(uname -s)-$(uname -m)-node$(node -v)-$(sha256sum pnpm-lock.yaml | cut -d' ' -f1)" ``` What does *not* belong: the branch name, the commit SHA, the run number, the date. None of them determine the content, and each of them guarantees a miss. ## Failure one — the key that never hits Symptoms: cache steps report a miss on every run; run times never improve; the cache store grows until eviction starts thrashing and even the legitimately reusable entries disappear. Causes are almost always over-specification — a commit SHA or timestamp in the key — or a digest computed over a file that changes on every run, such as a generated lockfile or a manifest the build itself rewrites. This failure is *loud in the metrics and silent in the logs*: nothing is broken, you are just paying full price. It is worth watching cache hit rate as an actual pipeline metric, because otherwise nobody notices for months. ## Failure two — the key that hits when it should not Symptoms: builds succeed, and ship the wrong thing. A job upgrades its runtime version but the key is only the lockfile digest, so it restores dependencies compiled against the old runtime. A cached build output is restored although a source file that fed it is not in the key, so a change never takes effect and the tests validate the previous revision. A native module cached on one OS image is restored on the next image generation. This is the dangerous one, because there is no signal. The pipeline is green, the diff looks right, and the artifact is wrong. Everything about key design is really about buying insurance against this case, which is why the correct instinct when in doubt is to add a component to the key: a wasted cache is a cost, a wrong cache is an incident. ## Partial and fallback restores widen the second failure Most platforms let you fall back to a related entry when the exact key misses — typically the newest entry whose key starts with a given prefix. This is genuinely useful: on a lockfile change you would rather start from last week's dependency store and download the delta than from nothing. But understand what you have accepted. A fallback restore is, by construction, content that does **not** match your inputs. It is only safe when the step that follows reconciles the restored content against the manifest — an installer that will still fetch what is missing and prune what should not be there. Never use a fallback restore for a cache of *build outputs*, where nothing downstream re-checks the content. ## Immutability and rotation On platforms where an entry is written once per key and never overwritten, a poisoned or stale entry persists until it is evicted — re-running the job does not fix it. That is exactly what the manual version prefix is for: changing the key is the only reliable way to abandon bad content. Design the key with that lever in it from the start, rather than discovering you need it during an incident. ## Verifying the design Two cheap checks. First, change one input at a time — bump the runtime version, touch the lockfile — and confirm the key changes each time; if it does not, you have found a wrong-hit path. Second, run periodically with caching disabled and compare outputs; a difference means an input is escaping the key.
- Which of the two failure modes would you rather have, and why?The key that never hits, every time. It costs money and minutes, and it is measurable — hit rate on a dashboard tells you immediately. The wrong hit costs correctness: a green pipeline that built against the wrong dependencies, with nothing in the logs or the diff to show it. So when a component is borderline, I put it in the key.
- When is it safe to fall back to a related cache entry instead of an exact key match?Only when the step after the restore reconciles the content against a pinned manifest — an installer that fetches what is missing and removes what should not be there. Then the fallback is just a warm start. For a cache of build outputs, where nothing re-checks the content, a fallback restore is a wrong hit by definition.
- You discover a cache entry contains bad content. How do you get rid of it?Change the key — bump the version segment in the prefix — so the whole namespace rotates and nothing restores the bad entry again. On platforms where an entry is immutable once written, re-running the job will not overwrite it, and waiting for eviction is not a plan. Manual purge tools exist but do not scale across branches.
- Why hash the lockfile rather than the dependency manifest?The manifest usually holds ranges — `^1.4`, `~2.0` — that resolve differently on different days, so two runs with an identical manifest digest can legitimately need different dependency trees. The lockfile pins exact versions and their checksums, so its digest is a genuine function of the content the cache will hold.
saying these in an interview costs you the question
- Puts the commit SHA or run number in the cache key
- Keys only on the lockfile and forgets the runtime version
- Thinks a low hit rate is the only cache failure
- Uses fallback restores for cached build outputs
- Expects a re-run to overwrite a bad cache entry