skip to content

In a CI pipeline, what is the difference between a cache and a build artifact, and why must a job still succeed when the cache is empty?

level: juniorimportance: should knowfreq 55%

answer

  1. speed versus required output
  2. eviction is normal, not a failure
  3. cold run must produce the same result
  4. cache the download store, not the install
  5. install step always runs

basics

~20 s

A cache is a best-effort copy of data a job could regenerate itself, kept only to make later runs faster; an artifact is a required output a later job or a release consumes. Every job must still succeed on an empty cache.

solid answer

~50 s

They look alike — both upload a directory and download it in a later run — but the contract is opposite. A **cache** holds regenerable data: downloaded dependencies, a package manager's local store, incremental compiler output. It is a speed optimisation, and the platform is free to evict it at any time for size limits, retention windows or a deleted branch. An **artifact** is a declared output of the job: the compiled binary, the test report, the SBOM. Something downstream depends on it existing, so its absence is a failure. That gives the practical rule: a cache miss must only ever cost time, never change the result. If the job would fail, or would build something different, because the cache was cold, the pipeline has a correctness bug rather than a performance one — and the usual shape of that bug is restoring `node_modules` from cache and skipping the install step that reconciles it with the lockfile.

code

bash · 3 lines
bash
# Cache ~/.npm; the install still runs every time and reconciles with the lockfile.
npm ci
# Cold cache: slow, correct. Warm cache: fast, identical result.

go deeper

for a junior

Be ready to say plainly that a cache only makes a run faster while an artifact is an output something else needs, and that an empty cache should slow a job down, never break it.

for a middle

Explain the mechanics behind the rule: eviction policies, size caps and retention windows mean cold is a normal state, and package managers are designed to read a cached store and still reconcile against the lockfile.

for a senior

Show how you catch the silent version — a scheduled cache-disabled run, and a review habit of asking which step would change its output if the cache vanished mid-pipeline.

for a principal

Own the policy: which data classes are allowed in a cache at all, how hand-offs between stages are standardised on artifacts, and how you keep teams from turning a best-effort store into an undeclared dependency of the release path.

## Two mechanisms that look identical Every CI platform offers two ways to move files out of one run and into another, and they are constantly confused because the mechanics are the same: a directory is archived, uploaded to central storage, and restored later. The difference is not in how they work but in what each one promises. A **cache** is a best-effort copy of data the job could produce again on its own. Typical contents: a dependency download directory (`~/.m2`, `~/.gradle/caches`, the npm cache under `~/.npm`, the pip wheel cache, the Go module cache), a package manager's content-addressed store, or an incremental build directory. Nothing about the pipeline's meaning depends on it. The platform may evict it whenever it likes — most impose a total size cap with least-recently-used eviction, plus a retention window measured in days, and caches tied to a branch usually vanish when the branch is deleted. An **artifact** is a declared output of the job. The compiled jar or binary, the test report the pipeline summary renders, the coverage file, the container image reference that a deploy stage will promote. Something downstream — another job, a release page, a human — consumes it, so if it is missing the pipeline is broken. ## Correctness is the dividing line The useful way to hold the distinction is by what happens on absence: - Cache absent → the job does more work and takes longer. Same result. - Artifact absent → the consuming job cannot run. Failure. That is why a cache must never be load-bearing. Concretely, a cold cache is not an exceptional condition to be handled; it is the normal state on a brand new branch, on the first run after an eviction, after a key change, and on any fresh ephemeral runner whose local disk starts empty. Caches also race: two jobs starting at the same time both miss, both compute, and both try to store — so a design that assumes "the previous job populated it" is already wrong. ## The classic mistake The most common cache defect in real pipelines is using a cache as an installer: ```bash # WRONG: the cache is the only thing that puts dependencies on disk if [ ! -d node_modules ]; then npm ci; fi ``` This has two failure modes at once. Cold, the directory is missing and, in the variants that omit the fallback entirely, the build simply fails. Warm, the job runs whatever an older commit's install produced, so dependencies drift away from the lockfile silently — nothing in the diff shows it, and the version that ships is not the version the lockfile describes. The correct shape always runs the real install and lets the cache make it fast: ```bash # RIGHT: install always runs; the cache only saves network and unpack time npm ci # succeeds cold (slow) and warm (fast), with the same result ``` Most package managers are built for exactly this: they read a lockfile, check what is already present in their local store, and fetch only the difference. Cache the *store*, not the resolved tree, whenever the tool gives you that choice. ## What goes where Cache: dependency downloads and package-manager stores, compiled third-party sources, prebuilt fixtures, incremental compilation directories (only when the key genuinely covers every input). Artifact: the build output that will be deployed or published, test and coverage reports, generated SBOMs, logs worth keeping, anything a later stage must consume. When the same bytes need to reach a deploy stage, that is an artifact hand-off, not a cache lookup — the deploy must not depend on a best-effort store. ## Proving cold-cache correctness Because the failure is silent, teams verify it on a schedule rather than by inspection: run the pipeline periodically with caching disabled, on a timer rather than on every push. If the cold run fails, or produces output that differs from the warm run, the cache was hiding a defect that would eventually ship — usually a dependency that no longer resolves, a tool that is no longer installed, or a generated file nobody regenerates any more. A last note worth saying out loud in an interview: a cache is not a security boundary and not durable storage. Anything you would be unhappy to lose, or unhappy for another job to read, does not belong in it.

  • How would you prove that your pipeline is still correct with a completely cold cache?
    Run the whole pipeline on a schedule with caching turned off — nightly or weekly rather than on every push — and treat any failure or output difference as a defect. It catches dependencies that no longer resolve, tools the cache was silently carrying, and generated files nobody regenerates. It costs one slow run and removes the class of bug that only appears on a new branch.
  • Two jobs start at the same moment and both miss the same cache key. What happens?
    Both do the full work, and both attempt to store an entry under that key. There is no coordination, so you pay the cost twice and one of the uploads is redundant. That is expected behaviour, not a bug — it is another reason a job can never assume a previous run populated the cache for it.
  • Where does a container image fit — cache or artifact?
    The built image is an artifact: something downstream deploys it, so it must exist. The builder's intermediate layer cache is a cache — its absence only costs build time. Keeping the two separate is what lets you promote the exact image you tested rather than rebuilding one that might differ.

saying these in an interview costs you the question

  • Says a cache miss should fail the job
  • Restores node_modules and skips the install step
  • Treats the cache as durable storage between stages
  • Passes a build output to the deploy stage through the cache
  • Assumes a fresh branch inherits a warm cache automatically

context