skip to content

In a monorepo build system, what does it mean for a build task to get a 'cache hit', and why does that make rebuilds faster?

level: juniorimportance: must knowfreq 70%

answer

  1. pure function analogy
  2. content-addressed hash
  3. cache key = inputs hash
  4. hit vs miss
  5. stale output risk

basics

~20 s

A cache hit means the build system already ran this exact task before with the exact same inputs, so it reuses the saved output instead of redoing the work — like reheating leftovers instead of cooking again.

solid answer

~30 s

Build systems like Bazel, Nx, or Turborepo track every task (compile, test, lint) as a function of its inputs — source files, dependencies, config, environment. Before running a task, they hash those inputs into a cache key and check if an output for that exact key already exists locally or remotely. If yes, that's a cache hit: the system copies the stored output back instead of re-executing the task, often reducing a step from minutes to milliseconds. This lets a monorepo with thousands of packages rebuild only what actually changed instead of everything on every CI run or local build.

go deeper

for a junior

Should be able to explain in plain language that unchanged work is skipped and reused, using an analogy; doesn't need to know hashing details.

for a middle

Should know that keys are computed from file content hashes, not timestamps, and name a couple of real inputs (deps, source files).

for a senior

Should be able to explain how an under-specified cache key produces stale results and how they'd debug/fix it in a real pipeline.

for a principal

Should connect caching correctness to trust boundaries — e.g. why remote/shared caches need integrity guarantees beyond what a single dev's local cache needs.

## A build task as a pure function Every build task—compiling a package, running a linter, executing a test suite—can be thought of as a **pure function**: given the same inputs, it always produces the same output. A build system that wants to avoid redundant work exploits this by first computing a **cache key** for a task before running it, then checking whether a cached result already exists for that exact key. - **Hit** — if a matching result is found, that's a cache hit: the system fetches the previously stored output (compiled artifacts, test logs, exit code) and hands it back to whoever asked, without ever spawning the compiler or test runner. - **Miss** — if no match exists, that's a cache miss, and the task actually executes; its output is then stored under that key so future requests for the same inputs can hit the cache. ## What goes into the key The inputs that go into the key typically include: - the content of every source file the task reads; - the versions and content hashes of its dependencies; - the command-line flags and environment variables that affect its behavior; - the toolchain version itself (compiler binary, linter version). Because the key is derived from file content (usually via a hash like `SHA-256`) rather than file modification timestamps, this technique is called **content-addressed caching**—two files with identical bytes produce the same hash regardless of when they were last touched, which makes the cache immune to things like a `git clone` resetting timestamps or a CI checkout re-writing every file's mtime. ## Why it matters at monorepo scale Why this matters becomes obvious at monorepo scale. A repository with hundreds of packages might have a full build/test/lint pipeline that takes 30–60 minutes from a cold start. Most day-to-day changes touch only a handful of those packages. Without caching, every CI run and even every local build command re-executes every task for every package, because the build tool has no way to know that most of the graph is unchanged. With content-addressed caching, only the tasks whose actual inputs changed miss the cache; everything downstream that transitively depends on unchanged inputs also gets a cache hit, because its own inputs (which include the hashes of its upstream dependencies' outputs) are unchanged too. The practical effect is that a single-line change to one leaf package can produce a build that finishes in seconds instead of tens of minutes, because only that package and its direct consumers actually re-run. ## The trade-off The trade-off is bookkeeping overhead and correctness risk. Computing hashes for every input of every task costs CPU time and I/O, though this is normally far cheaper than re-running the task itself. The bigger risk is **under-specification**: if the cache key omits an input that actually affects the output — a stray environment variable, an implicit read of a config file not declared as a dependency, a non-deterministic code generator — the build system will confidently return a cache hit for a task whose real behavior has changed. This produces a 'stale' or 'wrong' result that looks identical to a correct one until someone notices the shipped artifact doesn't match the source. Teams debug this by identifying the untracked input and adding it to the task's declared inputs, or by making the task's declared inputs stricter (sandboxing filesystem/network access) so nothing can sneak in unnoticed. ## Failure modes in production In production this shows up in a few recognizable ways. 1. **First**, 'why didn't my change take effect' bugs, where a developer changes a file that (unknown to the build graph) is read by a task but not declared as an input, so the task keeps serving a stale cached build. 2. **Second**, flaky CI where a task is legitimately non-deterministic (e.g. embeds a timestamp or random UUID in its output) and so never gets a stable cache key, defeating caching for that node and everything downstream of it. 3. **Third**, security and trust concerns for shared/remote caches, where a malicious or buggy CI job could poison the shared cache with an incorrect artifact that then gets served to every other developer and pipeline that computes the same key. ## Where it shows up A concrete example: `Bazel` popularized this model for large monorepos at Google, using strict sandboxing to guarantee hermeticity so cache keys can be trusted; `Nx` and `Turborepo` bring a lighter-weight version of the same idea to JS/TS monorepos, hashing each project's source files plus its dependency graph to decide whether to replay a cached build versus actually invoking `webpack`, `tsc`, or `jest`. In all three, the fundamental unit of reasoning is the same: a task node in the dependency graph, a hash of everything that could affect its output, and a lookup against a store of previous results keyed by that hash.

  • What happens if two different tasks happen to produce the same cache key by coincidence?
    In practice this is astronomically unlikely with a strong hash function like SHA-256, since the key space is enormous and hash collisions would require a targeted attack rather than accident. Build systems generally trust the hash's collision resistance rather than defending against it explicitly. The much more common and real risk is not hash collision but an under-specified key — a task with a missing input feeding a wrong hit.
  • Does a cache hit skip the task's side effects too, like writing to a database during a test?
    Yes, and that is precisely why cacheable tasks must be side-effect-free or hermetic — a cache hit only replays previously captured outputs and logs, it never re-executes the task's code. If a 'test' task mutates external state as a side effect, caching it silently skips that mutation on a hit, which is a correctness bug in the task definition, not in the cache.
  • How does a build tool decide a task is a 'leaf' with no cacheable output, like a deploy step?
    Deploy or publish steps are usually marked non-cacheable in the task graph configuration precisely because they have external side effects (pushing to a registry) that shouldn't be skipped just because the same code was built before. Build tools like Nx and Turborepo let you flag which tasks participate in caching per task type.

Like a photocopier that first checks if it already has an identical page on file — if the page's contents match exactly, it hands you the existing copy instead of running the original through the machine again.

saying these in an interview costs you the question

  • Believes caching just means 'don't rebuild if the file's timestamp is old'
  • Can't explain what could cause a cache hit to serve stale/wrong output
  • Thinks caching only applies to the exact file that changed, not its downstream consumers
  • No mention of inputs beyond source code (env vars, deps, toolchain version)

context