Walk through what actually goes into computing a task's cache key in a monorepo build system, and why leaving something out is dangerous.
answer
- input closure
- Merkle-DAG of hashes
- sandboxing enforces exhaustiveness
- false hit worse than miss
- env vars & config count as inputs
basics
~20 sThe cache key is a fingerprint made by hashing everything that could change the task's output — its source files, the versions of things it depends on, and its settings. If you forget to include something that matters, the cache can hand back a wrong, outdated answer without anyone noticing.
solid answer
~40 sA task's cache key is computed by hashing the closure of everything that can influence its output: the content of its own source files, the resolved versions/hashes of its declared dependencies (including transitive ones, often via their own output hashes), relevant environment variables and CLI flags, the toolchain/compiler version, and sometimes the OS/platform. These are combined into a single key. Correctness depends entirely on completeness of this input set — anything the task reads that isn't declared becomes an 'unhashed input,' and changes to it will not invalidate the cache, producing silently stale results. This is why build systems like Bazel enforce sandboxing: it's not just an optimization, it's how they guarantee the declared input list is actually exhaustive.
go deeper
Can name 'source files' as an input; doesn't need to reason about dependency propagation or sandboxing.
Should list several concrete input categories (source, deps, env, config, tool version) and explain hashing by content not timestamp.
Should explain the Merkle-DAG propagation through dependencies and why a false hit is worse than a miss, plus a concrete debugging approach.
Should connect input-completeness enforcement (sandboxing/hermeticity) to organizational trust in a shared remote cache and reproducibility guarantees.
## What a cache key is A **cache key** is the fingerprint a build system computes for a task before deciding whether it can skip running that task. Conceptually, a task is a function: given a set of inputs, it deterministically produces a set of outputs. To make caching sound, the key must summarize every input that can affect the output — nothing more, and critically, nothing less. Getting this input set right is the single hardest and most consequential part of building a correct incremental build/caching system. ## What goes into the input set Concretely, the input set usually spans several categories. 1. **First, the task's own source files** — for a compile task, that's every source file fed to the compiler, hashed by content (typically `SHA-1` or `SHA-256`), not by path or timestamp. 2. **Second, its dependencies**: for a package that depends on another package in the monorepo, the key must fold in either the dependency's source hashes or, more efficiently, the dependency's own already-computed output hash, so that a change deep in a transitive dependency propagates upward without needing to re-hash the entire subtree every time. This is why the dependency graph — the same graph used to determine 'what's affected' by a change — is exactly the structure caching rides on top of: each node's key incorporates its dependencies' keys, forming a **Merkle-DAG** where a single content change ripples deterministically to every downstream key. 3. **Third, the environment**: compiler/tool version, relevant environment variables (e.g. feature flags passed via env), command-line arguments, and configuration file contents (`tsconfig`, `eslint` config, a `Dockerfile`). 4. **Fourth, sometimes platform specifics** — OS, architecture — when output is genuinely platform-dependent, though many build systems intentionally exclude these to make caches portable across machines, accepting the risk that a task claiming platform-independence but secretly isn't will misbehave. ## Combining the pieces into one key All of these pieces get combined into one composite hash — often by hashing a canonical, sorted list of 'input path to content hash' pairs plus the tool/config fingerprints, then hashing that whole structure again into one final key. That final key is what gets looked up in the cache store (local disk, or a remote cache like a shared bucket, `Nx Cloud`, `Bazel Remote Cache`, or Turborepo's remote cache). ## Why it exists The reason this exists is straightforward: re-running expensive, deterministic work is wasted time, and at monorepo scale, the naive alternative (rebuild everything, always) becomes the dominant cost of every CI run and every developer's local iteration loop. Content-addressed keys let a build system make a very strong claim — 'this exact combination of inputs was already computed, here is the byte-identical result' — instead of a weaker heuristic like 'this file's timestamp looks unchanged.' ## The trade-off between too coarse and too narrow The core trade-off is between completeness of the input set and performance/complexity of computing it. | Key granularity | Consequence | |---|---| | A key that's too coarse (e.g., 'rebuild if anything in this package changed' without file-level granularity) | forces unnecessary cache misses, wasting the speed benefit | | A key that's too narrow — missing an actual input | produces false cache hits, which are far worse than false misses | A miss just costs time, but a false hit silently serves incorrect output as if it were correct, and nothing about the build's exit code or logs signals a problem. This is why serious build systems (`Bazel` foremost among them) pair cache-key computation with sandboxing/hermeticity enforcement: the task is executed in an environment where it physically cannot read anything outside its declared inputs, so if the build succeeds at all, the declared input list is provably exhaustive by construction, rather than by developer discipline alone. ## Failure modes in production Failure modes in production cluster around exactly this gap. - A team migrates a **codegen step** to read an extra JSON config file but forgets to declare it as an input; the cache keeps serving pre-migration output for weeks until someone notices a production behavior mismatch that doesn't match the deployed source. - A **CI environment variable** used for feature-flagging leaks into a task's behavior without being part of the key, so builds for two different flag states resolve to the same cache entry and one flag's users silently get the other's code. - A **non-deterministic step** — embedding a build timestamp, a random UUID, or unsorted map iteration order in its own output — technically has a stable key but an unstable output, which either defeats caching outright (every run looks unique) or, worse, means two logically-identical builds produce byte-different artifacts that fail an expected reproducibility check downstream. ## A real-world instance A concrete real-world instance: Bazel's remote caching model is built explicitly around this contract — every action declares its full input set, sandboxing enforces it can't cheat, and the resulting cache is trusted enough that large organizations share a single remote cache across thousands of engineers and CI machines, turning what would be a multi-hour full rebuild into a mostly cache-hit-driven build that finishes in minutes.
- Why do build systems fold a dependency's output hash into the consumer's key instead of re-hashing all of that dependency's source files every time?Re-hashing an entire transitive closure of source files for every downstream package would be expensive and redundant, since the dependency itself already computed and cached that summary. Using the dependency's own content-addressed output hash lets the key propagate cheaply and correctly — if the dependency's inputs didn't change, its output hash doesn't change, and neither does anything downstream.
- What's a practical way teams catch a missing/under-declared input before it causes a production incident?Running periodic 'clean build vs cached build' diffing — building once from a fully clean cache and once from a warm cache with identical source, then comparing artifacts byte-for-byte — surfaces silent divergence. Some teams also enforce sandboxing/hermetic execution so missing declarations fail loudly (file not found) instead of silently succeeding by reading an undeclared file.
- Should environment variables always be part of the cache key?Only the ones that actually influence the task's behavior — including every environment variable indiscriminately would cause needless cache misses whenever an unrelated env var changes. Build systems typically require explicit allowlisting of which env vars are 'sensed' inputs for a given task.
Like a recipe's exact fingerprint — if the cache key only records the ingredients but forgets the oven temperature, two identical-looking cakes baked at different temperatures get treated as the same dish.
saying these in an interview costs you the question
- Thinks cache keys are based on file paths or timestamps rather than content
- Doesn't mention dependency versions/hashes as part of the key
- No awareness that a missing input causes a silent wrong-answer, not a visible error
- Assumes sandboxing is just about security, not about proving key completeness