What's the practical difference between a local build cache and a remote/distributed build cache in a monorepo, and what does a team give up by adopting the remote one?
answer
- shared store vs per-machine
- CI always starts cold locally
- network cost vs recompute cost
- write-auth = trust boundary
- cache poisoning as supply-chain risk
basics
~20 sA local cache only remembers work done on your own machine; a remote cache is shared, so if a teammate or CI already built something, you can download that result instead of building it yourself. The trade-off is speed and sharing versus needing a network connection and trusting what others uploaded.
solid answer
~50 sA local cache stores task outputs on the same machine that ran them, keyed by content hash, so it only helps repeat builds on that one machine. A remote/distributed cache stores those same outputs in a shared service (S3, Bazel Remote Cache, Nx Cloud, Turborepo Remote Cache) that every developer's machine and every CI runner can read from and write to, so a cache miss on your laptop can still resolve as a hit if any other machine already computed that exact key — this is huge for CI, where every job starts from a cold, empty local cache. The costs are network latency (uploading/downloading artifacts can outweigh the savings for cheap tasks), infrastructure to run and secure the cache service, and a trust boundary: a remote cache can serve a bad or malicious artifact to everyone unless writes are authenticated and integrity-checked.
go deeper
Should grasp that 'remote' means shared across machines, via the analogy, without needing to reason about trust boundaries.
Should explain why CI benefits disproportionately from remote caching versus local, and name at least one real product.
Should reason about when remote caching is a net loss (small tasks, large artifacts) and design a basic read/write trust split.
Should treat cache poisoning as a supply-chain security concern and connect cache-hit-rate monitoring to build system health/ROI decisions at org scale.
## The same problem at a different scope A local cache and a remote cache solve the same problem — avoid re-running a task whose inputs haven't changed — but they differ in scope: who can benefit from a previously computed result. | Cache | Where the store lives | |---|---| | Local | on the single machine that executed the task, typically as a directory on disk keyed by content hash | | Remote (or distributed) | a shared service that many machines read from and write to over the network | ## What a local cache covers A local cache lives on the single machine that executed the task. It helps exactly one scenario well: a developer or CI job re-running the same command multiple times on the same machine without the relevant inputs changing, e.g. running the test suite twice in a row, or switching git branches and back. Its big limitation is that every fresh environment starts with an empty cache and gets zero benefit, no matter how many times that exact task has already been computed elsewhere by someone else on the team: - a new CI runner spun up per job; - a new hire's laptop; - a Docker container rebuilt from scratch. ## How a remote cache closes that gap A remote (or distributed) cache addresses exactly that gap by centralizing the cache store into a shared service that many machines read from and write to over the network — commonly backed by object storage fronted by a caching service (`Bazel Remote Cache` server, `Nx Cloud`, `Turborepo Remote Cache`, `BuildBuddy`). The workflow becomes: 1. Before running a task, compute its content-addressed key as usual, but instead of only checking a local directory, query the remote service. 2. If any machine anywhere — a teammate's laptop an hour ago, or last night's CI run — already computed and uploaded a result for that exact key, the current machine downloads it instead of executing the task. 3. After a real cache miss and execution, the new result gets uploaded back so the next machine (often the very next CI job in the pipeline) benefits. This matters enormously in CI, which is the case local caching can't help at all: every job typically starts cold. In a monorepo running the same tasks repeatedly across many PRs that mostly touch the same shared code, a warm remote cache can turn a CI pipeline that would take 20–40 minutes into one that finishes in a couple of minutes, because most of the graph resolves to cache hits computed by a previous run. ## The trade-offs The trade-offs are real and multi-dimensional. 1. **First, network cost**: downloading a cached artifact is only a win if it's cheaper than re-running the task; for a task that takes 200ms to execute but produces a 50MB artifact, fetching over a slow network connection can be a net loss, so mature systems set thresholds or heuristics about which tasks are worth caching remotely at all. 2. **Second, operational cost**: running a remote cache service means standing up and maintaining infrastructure (or paying for a hosted one), managing storage growth and eviction policy, and monitoring cache hit rates to know if it's actually paying for itself. 3. **Third, and most consequential, trust**: a local cache can only be poisoned by bugs in your own build, but a shared remote cache is a trust boundary — if writes aren't authenticated, or if a misconfigured/malicious CI job uploads an incorrect artifact under a legitimate-looking key, every subsequent machine that queries that key downloads the bad result, silently, with no visible failure. This is why production remote-cache setups require write authentication (only trusted CI, not arbitrary PR branches from forks, can write), and often checksum/integrity verification on download. ## Failure modes Failure modes show up as a specific, recognizable class of incident: - **'It works on my machine but CI is broken'** (or vice versa) because one side has a stale/incorrect remote cache entry the other doesn't share. - **A security incident** where an attacker with write access to the cache (e.g. via a compromised CI credential) poisons a widely-shared key and gets malicious code distributed to every subsequent build that reads it, without touching source control at all — a known supply-chain attack vector, which is why untrusted PR-triggered CI jobs are commonly restricted to read-only cache access. - **A more mundane failure** — remote cache staleness after a genuine bug fix in the build tool itself changes what a task's inputs should include; old cache entries computed under the previous, buggy key scheme get served as if valid until the cache is explicitly bumped/invalidated (often via a cache-version salt in the key). ## Where it shows up Concretely: Google's internal build system and its open-source descendant `Bazel` popularized remote caching (and remote execution) at massive scale, sharing a single cache across the whole engineering org; in the JS/TS ecosystem, Vercel's `Turborepo Remote Cache` and `Nx Cloud` bring the same model to typical web monorepos, where teams report CI times dropping dramatically once the cache hit rate on the main branch stabilizes at a high level.
- Why might a team deliberately choose NOT to cache a particular task remotely, even though it's technically cacheable?If a task is cheap to execute (sub-second) but produces a large artifact, the network round-trip to fetch it can be slower than just re-running it, so caching it remotely is a net loss. Teams also sometimes skip remote caching for tasks with side effects or weak hermeticity guarantees, where they don't fully trust that the declared input set is complete.
- How would you design write access to a remote cache to prevent a compromised PR-triggered CI job from poisoning it?Restrict write permissions to a small set of trusted pipelines — typically only CI runs on the protected main/trunk branch after merge — while giving untrusted PR builds read-only access to the cache. This means a malicious or buggy fork PR can benefit from existing cache entries but can never inject a poisoned one that other builds would later trust.
- If CI cache hit rate suddenly drops from 90% to near 0%, what's a likely cause?A common cause is a change to something globally included in every task's cache key — a toolchain/compiler version bump, a shared config file edit, or a lockfile-wide dependency update — which invalidates a huge swath of the graph at once. Another cause is a cache-version salt being bumped deliberately after a bug fix in the build tool's own key computation.
Local cache is like your own fridge leftovers; remote cache is like a shared work-kitchen freezer everyone can pull a pre-made meal from — great when someone already cooked it, risky if someone put something spoiled in the freezer and labeled it wrong.
saying these in an interview costs you the question
- Thinks local and remote caches are functionally identical, just 'in the cloud'
- Doesn't mention that CI runners start with an empty local cache every time
- No awareness that a shared cache is a write-trust/security boundary
- Assumes remote caching is always strictly faster with no network cost consideration