skip to content

A team enables shared remote caching in their monorepo build tool. Soon after, engineers start seeing CI report a task as passing via a cache hit, even though the same code fails when run fresh locally. What's the likely root cause of this class of bug, and how do you prevent it going forward?

level: seniorimportance: should knowfreq 45%

answer

  1. cache key != actual inputs when hermeticity leaks
  2. hidden inputs: env vars, clock, network, randomness
  3. sandboxing blocks undeclared reads structurally
  4. force cache bypass to diagnose
  5. convention-based prevention degrades with team size

basics

~20 s

Some hidden input the task depends on (like an environment variable, network response, or timestamp) isn't included in the cache key. So a run that should be treated as different from a previous one gets matched to it anyway, and the old, wrong result gets replayed instead of rerunning.

solid answer

~50 s

This is a hermeticity violation: the cache key is computed from a declared input set (source hashes, command, declared deps), but the task's actual behavior depends on something outside that set — an environment variable, system clock, network call, filesystem state outside the sandbox, or non-deterministic output like a random seed or map iteration order. Because the undeclared input differs between the cached run and the current one but isn't part of the key, the tool concludes 'inputs match' and replays the old (possibly now-wrong) result instead of re-executing. Prevention means: enumerating and declaring every real input for hermetic-sensitive tasks; running affected tasks inside a sandbox that structurally blocks undeclared reads/network access (Bazel's approach) rather than relying on developer discipline; treating flaky or environment-sensitive tests as a signal to fix the test rather than to loosen caching; and, when a stale hit is suspected, being able to force a cache bypass to confirm before trusting the tool again.

go deeper

for a junior

Should recognize that a 'passing' cached result can sometimes be wrong and that this is different from a normal test failure.

for a middle

Should identify at least one concrete hidden-input example (env var, timestamp, network) as a plausible cause.

for a senior

Should describe the diagnostic process (force cache bypass, audit declared vs actual inputs) and name sandboxing as the structural prevention, distinct from convention-based discipline.

for a principal

Should weigh the enforcement-cost-vs-risk trade-off across the whole org, decide when investing in sandboxed execution is worth it, and design an incident response/postmortem process for when a stale cache hit reaches production undetected.

## The bug class This bug class — "CI says green via cache, reality says red" — is one of the most dangerous failure modes in a shared-caching build system precisely because it looks identical to a healthy fast build from the outside. Understanding it requires being precise about what a **cache key** actually represents versus what a task's **real inputs** are. ## What the key represents, and what it misses The orchestrator computes a cache key from a declared input set: it hashes the source files the task lists as inputs, the transitive hashes of its declared dependencies, the exact command and flags, and sometimes the toolchain version. That key is a proxy for "everything that could affect this task's output" — but it's only accurate if the declared set genuinely matches the actual set of things the task's behavior depends on. A task can secretly depend on something outside that declared set: - an environment variable read by the test framework to toggle a feature flag; - the system clock via `Date.now()` embedded in an assertion; - a live network call to a staging service whose response varies; - unordered map/set iteration in a language where that's not guaranteed; - a random seed not fixed for the test run. In each of these, the cache key stays identical across runs where that hidden factor actually differs. The tool, seeing an unchanged key, concludes it already has the answer and replays the old result instead of re-executing, and if the hidden factor's new value would have produced a failure, that failure is masked. It doesn't just slow down bug discovery — it removes any signal that a bug exists at all, until someone hits the actual broken behavior in production and starts asking why CI didn't catch it, at which point "well it did pass" is technically true and deeply unhelpful. ## Why the risk exists at all This exists as a risk precisely because caching's whole value proposition is trusting a hash instead of re-executing — the faster and more aggressive the caching, the more damage an undetected hermeticity gap does, since a single bad cache entry gets replayed for every subsequent matching key until someone invalidates it or the input set finally changes for an unrelated reason. ## The trade-off: enforcement cost versus risk tolerance The trade-off engineering teams navigate here is **enforcement cost versus risk tolerance**. The strongest prevention is structural: run tasks inside a sandbox that only exposes declared inputs and blocks everything else by construction — - no network access, - no filesystem access outside the declared input set, - environment variables scrubbed to an explicit allow-list. Bazel does this by default for build actions and can be configured to sandbox test execution too, which is precisely why organizations with strict correctness requirements lean on it. The cost is real: sandboxing breaks tests or build steps that legitimately need network access (e.g., integration tests against a real staging API) unless those are explicitly carved out as non-cacheable or given a controlled, mocked substitute — which is more upfront engineering work than just letting the test run unconstrained. ## Convention instead of enforcement Lighter tools without built-in sandboxing (Nx, Turborepo) instead rely on convention and configuration: task authors declare `inputs`/`outputs` explicitly, and the org has to establish a practice of treating any environment- or network-dependent test as something to either fix (mock the dependency, fix the seed) or explicitly mark as non-cacheable / always-run. This is cheaper to set up but depends on ongoing engineering discipline rather than a structural guarantee, so it degrades as team size grows and fewer people remember the convention. ## Diagnosing an actual incident Diagnosing an actual incident typically starts by: 1. **forcing a cache bypass** (most tools offer a `--no-cache` or equivalent flag) to confirm the task genuinely behaves differently fresh versus cached; 2. then **auditing the task's declared inputs** against what it actually reads/calls at runtime — commonly surfaced by grepping the task's source for `process.env`, `Date.now()`, `Math.random()`, or outbound HTTP calls, and checking whether each is either fixed/mocked or explicitly declared as a cache-relevant input. ## Where it shows up A concrete, well-known real-world pattern of this exact bug class: teams using Bazel without sandboxing enabled for a given target (a common misconfiguration when a team disables sandboxing to "fix" a build that fails inside the sandbox, rather than fixing the underlying hermeticity violation) have reported exactly this symptom — a target that reads local user-specific config outside its sandboxed inputs, producing cache hits that mask a bug only reproducible with a clean environment, discovered days later when a new hire's fresh machine failed a "passing" target. The fix in that pattern is almost always to restore sandboxing and explicitly declare the missing input, not to disable caching altogether.

  • How would you retrofit sandboxed execution onto a task that currently makes a legitimate network call, without losing cache correctness?
    Either mark the task as explicitly non-cacheable/always-run so it never gets a false hit, or replace the live network dependency with a recorded/mocked response that's itself a declared, versioned input file — turning a non-deterministic external call into a deterministic, hashable local fixture that caching can safely key on.
  • Why is 'just disable sandboxing to unblock the build' a dangerous quick fix?
    Disabling sandboxing removes the structural guarantee that catches undeclared inputs, so the build starts working again for the wrong reason — because it's now silently reading something outside its declared inputs, which is exactly the hermeticity gap that produces stale, incorrect cache hits later. It trades a loud, fixable build failure now for a quiet, hard-to-diagnose correctness bug later.
  • What's a lightweight way to catch this class of bug before it reaches a shared remote cache used by the whole team?
    Run the same task twice in a row locally with slightly different environment conditions (different env var values, different system time, or airplane mode for network-dependent tasks) and confirm the output is identical; a difference despite an unchanged declared cache key is a direct signal of a hermeticity gap, catchable in a pre-merge check before it ever pollutes the shared cache.

Like a forger who photocopies an old signed contract and slaps today's date on the copy — anyone checking only the signature (the declared 'input') sees a match and approves it, never noticing the actual terms underneath silently changed.

saying these in an interview costs you the question

  • Blames the caching tool itself as buggy rather than looking for an undeclared input in the task
  • Suggests disabling caching entirely as the fix rather than fixing hermeticity
  • Can't name concrete examples of hidden inputs (env vars, clock, network, randomness)
  • Doesn't mention sandboxing or explicit input declaration as a structural prevention
  • Assumes this can't happen with a mature tool like Bazel without qualifying that sandboxing has to actually be enabled and configured correctly

context