skip to content

What does it mean for a build task to be 'hermetic', and what specifically breaks when a task isn't?

level: seniorimportance: must knowfreq 65%

answer

  1. pure function w/ no hidden inputs
  2. sandboxing enforces exhaustiveness
  3. hidden input -> silent stale hit
  4. nondeterministic output -> perpetual miss or reproducibility break
  5. reproducible builds tooling

basics

~20 s

A hermetic task only uses the inputs it's told about — no hidden reads from the network, system clock, or random files — so running it twice with the same inputs always gives the same result. If a task secretly depends on something outside that list, caching it becomes unsafe.

solid answer

~60 s

Hermeticity means a task's output is a pure function of its declared inputs and nothing else — no ambient filesystem access outside declared paths, no network calls, no reliance on system time, machine hostname, locale, or unordered iteration that varies run to run. This matters because caching's entire correctness argument rests on 'same inputs -> same key -> safe to reuse output,' and that argument only holds if the declared input set is actually exhaustive. Non-hermetic tasks break this in two ways: (1) they read something real that isn't in the key, so a change to that hidden input doesn't invalidate the cache, serving stale results; or (2) they produce genuinely different output on each run even with identical inputs (timestamps, random IDs, unordered map serialization), which defeats caching or produces artifacts that fail downstream reproducibility checks. Build systems like Bazel enforce hermeticity via sandboxing — running the task in a restricted environment where non-declared reads simply fail — turning a silent correctness bug into a loud build failure.

go deeper

for a junior

Should understand hermetic loosely as 'always gives the same result for the same input' without needing sandboxing details.

for a middle

Should be able to name a concrete non-hermetic example (e.g. embedding a timestamp) and say why it's a problem for caching.

for a senior

Should clearly separate the two failure modes (hidden input vs non-deterministic output) and connect sandboxing to enforcement, not just convention.

for a principal

Should weigh the cost of full sandboxing against risk-based adoption and connect hermeticity to reproducible-build/supply-chain concerns beyond caching alone.

## Why hermeticity is what makes caching trustworthy Hermeticity is the property that makes content-addressed caching trustworthy in the first place. A cache system's entire safety argument is: 'I computed a key from everything that affects this task's output, so if I see that key again, I can hand back the old result instead of recomputing.' That argument is only true if the input set really is everything that affects the output. A **hermetic task** is one where that's actually the case — it reads only its declared inputs (specific source files, specific declared dependencies) and produces output that's a pure, deterministic function of those inputs, with no side channels: - no arbitrary filesystem reads outside its sandbox; - no network requests; - no dependence on wall-clock time, machine hostname, environment variables that weren't explicitly declared, or iteration order of an unordered data structure like a hash map. ## Two distinct failure classes Two distinct things can go wrong when a task isn't hermetic, and it's worth separating them because they produce different symptoms. ## The first: a hidden real input The first is a hidden real input: the task's output genuinely depends on something — say, a config file at a fixed absolute path outside the declared source tree, or a version of a system library installed globally — that isn't part of the cache key. Here, the task is deterministic in isolation but the build system's model of its dependencies is incomplete. The failure mode is **silent staleness**: someone changes that hidden input, the task's actual output should change, but the cache key doesn't change, so the cache confidently serves the old, now-wrong result. Nothing errors; the build reports success; the artifact is simply incorrect relative to current source. This is the more insidious of the two failure classes because it's invisible until someone independently notices behavior doesn't match code. ## The second: genuine non-determinism The second failure is genuine non-determinism: the task, given byte-identical inputs, produces different byte-level output on different runs — because it embeds the current timestamp, generates a random UUID, serializes a hash map in whatever order the runtime happens to iterate it, or depends on floating-point results that vary slightly across CPU architectures. Here the input side is fine (the key is stable), but the output side is unstable. This either defeats caching outright — every run 'misses' because the tool compares output hashes and never sees a repeat, wasting all the caching infrastructure's overhead for zero benefit — or, in caching systems that trust the declared key blindly, it produces two logically-equivalent artifacts that are byte-different, which breaks anything downstream that expects reproducible builds: - binary diffing for release verification; - deterministic image layers; - supply-chain attestation systems that compare a rebuilt artifact's hash against a published one to prove it wasn't tampered with. ## Why it is enforced rather than hoped for Why hermeticity is treated as a first-class engineering concern rather than an afterthought is that developer discipline alone can't reliably guarantee it — it's very easy to accidentally add an ambient read (a codegen script that reads the current time for a 'build info' banner, a linter that reads a config from a home directory instead of the repo) without anyone noticing during code review. This is why Bazel's approach is to enforce hermeticity structurally: each action executes inside a sandbox (a restricted filesystem view and, optionally, network namespace) that physically cannot see anything the build definition didn't explicitly declare as an input. If a task tries to read an undeclared file, the read fails immediately and loudly — the build breaks — instead of silently succeeding and producing an unhashed dependency. This converts hermeticity from 'hopefully true if everyone is careful' into 'provably true because the sandbox makes violations impossible to miss.' ## What the rigor costs The cost of this rigor is real: sandboxing has overhead (spinning up an isolated environment per action), and it requires build authors to be exhaustive and explicit about every input a task needs, which is more upfront work than letting a script read whatever it wants. Teams without Bazel-grade sandboxing (most JS/TS monorepo tools like `Nx` or `Turborepo`) generally rely on convention and best-effort input declaration rather than enforced isolation, trading some correctness guarantee for much lower adoption friction — appropriate for many teams, but it means the two failure modes above are always a latent risk rather than a structurally prevented one. ## Where it shows up A concrete real-world example: reproducible-builds efforts (used by projects like Debian, and by supply-chain security tools) treat non-determinism as a bug class worth dedicated tooling to hunt down — tools like `diffoscope` exist specifically to diff two builds of 'the same' source and surface exactly which non-hermetic input (a timestamp, a build path baked into debug symbols, an unordered archive member list) caused a byte-level difference.

  • How would you notice, in practice, that a task is producing non-deterministic output rather than just being slow?
    A direct test is running the exact same task twice with identical, unchanged inputs and diffing the two output artifacts byte-for-byte — if they differ despite identical inputs, that's non-determinism, not staleness. Some CI pipelines run this check periodically as a 'determinism gate' specifically to catch it before it silently defeats caching.
  • Why might a team accept some non-hermetic tasks rather than fixing all of them immediately?
    Fully sandboxing every task is significant upfront engineering investment, and for a task that's rarely re-run or cheap to execute, the ROI of enforcing strict hermeticity may not be worth it. Teams often triage: hermeticity matters most for expensive, frequently-invalidated tasks sitting high in the dependency graph, where staleness or wasted recomputation has the biggest blast radius.
  • Does hermeticity only matter for caching, or does it have value even without a cache?
    It has independent value for reproducibility and debuggability — a hermetic build lets you say 'this artifact was built from exactly this source, nothing else,' which matters for security attestation, bisecting regressions, and giving any two engineers on any two machines confidence they'll get the same result. Caching is the performance payoff, but hermeticity's correctness guarantee stands on its own.

Like a chemistry experiment run in an open room versus a sealed clean room — in the open room, dust, humidity, or someone bumping the table can quietly change the result without being recorded as a variable; the sealed room guarantees the only things that could have affected the outcome are the ones you deliberately put in.

saying these in an interview costs you the question

  • Conflates 'hermetic' with just 'the tests pass'
  • Doesn't distinguish hidden-input staleness from output non-determinism as two separate failure modes
  • Thinks sandboxing is only for security, not for proving input completeness
  • No example of a common non-hermetic pattern (timestamps, random IDs, unordered iteration)

context