skip to content

A build splits dependencies into a separate `vendor` chunk with a content-hashed filename, but that hash changes on nearly every deploy even when no dependency was upgraded. Why does that undermine the caching strategy, and what usually causes it?

level: middleimportance: should knowfreq 42%

answer

  1. split only pays if the URL survives
  2. hash is honest; bytes really changed
  3. ids assigned by build order shift
  4. chunk map inside the chunk churns everything
  5. strip timestamps from long-lived files

basics

~20 s

A hash that changes without a real content change destroys the point of splitting out vendor code: returning users re-download it every release. Usual causes are renumbered internal module identifiers, embedded build metadata, or the chunk-URL map living inside the chunk.

solid answer

~50 s

Separating dependencies into their own file is only worth doing if that file's URL survives releases — that is the entire return on the extra request. If its hash moves every deploy, returning visitors re-download the largest, least-changing part of the app every time, and you have paid for a split that buys nothing. The hash is honest, so the bytes really are different; the job is to find out why. The classic causes are internal module or chunk identifiers that are assigned by build order and shift whenever application code moves, build metadata baked into the file such as a timestamp or commit id in a banner, the loader's chunk-name-to-hashed-URL map being emitted inside the vendor chunk so it changes whenever any other chunk's hash changes, and a grouping rule loose enough that application-adjacent code lands in the vendor bucket. The fix is to make those inputs deterministic and to move the chunk map into its own tiny file.

go deeper

for a junior

Know that a hashed filename only helps a returning visitor if the file keeps the same name between releases, and that a different hash always means the file's bytes really are different.

for a middle

Be able to name concrete causes of spurious churn — order-dependent internal ids, embedded build metadata, the chunk map emitted inside the chunk — and say how you would confirm each by diffing two builds.

for a senior

Show that you would quantify it before acting: chunk size times returning-user rate times deploy frequency, verified in field data, and that you would group dependencies by how often they change rather than by where they came from.

for a principal

Own the guarantee itself — reproducible builds, per-file volatility tiers, and a check that catches hash churn on long-lived artifacts in CI, so that the caching promise does not quietly decay release by release.

## What a stable hash is actually worth Splitting dependency code away from application code costs something: an extra request, an extra entry in the build, more moving parts. It is justified by one payoff — dependency code changes rarely and application code changes constantly, so keeping them in separate files means a normal release invalidates only the small, volatile half. Returning users re-download tens of kilobytes instead of hundreds. That payoff is expressed entirely through the filename. If `vendor.<hash>.js` gets a new hash on every deploy, its URL is new every deploy, every cache misses, and every returning visitor downloads the whole dependency payload again. You now have the cost of the split with none of the benefit, and — worse — the metric looks fine, because the file is being served correctly and quickly. It just should not have been requested at all. So when you see the hash churn, do not reach for the hash function. A content hash is honest by construction: it changed because the bytes changed. The question is why bytes you believe to be unchanged are, in fact, different. ## Why the bytes differ when the dependencies did not **Non-deterministic internal identifiers.** Bundlers rewrite each module to an internal id and each chunk to an internal id. If those ids are assigned in traversal or discovery order, adding one import anywhere in application code renumbers modules and shifts ids inside the dependency chunk too. Nothing about the dependency changed; its numbering did. Bundlers grew deterministic identifier modes — ids derived from a stable property such as module path rather than from build order — precisely because this was the single most common cause of spurious churn. **Build metadata baked into the file.** A banner comment carrying a build timestamp, a commit sha, a version string, or a source-map comment naming a file whose own name changed, all put per-release bytes inside a file that is otherwise identical. Any one of them is enough to move the hash. **The chunk map living in the wrong file.** Something has to know that the chunk called `route-detail` currently lives at `route-detail.3c19aa7e.js`. That mapping changes whenever *any* chunk's hash changes, which is to say every deploy. If the loader runtime holding that map is emitted inside the vendor chunk, the vendor chunk inherits the churn of the entire build. Extracting the runtime and its map into its own tiny file — a file everyone expects to change every deploy — isolates the volatility where it belongs. **A grouping rule that is too broad.** "Everything from the dependency directory" sounds stable, but it sweeps in code whose presence depends on what application code imports. A new import can pull a new dependency into the bucket, or dead-code elimination can drop one, and the chunk's contents genuinely change. **Toolchain drift.** A bundler or minifier upgrade legitimately changes output bytes. This one is real churn, it is a one-off per upgrade, and it is not a bug. ## How to confirm which one it is Build the same commit twice and diff the outputs: if two builds of *identical source* produce different bytes, you have non-determinism (metadata, timestamps, ordering) rather than a content change. Then build two adjacent commits and diff those: if the only differences are renumbered identifiers or a banner line, the churn is spurious. Reading the first and last few hundred bytes of the chunk usually finds a banner or an embedded map in seconds. ## What to do, and when not to bother Make identifiers deterministic, strip per-build metadata from long-lived files, emit the chunk map as its own artifact, and define the dependency grouping by stability rather than by directory — the framework core that changes twice a year does not belong in the same file as a charting library you upgrade monthly. And size the problem before you fix it. Churn on a 20 kB chunk costs a returning visitor almost nothing. Churn on a 400 kB chunk, for a site that deploys daily and whose users return weekly, is a repeated several-hundred-kilobyte download for the entire returning population — the kind of number worth putting in a ticket. The decision is bytes times returning-user rate times deploy frequency, not tidiness.

  • How would you prove the churn is spurious rather than a real content change?
    Build the same commit twice: identical source producing different bytes proves non-determinism — a timestamp, a commit id, or ordering-dependent output. Then diff two adjacent releases' chunks; if the only differences are renumbered internal identifiers or a banner line, nothing semantically changed. Both checks take minutes and tell you which class of cause to chase before touching any build settings.
  • Where does the mapping from a chunk's logical name to its hashed URL live, and why does that placement matter?
    It lives in the loader runtime the build emits, because something must resolve `route-detail` to `route-detail.3c19aa7e.js` at runtime. That map changes on every deploy by definition. If it is emitted inside a long-lived chunk, it drags that chunk's hash along with it; giving it its own small file confines the churn to a file nobody expected to cache.
  • Is chunk hash churn always worth fixing?
    No — weigh chunk size against returning-user rate and deploy frequency. A 20 kB chunk churning daily is noise. A 400 kB chunk churning daily, on a site whose users return weekly, means the entire returning population re-downloads it repeatedly for no reason. Measure the bytes first; the fix is cheap but not free, and the payoff scales entirely with those three numbers.

saying these in an interview costs you the question

  • Thinks content hashes include the build timestamp by design
  • Says the hash is unstable, so hashing cannot be trusted
  • Assumes any deploy must invalidate every chunk anyway
  • Believes revalidation makes the re-download nearly free
  • Treats one big vendor bucket as automatically stable

context