skip to content

A web app's build places every dependency from node_modules into one large shared vendor chunk, separate from the application code. What is the caching argument for that layout, and what tends to break it on a real project?

level: middleimportance: should knowfreq 54%

answer

  1. hashed filenames, long cache lifetime
  2. dependencies churn slower than app code
  3. one bump rewrites the whole chunk
  4. node_modules is a source, not a need
  5. group by change cadence instead

basics

~20 s

The argument is cache lifetime: dependencies change less often than app code, so a separate vendor chunk stays cached across deploys. It breaks because one dependency bump rewrites the whole chunk's content hash, invalidating megabytes for every returning user.

solid answer

~50 s

The theory is that application code changes every deploy while dependencies change monthly, so isolating dependencies gives returning visitors a large file they rarely have to re-download. With content-hashed filenames and long cache lifetimes, a normal deploy then touches only the small app chunks. In practice the single-vendor layout fails on two fronts. First, granularity: the chunk's hash is a function of all its contents, so upgrading one small library invalidates the entire vendor file — and on a project with active dependency automation that happens weekly. Second, first-visit cost: everything every route depends on is now in one blocking chunk, so a visitor who only sees the login page still downloads the charting library. The better layout groups dependencies by change cadence and by who needs them — a stable core-framework chunk, a chunk for large route-specific libraries loaded with those routes — rather than one bucket defined by "came from node_modules".

code

bash · 6 lines
bash
# Which files does a returning user actually re-download after this deploy?
git stash && npm run build && cp -r dist /tmp/build-prev
git stash pop && npm run build

# Emitted filenames that differ = the invalidation blast radius
diff <(ls /tmp/build-prev/assets | sort) <(ls dist/assets | sort)

go deeper

for a junior

Know that build output filenames contain a content hash so files can be cached for a long time, and that a vendor chunk is meant to hold code that changes less often than your own.

for a middle

Explain that the hash covers the whole chunk, so any single dependency upgrade invalidates all of it, and that lumping route-specific libraries into a global chunk penalises first-time visitors.

for a senior

Demonstrate that you validate the layout rather than assume it: diff emitted filenames between two builds to see the real invalidation blast radius, and re-group dependencies by update cadence and by which routes need them.

for a principal

Own the tradeoff curve between cache granularity and request/compression overhead, and tie the chunking policy to deploy frequency and dependency-update automation so the strategy still holds as the team's release cadence changes.

## The idea behind a vendor chunk Production builds emit filenames containing a hash of the file's contents (`main.9f3c1a.js`). That lets the server serve them with a very long cache lifetime and treat a new deploy as a new filename rather than a cache invalidation problem. The moment content changes, the URL changes; the moment it does not, returning users reuse the copy they already have. A vendor chunk is an attempt to maximise the second case. Application code churns constantly — every feature branch merged rewrites it. Dependencies churn slowly. If both live in one file, the file's hash changes on every deploy and every returning user re-downloads the framework they already had. Separating them means a routine deploy changes only the small application chunks. When dependencies really are stable and deploys are frequent, this is a genuine win measured on **repeat** visits: the returning-visitor byte count drops to the size of the app chunks. ## Why the single-bucket version disappoints **Hashes are all-or-nothing.** The chunk's hash covers its entire contents. Bump one utility library by a patch version and the whole vendor chunk — framework, date library, charting, everything — gets a new filename and is re-downloaded by every user. Projects with automated dependency updates merge such bumps constantly, so the "rarely changes" premise quietly stops holding. The bigger the chunk, the more expensive each of those invalidations is. **It fights code splitting.** "Everything from node_modules" is a rule about where a file came from, not about when it is needed. A rich-text editor used only by the admin route, a charting library used on one dashboard, and an internationalisation bundle for a locale most users never select all end up in the chunk that every visitor must download before anything renders. This makes first-visit performance worse in exchange for a repeat-visit gain — usually a bad trade, because first-time visitors have nothing cached and are the ones deciding whether to stay. **Cache hit rates are lower than assumed.** HTTP cache entries are evicted under storage pressure, users arrive from different origins and devices, and browser storage partitioning means a shared library fetched on another site does not help yours. Sizing a strategy around a hypothetical warm cache overestimates the benefit. ## A better layout Stop using "came from node_modules" as the grouping key. Group by **change cadence** and **who needs it**: - **A framework core chunk** — the handful of packages that every route needs and that you upgrade deliberately, a few times a year. This is the chunk that genuinely earns a long cache life. - **Route-scoped dependency chunks** — a heavy library used by one route belongs with that route's dynamic import, not in a global chunk. It is then downloaded only by users who reach that route, and its churn cannot invalidate anything else. - **A shared chunk for genuinely common non-framework code** — extracted when two or more route chunks need it, and only when it is big enough to matter. Most bundlers expose controls for this grouping; the strategic decision — which packages belong in which bucket — is yours regardless of tool. ## How to check whether it is working Do not reason about this in the abstract; measure it: - **Simulate a deploy.** Build the current commit, build the previous one, and diff the emitted filenames. The set of files whose hash changed is exactly what a returning user re-downloads. If a one-line copy change rewrites 400 KB, the layout is wrong. - **Check first-visit weight per route.** For the landing route, add up entry plus vendor plus route chunk. If that number includes libraries the landing route never calls, the vendor bucket is doing harm. - **Watch for hash instability.** Chunks whose hash changes even when their contents did not (from build-order or module-id churn) silently destroy cache hits; if you see it, the fix is deterministic module identifiers rather than more splitting. ## The tradeoff to state out loud Every chunk boundary trades cache granularity against request count and compression efficiency. One chunk per package maximises cache reuse and produces hundreds of tiny, poorly-compressed files. One chunk for everything minimises requests and maximises invalidation blast radius. The answer is a small number of buckets chosen by update cadence, validated by diffing hashes across two real builds.

  • Why do content-hashed filenames matter for this strategy at all?
    They decouple deploys from cache invalidation. Because the URL changes whenever the bytes change, you can serve every chunk with a very long max-age and never worry about a stale copy: a new build simply references new filenames from the HTML. Without hashing you would need short lifetimes or revalidation, and the whole "keep vendor cached across deploys" argument collapses.
  • When is a single vendor chunk actually the right call?
    On a small app whose dependencies are a stable framework core that every route needs anyway, and which deploys infrequently. There the grouping key "from node_modules" happens to coincide with "needed everywhere and rarely changes", so the layout is correct by accident. The moment a heavy route-specific library lands in it, that stops being true.
  • A chunk's hash changes on every build even though its source did not. What is going on?
    Usually non-deterministic module identifiers or build ordering leaking into the emitted output, so byte-identical logic produces different bytes. It destroys cache hits invisibly — users re-download unchanged code every deploy. Diagnose by building the same commit twice and diffing the output, then switch the build to deterministic module ids or stable chunk naming.
  • How does this interact with how often the team deploys?
    Directly. The value of isolating stable code scales with deploy frequency: a team shipping ten times a day gets a large repeat-visit saving from a well-scoped framework chunk, while a team deploying monthly gains little and should optimise first-visit weight instead. Deploy cadence should be an explicit input to the chunking decision.

saying these in an interview costs you the question

  • Assumes vendor code never changes, so the chunk stays cached forever
  • Groups by node_modules rather than by what a route needs
  • Ignores that one dependency bump invalidates the entire chunk
  • Optimizes repeat visits while making the first visit heavier
  • Believes a shared CDN copy of a library is cached across sites

context